QuestionStep 1 · Prefix caching
A customer’s agent puts the current time (Current time: …) at the very top of its system prompt. The system prompt is the fixed instructions sent at the start of every request. What do you tell them?
Show answerHide answer
In plain words
Move the time to the end. Put everything that never changes first and anything that changes last, so every request starts with the same text the server can reuse.
Picture it
A bookmark only helps if every page before it is the same as last time. Write the current time on page 1 and the bookmark is useless: you start the book again on every visit.
With real numbersthe lesson’s script and example printout, lesson 06
- Opening first, nothing that changes above it (the lesson’s stable-first run, question at the end): call 1 takes 740 ms, the rest 36 ms (the middle value). 740 ÷ 36 = about 20 times sooner.
- A line with the time and the question added above the opening (variable first): call 1 takes 744 ms, calls 2 to 12 about 712 ms. 744 ÷ 712 = 1.04, so almost no gain.
- The only difference between the two runs: the time-first run adds that one line above the opening.
- The fix changes how the prompt is built. It needs no new server and no new setting.
Words to know
- System prompt
- The fixed instructions at the start of every request. Example: the script’s 3,028-token support-agent opening.
- Stable first, variable last
- Build the prompt with the parts that never change at the top. Example: rules and tools first, the time and the question last.
- Prefix
- The start of a prompt, up to the first token that differs from last time. Example: a new timestamp at the top leaves no shared prefix.
Go deeper: the engineer version
The kit's question
The customer’s agent puts Current time: … at the top of the system prompt. What do you tell them?
The kit's answer
Move it to the end. Stable content first, variable content last.
More detail: Prefix caches match token by token from the start. llama.cpp and SGLang reuse the longest identical prefix; vLLM and the practice server hash fixed-size blocks, each key covering everything before it (64-token blocks in mock_server.py), so one changed token invalidates every block after it. Per-request values such as the time or the user’s name belong after the stable part, ideally in the last user message.
How did you do?