Prefix caching
course 8 of 22
lesson 06 · lessons/06-prefix-caching
35 min About $0.03 Mac, then Fireworks
What you will do and why
Agents resend the same long opening with every request. When the server reuses it, replies start sooner and, on Fireworks, cost less.
Why it matters: In the lesson’s example printout, call 1 waits 740 ms (thousandths of a second) for its first word. With the 3,000-token opening reused (a token is a word piece, about three quarters of a word), the next calls wait about 36 ms: about 20 times less.
You are done when: You have three things. On llama.cpp: the wait for the first word with the opening first and with the time first. On Fireworks: the opening-first wait, and a usage line where cached_tokens covers most of prompt_tokens.
Prefix Caching · 2:37
Download mp4 (9.6 MB)Chapters
In this video Why agent traffic is mostly repeats, how servers reuse the repeated part, and what that does to the wait and the bill.
3 key points
Most of what an agent sends is the same opening again.
In the video’s 20-call example, 66,500 of the 72,000 tokens sent repeat the opening: 92%. The video shows 93% because it rounds 66,500 up to 67,000.
Reusing the opening cuts both the wait and the bill.
In the video’s session, 95 of 100 calls see their first word within 0.35 seconds instead of 1.9. The input costs about a fifth of a cent instead of 2 cents.
Put the parts that never change first, or there is nothing to reuse.
One changing line at the top, such as the time, makes every call new: calls 2 to 12 took 712 ms instead of 36 ms in the lesson’s sample.
In plain words
Section titled “In plain words”An agent sends the same long opening with every request: its instructions, its tools (actions it may ask for, such as a refund) and its rules. Only the last line, the question, changes. With prefix caching the server keeps its notes on that opening (the KV cache from Day 4). It then reads only the new part, so the reply starts much sooner.
Picture it
A bookmark: if the first 3,000 words of a book are the same as last time, you start reading at word 3,001. It only works if every word before the bookmark matches. Change one word on page 1 and you start again from the top.
With real numbersthe lesson’s script and example printout (lesson 06), and the Prefix Caching video’s prices
- The shared opening: 3,028 tokens by the script’s estimate (12,115 characters ÷ 4; about 2,300 words), sent with each of 12 different questions.
- Opening first, call 1 (nothing saved yet): 740 ms before the first word.
- Calls 2 to 12 (opening reused): 36 ms, the middle value. 740 ÷ 36 = about 20 times sooner.
- A line with the time of the request (and the question) added above the opening: call 1 takes 744 ms, calls 2 to 12 about 712 ms. 744 ÷ 712 = 1.04, so almost no gain.
- On Fireworks the reused part also costs less. The video’s example prices: $0.30 per million tokens, or $0.006 cached. $0.006 ÷ $0.30 = 2% of the price.
Words to know
- Prefix
- The start of a prompt, up to the first token that differs from last time. Example: the shared 3,028-token opening.
- TTFT (time to first token)
- How long until the first token (word piece) of the reply appears. Example: 740 ms for call 1, 36 ms for the calls after it.
- Agent
- A program that calls a model many times, using tools, to finish a task. Example: the lesson’s support agent with 5 tools.
- Cached tokens
- Prompt tokens the server reused instead of reading again. Fireworks reports them as
cached_tokensand bills them at a discount.
From the lesson
lessons/06-prefix-caching/README.md
What and why
Section titled “What and why”An agent sends the same 3,000-token preamble (instructions, tool definitions, policy) on every call and changes only the last line. Without a prefix cache the server re-reads that preamble every time. With one, it keeps the KV cache from last time and reads only the new part.
It works like a bookmark: if the first 3,000 words are identical, start reading at word 3,001. It only works if the text is identical from the very first token. Put a timestamp at the top and every call looks new.
stable-first [■■■■■■■■■■ preamble (cached) ■■■■■■■■■■][question] → TTFT ~20× lowervariable-first [time+question][■■■■■■■ preamble ■■■■■■■] → nothing reusable| where | name | what it saves |
|---|---|---|
| vLLM | automatic prefix caching (block hashes) | TTFT, GPU time |
| SGLang | RadixAttention (a tree of shared prefixes) | TTFT, very good for branching agents |
| llama.cpp | slot cache reuse | TTFT |
| Fireworks | prompt caching, on by default, with session affinity via user |
TTFT and the bill: cached input is heavily discounted |
Read the code first
Section titled “Read the code first”prefix_bench.py → messages()puts the good pattern and the anti-pattern side by side.extra = {"user": ...}shows session affinity: the same user stays on the same replica, and so the same cache.felab/mock_server.py → uncached_tokens()is a 15-line model of block-hash prefix caching.
cd lessons/06-prefix-cachingpython prefix_bench.py --target mockpython prefix_bench.py --target mock --order variable-first
bash ../02-three-local-servers/serve_llamacpp.sh # other terminal. Add Q4_K_M 16384 4 after .sh (see the note below)python prefix_bench.py --target llamacpppython prefix_bench.py --target llamacpp --order variable-first
bash mlx_prompt_cache.sh # the MLX by-hand version
python prefix_bench.py --target fireworks --usage # Fireworks: cached_tokens in the usage blockWhat each command does
python prefix_bench.py --target mockSends 12 different questions, each after the same 3,028-token opening, to the practice server (free, offline). Look for call 1 marked (cold) and calls 2 to 12 far faster. The practice server reads 2,500 tokens a second, so call 1 takes about 1.2 seconds (3,028 ÷ 2,500), more than the lesson’s 740 ms. Calls 2 to 12 take about 30 ms (15 ms plus the few new tokens), so its speed-up comes out near 40.
python prefix_bench.py --target mock --order variable-firstThe same 12 questions, but with the time of the request (and the question) at the very top. Look for every call about as slow as call 1. speedup_x (call 1’s time divided by the middle of the others) should be about 1.0.
bash ../02-three-local-servers/serve_llamacpp.shStarts the real llama.cpp server from Day 2. Run it in a second terminal, from this lesson’s folder (
cd lessons/06-prefix-cachinginside03-labs). AddQ4_K_M 16384 4after.sh: the 4-bit model, room for 16,384 tokens, 4 seats (4,096 each). Without them it has 8,192 tokens over 4 seats, 2,048 each, too few for the 3,000-token opening (Day 2’s trap). Look for the line saying it is listening.python prefix_bench.py --target llamacppThe same test on the real model on your Mac. Look for a slow call 1 and calls 2 to 12 far faster. On Day 4’s example Mac, which reads about 1,050 tokens a second, call 1 takes roughly 3 seconds (3,028 ÷ 1,050). Write down first_ms (call 1) and rest_p50_ms (the middle of calls 2 to 12).
python prefix_bench.py --target llamacpp --order variable-firstThe time at the top again, on the real model. Look for every call about as slow as call 1 and a speedup_x near 1.0: llama.cpp only reuses text that matches from the very first token.
bash mlx_prompt_cache.shThe same idea done by hand with MLX (Apple’s own engine). Step 1 answers cold, reading everything. Step 2 reads the opening once and saves its notes to a file. Step 3 answers from that file. Look for the
Prompt:line: about 3,000 tokens in step 1, only the short question in step 3.python prefix_bench.py --target fireworks --usageThe same test on Fireworks (paid, about $0.03), then one extra call that prints the token counts. Every call carries the same session name (the
uservalue), so all reach the same copy of the model. Look for theusage:line near the end, just above theCost checkline: cached_tokens should cover most of prompt_tokens.
What you should see
Section titled “What you should see”How to read it
Each run prints one row, after a target column naming the server: stable-first puts the opening at the top, variable-first the time. first_ms is call 1, rest_p50_ms the middle of calls 2 to 12, and speedup_x the first divided by the second. In the lesson’s sample, stable-first falls from 740 to 36 ms (about 20 times) and variable-first stays slow, 744 then 712 ms (about 1).
order first_ms rest_p50_ms speedup_xstable-first 740.0 36.0 20.6variable-first 744.0 712.0 1.0 ← one changed token at the top kills itWith --usage on Fireworks, cached_tokens should cover most of prompt_tokens.
Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
A customer’s agent puts the current time (Current time: …) at the very top of its system prompt. The system prompt is the fixed instructions sent at the start of every request. What do you tell them?Show answerHide
In plain words
Move the time to the end. Put everything that never changes first and anything that changes last, so every request starts with the same text the server can reuse.
Picture it
A bookmark only helps if every page before it is the same as last time. Write the current time on page 1 and the bookmark is useless: you start the book again on every visit.
With real numbersthe lesson’s script and example printout, lesson 06
- Opening first, nothing that changes above it (the lesson’s stable-first run, question at the end): call 1 takes 740 ms, the rest 36 ms (the middle value). 740 ÷ 36 = about 20 times sooner.
- A line with the time and the question added above the opening (variable first): call 1 takes 744 ms, calls 2 to 12 about 712 ms. 744 ÷ 712 = 1.04, so almost no gain.
- The only difference between the two runs: the time-first run adds that one line above the opening.
- The fix changes how the prompt is built. It needs no new server and no new setting.
Words to know
- System prompt
- The fixed instructions at the start of every request. Example: the script’s 3,028-token support-agent opening.
- Stable first, variable last
- Build the prompt with the parts that never change at the top. Example: rules and tools first, the time and the question last.
- Prefix
- The start of a prompt, up to the first token that differs from last time. Example: a new timestamp at the top leaves no shared prefix.
Go deeper: the engineer version
The kit's question
The customer’s agent puts Current time: … at the top of the system prompt. What do you tell them?
The kit's answer
Move it to the end. Stable content first, variable content last.
More detail: Prefix caches match token by token from the start. llama.cpp and SGLang reuse the longest identical prefix; vLLM and the practice server hash fixed-size blocks, each key covering everything before it (64-token blocks in mock_server.py), so one changed token invalidates every block after it. Per-request values such as the time or the user’s name belong after the stable part, ideally in the last user message.
A customer runs several identical copies of the model (replicas) behind one address. Why does it matter to send a user value (a session name) with each request, so each session keeps going to the same copy (session affinity)?Show answerHide
In plain words
Each copy keeps its own saved notes, and the copies do not share them. A request that lands on a copy that has not seen this conversation starts cold: that copy must read the conversation again.
Picture it
A chain of cafes where each barista remembers your usual order. Go back to the same barista and your coffee comes fast. Walk up to a new one and you explain it all again.
With real numberslesson 06’s example printout and Day 3’s sizing question
- Warm, in the lesson’s sample (the copy has already read the text): 36 ms to the first word.
- Cold (the copy must read it all): 740 ms, about 20 times longer.
- Day 3’s example service had 18 copies. Sent at random, a conversation’s next request would reach the copy holding its history only 1 time in 18.
- The script sends
user=lesson-06-sessionon every Fireworks call, so they all reach the same copy.
Words to know
- Replica
- One complete running copy of the server and model. Example: 18 copies for 120 people on Day 3.
- Session affinity
- Sending every request of one session to the same copy, so it finds its saved notes. Example: the
userfield. - Cold and warm
- Cold: nothing saved yet, so everything is read; warm: the saved notes are reused. Example: 740 ms cold, 36 ms warm.
Go deeper: the engineer version
The kit's question
Why does user / session affinity matter on a multi-replica deployment?
The kit's answer
Each replica holds its own cache, so a request routed to a different replica starts cold.
More detail: Each replica holds its own KV cache in its own GPU memory; nothing is shared between replicas. A shared system prompt warms on each replica after its first request there, but a session’s own history (earlier turns, tool results) is cached only where it was served. Without affinity those turns land on different replicas, each pays full prefill, and the hit rate falls. Caches also expire, so the first call after an idle spell is slow (the Prefix Caching video).
What is the business case for prefix caching: the argument, in money and speed, you would give the customer?Show answerHide
In plain words
Most of an agent’s input is the same opening again, so most of the input bill moves to the much cheaper cached price. The reply also starts several times sooner.
Picture it
A print shop charges full price to set up a page and a small fee for every extra copy. An agent keeps ordering the same page; with caching it pays for the setup once and copy prices after that.
With real numbersthe Prefix Caching video’s example prices, lesson 06’s answer (90% repeated) and the Day 10 memo script
- An agent whose input is 90% repeated opening, per 1 million input tokens:
- Without caching: 1,000,000 × $0.30 per million = $0.30.
- With caching: 100,000 new tokens × $0.30 per million = $0.03, plus 900,000 cached × $0.006 per million = $0.0054.
- Total $0.0354 instead of $0.30: about 12% of the bill, so 88% saved.
- At 50% off instead (what the Day 10 memo assumes until you check): $0.03 + 900,000 × $0.15 per million = $0.165, so 45% saved.
- Lesson example printout: 740 ms becomes 36 ms (the middle call), about 20 times sooner.
- The video’s session: 1.9 s becomes 0.35 s (the time 95 of 100 calls stay under), about 5 times sooner.
Words to know
- Business case
- The argument, in money and results, for making a change. Example: 88% off the input bill at the video’s prices.
- Input tokens
- The tokens you send to the model; providers bill them per million. Example: $0.30 per million in the video’s example.
- Cached tokens
- Input tokens the server reused; Fireworks bills them at a discount. Example: $0.006 instead of $0.30 per million.
- Order of magnitude
- About 10 times. Example: 740 ms down to 36 ms is more than one order of magnitude.
Go deeper: the engineer version
The kit's question
What is the business case?
The kit's answer
For an agent that is 90% repeated context, most of the input bill becomes cached-token pricing, and TTFT drops by an order of magnitude.
More detail: Cached-token pricing is set per model. The video’s $0.30 and $0.006 per million match the lab book’s figures for DeepSeek V4.1 Flash on Fireworks (98% off); the Day 10 memo script assumes only 50% off until you check the price page. Your own run uses the script’s default model, gpt-oss-120b, which has its own prices; read them off the price page before quoting a saving. The TTFT gain depends on the workload: 20x (call 1 against the median of calls 2 to 12) in the lesson’s sample, 5.4x at p95 in the video’s measured session (1.9 s to 0.35 s). Quote the customer’s own before-and-after with the cache hit rate; a cache that is never hit only takes memory away from concurrency (the video).
Explain what you learned
Section titled “Explain what you learned”Question Our agent bill is too high. What would you do first?
One clear answer
Their system prompt is identical on every call, so most of the input bill is cacheable. Stable first, variable last, session affinity on, and I’ll show the cached-token count in the usage block.
What this means
- “Their system prompt is identical on every call”: The customer’s fixed opening (instructions, tools, rules) is the same text on every request. In the lesson it is 3,028 tokens, sent with every question.
- “so most of the input bill is cacheable”: Most of what they pay for the model to read can be reused and billed at the cached price. In the video’s 20-call session, 92% of the input tokens are repeats.
- “Stable first, variable last”: Put what never changes at the top and what changes (the time, the question) at the bottom, so the opening matches every time. With the time at the top, calls 2 to 12 took 712 ms instead of 36 ms in the lesson’s sample.
- “session affinity on”: Send a session id (the
userfield) so each conversation keeps going to the same copy of the model, the one holding its saved notes. - “and I’ll show the cached-token count in the usage block”: I prove it with the counts Fireworks returns with each answer:
cached_tokensnext toprompt_tokens. Most of the prompt should be cached.
Your numbersSaved on this device and collected in the Day 5 wrap-up.
Hint: From python prefix_bench.py --target llamacpp: the first_ms and rest_p50_ms columns. Write both, for example “740 / 36”.
Hint: The speedup_x column of the --order variable-first run. Expect about 1.0.
Hint: From python prefix_bench.py --target fireworks --usage: the first_ms and rest_p50_ms columns. Write both.
Hint: From the usage: line near the end, just above the Cost check line. Write both numbers, cached first.
Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.
Done when
Section titled “Done when”You have three things. On llama.cpp: the wait for the first word with the opening first and with the time first. On Fireworks: the opening-first wait, and a usage line where cached_tokens covers most of prompt_tokens.
Stuck?
Section titled “Stuck?”Connection refused.- That server is not running.
make check(from the03-labsfolder) shows what is up. Start the practice server withmake mock &, or start llama.cpp as in the run notes and wait until it says it is listening. - Opening first on the practice server shows a speed-up near 1, and call 1 is already fast.
- The practice server still remembers the opening from an earlier run: this command run twice, or
make smokewhile it was up. Restart it:pkill -f felab.mock_server, thenmake mock &from the03-labsfolder, and run again. No such file or directorywhen you start llama.cpp.- The command’s path starts from this lesson’s folder. From the
03-labsfolder runcd lessons/06-prefix-caching, then start it again. - llama.cpp returns an error about the context size, or its opening-first run shows no speed-up.
- Each of its 4 seats holds only 2,048 tokens by default, less than the 3,000-token opening. Stop it (Ctrl+C) and, from this lesson’s folder, start it with
bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 4(4,096 tokens per seat). FIREWORKS_API_KEY is not set.- Put your key in the
.envfile in03-labs, as on Day 1, or export it in this terminal. cached_tokens=0orcached_tokens=None.- Run the Fireworks command again right away: caches expire when idle (the Prefix Caching video). If it still says None, this model does not report it; pass another with
--model. - A Fireworks model id is not found.
- Model ids change. Copy a current one from the Fireworks model library and pass it with
--model. mlx_lm.generate: command not found.- The Python environment from Day 1 is not active. From the
03-labsfolder runsource .venv/bin/activate, then try again.
Go deeper: the lab book's local lab for this step Go deeper: the lab book's Fireworks lab for this step All fixes
Code in this step
Section titled “Code in this step”mlx_prompt_cache.sh Bash · 25 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 06 · MLX variant — the same idea, done by hand.## mlx_lm.cache_prompt runs prefill ONCE on the preamble and saves the KV cache# to disk. mlx_lm.generate then loads that file and only prefills the new question.# This is literally what a server-side prefix cache does, made visible.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"MODEL=${MLX_MODEL:-mlx-community/Meta-Llama-3.1-8B-Instruct-4bit}
# Dump the same preamble prefix_bench.py usespython -c "from prefix_bench import PREAMBLE; print(PREAMBLE)" > preamble.txt
echo "▸ 1. cold: prefill the ~3k-token preamble + question (watch 'Prompt: … tokens-per-sec')"time mlx_lm.generate --model "$MODEL" --max-tokens 40 \ --prompt "$(cat preamble.txt) Question: How do I rotate my API key?"
echo "▸ 2. cache the preamble once"mlx_lm.cache_prompt --model "$MODEL" --prompt "$(cat preamble.txt)" --prompt-cache-file preamble.safetensors
echo "▸ 3. warm: only the question is prefilled"time mlx_lm.generate --prompt-cache-file preamble.safetensors --max-tokens 40 \ --prompt " Question: How do I rotate my API key?"prefix_bench.py Python · 88 lines
"""Lesson 06 — prefix caching on an agent-shaped workload.
Agents resend the same big preamble (instructions + tool definitions + docs) onevery call and change only the last few lines. If the server remembers the KVcache for that preamble, it skips re-reading it: TTFT collapses and, on Fireworks,cached input tokens are billed at a steep discount.
This script sends 12 different questions two ways:
--order stable-first [ 3k-token preamble ][ question ] ← cache-friendly --order variable-first [ question + timestamp ][ preamble ] ← one changed token early invalidates everything after it
python prefix_bench.py --target mock python prefix_bench.py --target mock --order variable-first python prefix_bench.py --target llamacpp # llama-server reuses slot caches python prefix_bench.py --target fireworks --usage # ~$0.03; prints cached_tokens"""import argparseimport datetime as dt
from felab import add_target_args, banner, client, percentile, record, resolve, stream_once, table
TOOLS = "\n".join( f"- tool `{name}`: {desc}. Arguments: JSON object with fields id (string), limit (int), filters (object)." for name, desc in [("search_tickets", "full-text search over support tickets"), ("get_invoice", "fetch an invoice by id"), ("refund", "issue a refund up to the limit"), ("status_page", "read current incident status"), ("escalate", "page the on-call engineer")])POLICY = ("Refunds over $500 need manager approval. Outages affecting more than 5% of users are SEV-2. " "Never reveal internal ticket ids. Answer in under 80 words. ") * 75PREAMBLE = f"You are the support agent for Acme Cloud.\n\n## Tools\n{TOOLS}\n\n## Policy\n{POLICY}"
QUESTIONS = [ "A customer was charged twice this month. What do I do?", "Is there an outage in eu-west right now?", "How do I rotate my API key?", "Customer wants a $900 refund — can I approve it?", "What counts as a SEV-2?", "Can I share the internal ticket id with the customer?", "Which tool finds old tickets about SSO?", "Latency spiked at 09:00 — who do I page?", "The customer asks for their March invoice.", "How long should my answer be?", "Card declined on renewal, launch is tomorrow.", "Summarise the refund policy in one line.",]
def messages(q: str, order: str) -> list[dict]: if order == "stable-first": return [{"role": "system", "content": PREAMBLE}, {"role": "user", "content": q}] # anti-pattern: something that changes every call goes FIRST stamp = dt.datetime.now().isoformat(timespec="microseconds") return [{"role": "system", "content": f"Request time {stamp}. User asks: {q}\n\n{PREAMBLE}"}, {"role": "user", "content": q}]
def main() -> None: ap = add_target_args(argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)) ap.add_argument("--order", choices=["stable-first", "variable-first"], default="stable-first") ap.add_argument("--usage", action="store_true", help="one extra non-streamed call to print cached_tokens") a = ap.parse_args() t = resolve(a) banner(t) cli = client(t)
# `user` = session affinity on Fireworks: keeps you on the replica that holds your cache. extra = {"user": "lesson-06-session"} if t.is_paid else {} print(f"preamble ≈ {len(PREAMBLE) // 4:,} tokens · order = {a.order}\n")
ttft = [] for i, q in enumerate(QUESTIONS): s = stream_once(cli, t.model, messages(q, a.order), max_tokens=48, **extra) ttft.append(s.ttft_ms) print(f" call {i + 1:2d} TTFT {s.ttft_ms:7.0f} ms {'(cold)' if i == 0 else ''}")
row = {"target": t.name, "order": a.order, "first_ms": ttft[0], "rest_p50_ms": percentile(ttft[1:], 50), "speedup_x": ttft[0] / max(1e-6, percentile(ttft[1:], 50))} print("\n" + table([row])) record("06-prefix", row)
if a.usage: r = cli.chat.completions.create(model=t.model, max_tokens=8, messages=messages("One more question.", a.order), **extra) u = r.usage cached = getattr(getattr(u, "prompt_tokens_details", None), "cached_tokens", None) print(f"\nusage: prompt_tokens={u.prompt_tokens} cached_tokens={cached}") print("Cost check: cached input is billed at a fraction of normal input — see the model's pricing page.")
if __name__ == "__main__": main()