Skip to content

Prefix caching

course 8 of 22

lesson 06 · lessons/06-prefix-caching

35 min About $0.03 Mac, then Fireworks

What you will do and why

Agents resend the same long opening with every request. When the server reuses it, replies start sooner and, on Fireworks, cost less.

Why it matters: In the lesson’s example printout, call 1 waits 740 ms (thousandths of a second) for its first word. With the 3,000-token opening reused (a token is a word piece, about three quarters of a word), the next calls wait about 36 ms: about 20 times less.

You are done when: You have three things. On llama.cpp: the wait for the first word with the opening first and with the time first. On Fireworks: the opening-first wait, and a usage line where cached_tokens covers most of prompt_tokens.

Prefix Caching · 2:37

Download mp4 (9.6 MB)

Chapters

In this video Why agent traffic is mostly repeats, how servers reuse the repeated part, and what that does to the wait and the bill.

3 key points

  1. Most of what an agent sends is the same opening again.

    In the video’s 20-call example, 66,500 of the 72,000 tokens sent repeat the opening: 92%. The video shows 93% because it rounds 66,500 up to 67,000.

  2. Reusing the opening cuts both the wait and the bill.

    In the video’s session, 95 of 100 calls see their first word within 0.35 seconds instead of 1.9. The input costs about a fifth of a cent instead of 2 cents.

  3. Put the parts that never change first, or there is nothing to reuse.

    One changing line at the top, such as the time, makes every call new: calls 2 to 12 took 712 ms instead of 36 ms in the lesson’s sample.

An agent sends the same long opening with every request: its instructions, its tools (actions it may ask for, such as a refund) and its rules. Only the last line, the question, changes. With prefix caching the server keeps its notes on that opening (the KV cache from Day 4). It then reads only the new part, so the reply starts much sooner.

Picture it

A bookmark: if the first 3,000 words of a book are the same as last time, you start reading at word 3,001. It only works if every word before the bookmark matches. Change one word on page 1 and you start again from the top.

With real numbersthe lesson’s script and example printout (lesson 06), and the Prefix Caching video’s prices

  • The shared opening: 3,028 tokens by the script’s estimate (12,115 characters ÷ 4; about 2,300 words), sent with each of 12 different questions.
  • Opening first, call 1 (nothing saved yet): 740 ms before the first word.
  • Calls 2 to 12 (opening reused): 36 ms, the middle value. 740 ÷ 36 = about 20 times sooner.
  • A line with the time of the request (and the question) added above the opening: call 1 takes 744 ms, calls 2 to 12 about 712 ms. 744 ÷ 712 = 1.04, so almost no gain.
  • On Fireworks the reused part also costs less. The video’s example prices: $0.30 per million tokens, or $0.006 cached. $0.006 ÷ $0.30 = 2% of the price.

Words to know

Prefix
The start of a prompt, up to the first token that differs from last time. Example: the shared 3,028-token opening.
TTFT (time to first token)
How long until the first token (word piece) of the reply appears. Example: 740 ms for call 1, 36 ms for the calls after it.
Agent
A program that calls a model many times, using tools, to finish a task. Example: the lesson’s support agent with 5 tools.
Cached tokens
Prompt tokens the server reused instead of reading again. Fireworks reports them as cached_tokens and bills them at a discount.

From the lesson

lessons/06-prefix-caching/README.md

An agent sends the same 3,000-token preamble (instructions, tool definitions, policy) on every call and changes only the last line. Without a prefix cache the server re-reads that preamble every time. With one, it keeps the KV cache from last time and reads only the new part.

It works like a bookmark: if the first 3,000 words are identical, start reading at word 3,001. It only works if the text is identical from the very first token. Put a timestamp at the top and every call looks new.

stable-first [■■■■■■■■■■ preamble (cached) ■■■■■■■■■■][question] → TTFT ~20× lower
variable-first [time+question][■■■■■■■ preamble ■■■■■■■] → nothing reusable
where name what it saves
vLLM automatic prefix caching (block hashes) TTFT, GPU time
SGLang RadixAttention (a tree of shared prefixes) TTFT, very good for branching agents
llama.cpp slot cache reuse TTFT
Fireworks prompt caching, on by default, with session affinity via user TTFT and the bill: cached input is heavily discounted
Terminal window
cd lessons/06-prefix-caching
python prefix_bench.py --target mock
python prefix_bench.py --target mock --order variable-first
bash ../02-three-local-servers/serve_llamacpp.sh # other terminal. Add Q4_K_M 16384 4 after .sh (see the note below)
python prefix_bench.py --target llamacpp
python prefix_bench.py --target llamacpp --order variable-first
bash mlx_prompt_cache.sh # the MLX by-hand version
python prefix_bench.py --target fireworks --usage # Fireworks: cached_tokens in the usage block

What each command does

  1. python prefix_bench.py --target mock

    Sends 12 different questions, each after the same 3,028-token opening, to the practice server (free, offline). Look for call 1 marked (cold) and calls 2 to 12 far faster. The practice server reads 2,500 tokens a second, so call 1 takes about 1.2 seconds (3,028 ÷ 2,500), more than the lesson’s 740 ms. Calls 2 to 12 take about 30 ms (15 ms plus the few new tokens), so its speed-up comes out near 40.

  2. python prefix_bench.py --target mock --order variable-first

    The same 12 questions, but with the time of the request (and the question) at the very top. Look for every call about as slow as call 1. speedup_x (call 1’s time divided by the middle of the others) should be about 1.0.

  3. bash ../02-three-local-servers/serve_llamacpp.sh

    Starts the real llama.cpp server from Day 2. Run it in a second terminal, from this lesson’s folder (cd lessons/06-prefix-caching inside 03-labs). Add Q4_K_M 16384 4 after .sh: the 4-bit model, room for 16,384 tokens, 4 seats (4,096 each). Without them it has 8,192 tokens over 4 seats, 2,048 each, too few for the 3,000-token opening (Day 2’s trap). Look for the line saying it is listening.

  4. python prefix_bench.py --target llamacpp

    The same test on the real model on your Mac. Look for a slow call 1 and calls 2 to 12 far faster. On Day 4’s example Mac, which reads about 1,050 tokens a second, call 1 takes roughly 3 seconds (3,028 ÷ 1,050). Write down first_ms (call 1) and rest_p50_ms (the middle of calls 2 to 12).

  5. python prefix_bench.py --target llamacpp --order variable-first

    The time at the top again, on the real model. Look for every call about as slow as call 1 and a speedup_x near 1.0: llama.cpp only reuses text that matches from the very first token.

  6. bash mlx_prompt_cache.sh

    The same idea done by hand with MLX (Apple’s own engine). Step 1 answers cold, reading everything. Step 2 reads the opening once and saves its notes to a file. Step 3 answers from that file. Look for the Prompt: line: about 3,000 tokens in step 1, only the short question in step 3.

  7. python prefix_bench.py --target fireworks --usage

    The same test on Fireworks (paid, about $0.03), then one extra call that prints the token counts. Every call carries the same session name (the user value), so all reach the same copy of the model. Look for the usage: line near the end, just above the Cost check line: cached_tokens should cover most of prompt_tokens.

How to read it

Each run prints one row, after a target column naming the server: stable-first puts the opening at the top, variable-first the time. first_ms is call 1, rest_p50_ms the middle of calls 2 to 12, and speedup_x the first divided by the second. In the lesson’s sample, stable-first falls from 740 to 36 ms (about 20 times) and variable-first stays slow, 744 then 712 ms (about 1).

order first_ms rest_p50_ms speedup_x
stable-first 740.0 36.0 20.6
variable-first 744.0 712.0 1.0 ← one changed token at the top kills it

With --usage on Fireworks, cached_tokens should cover most of prompt_tokens.

3 questions. Say your answer out loud, then tap to check it.

A customer’s agent puts the current time (Current time: …) at the very top of its system prompt. The system prompt is the fixed instructions sent at the start of every request. What do you tell them?Show answerHide

In plain words

Move the time to the end. Put everything that never changes first and anything that changes last, so every request starts with the same text the server can reuse.

Picture it

A bookmark only helps if every page before it is the same as last time. Write the current time on page 1 and the bookmark is useless: you start the book again on every visit.

With real numbersthe lesson’s script and example printout, lesson 06

  • Opening first, nothing that changes above it (the lesson’s stable-first run, question at the end): call 1 takes 740 ms, the rest 36 ms (the middle value). 740 ÷ 36 = about 20 times sooner.
  • A line with the time and the question added above the opening (variable first): call 1 takes 744 ms, calls 2 to 12 about 712 ms. 744 ÷ 712 = 1.04, so almost no gain.
  • The only difference between the two runs: the time-first run adds that one line above the opening.
  • The fix changes how the prompt is built. It needs no new server and no new setting.

Words to know

System prompt
The fixed instructions at the start of every request. Example: the script’s 3,028-token support-agent opening.
Stable first, variable last
Build the prompt with the parts that never change at the top. Example: rules and tools first, the time and the question last.
Prefix
The start of a prompt, up to the first token that differs from last time. Example: a new timestamp at the top leaves no shared prefix.
Go deeper: the engineer version

The kit's question

The customer’s agent puts Current time: … at the top of the system prompt. What do you tell them?

The kit's answer

Move it to the end. Stable content first, variable content last.

More detail: Prefix caches match token by token from the start. llama.cpp and SGLang reuse the longest identical prefix; vLLM and the practice server hash fixed-size blocks, each key covering everything before it (64-token blocks in mock_server.py), so one changed token invalidates every block after it. Per-request values such as the time or the user’s name belong after the stable part, ideally in the last user message.

A customer runs several identical copies of the model (replicas) behind one address. Why does it matter to send a user value (a session name) with each request, so each session keeps going to the same copy (session affinity)?Show answerHide

In plain words

Each copy keeps its own saved notes, and the copies do not share them. A request that lands on a copy that has not seen this conversation starts cold: that copy must read the conversation again.

Picture it

A chain of cafes where each barista remembers your usual order. Go back to the same barista and your coffee comes fast. Walk up to a new one and you explain it all again.

With real numberslesson 06’s example printout and Day 3’s sizing question

  • Warm, in the lesson’s sample (the copy has already read the text): 36 ms to the first word.
  • Cold (the copy must read it all): 740 ms, about 20 times longer.
  • Day 3’s example service had 18 copies. Sent at random, a conversation’s next request would reach the copy holding its history only 1 time in 18.
  • The script sends user = lesson-06-session on every Fireworks call, so they all reach the same copy.

Words to know

Replica
One complete running copy of the server and model. Example: 18 copies for 120 people on Day 3.
Session affinity
Sending every request of one session to the same copy, so it finds its saved notes. Example: the user field.
Cold and warm
Cold: nothing saved yet, so everything is read; warm: the saved notes are reused. Example: 740 ms cold, 36 ms warm.
Go deeper: the engineer version

The kit's question

Why does user / session affinity matter on a multi-replica deployment?

The kit's answer

Each replica holds its own cache, so a request routed to a different replica starts cold.

More detail: Each replica holds its own KV cache in its own GPU memory; nothing is shared between replicas. A shared system prompt warms on each replica after its first request there, but a session’s own history (earlier turns, tool results) is cached only where it was served. Without affinity those turns land on different replicas, each pays full prefill, and the hit rate falls. Caches also expire, so the first call after an idle spell is slow (the Prefix Caching video).

What is the business case for prefix caching: the argument, in money and speed, you would give the customer?Show answerHide

In plain words

Most of an agent’s input is the same opening again, so most of the input bill moves to the much cheaper cached price. The reply also starts several times sooner.

Picture it

A print shop charges full price to set up a page and a small fee for every extra copy. An agent keeps ordering the same page; with caching it pays for the setup once and copy prices after that.

With real numbersthe Prefix Caching video’s example prices, lesson 06’s answer (90% repeated) and the Day 10 memo script

  • An agent whose input is 90% repeated opening, per 1 million input tokens:
  • Without caching: 1,000,000 × $0.30 per million = $0.30.
  • With caching: 100,000 new tokens × $0.30 per million = $0.03, plus 900,000 cached × $0.006 per million = $0.0054.
  • Total $0.0354 instead of $0.30: about 12% of the bill, so 88% saved.
  • At 50% off instead (what the Day 10 memo assumes until you check): $0.03 + 900,000 × $0.15 per million = $0.165, so 45% saved.
  • Lesson example printout: 740 ms becomes 36 ms (the middle call), about 20 times sooner.
  • The video’s session: 1.9 s becomes 0.35 s (the time 95 of 100 calls stay under), about 5 times sooner.

Words to know

Business case
The argument, in money and results, for making a change. Example: 88% off the input bill at the video’s prices.
Input tokens
The tokens you send to the model; providers bill them per million. Example: $0.30 per million in the video’s example.
Cached tokens
Input tokens the server reused; Fireworks bills them at a discount. Example: $0.006 instead of $0.30 per million.
Order of magnitude
About 10 times. Example: 740 ms down to 36 ms is more than one order of magnitude.
Go deeper: the engineer version

The kit's question

What is the business case?

The kit's answer

For an agent that is 90% repeated context, most of the input bill becomes cached-token pricing, and TTFT drops by an order of magnitude.

More detail: Cached-token pricing is set per model. The video’s $0.30 and $0.006 per million match the lab book’s figures for DeepSeek V4.1 Flash on Fireworks (98% off); the Day 10 memo script assumes only 50% off until you check the price page. Your own run uses the script’s default model, gpt-oss-120b, which has its own prices; read them off the price page before quoting a saving. The TTFT gain depends on the workload: 20x (call 1 against the median of calls 2 to 12) in the lesson’s sample, 5.4x at p95 in the video’s measured session (1.9 s to 0.35 s). Quote the customer’s own before-and-after with the cache hit rate; a cache that is never hit only takes memory away from concurrency (the video).

Question Our agent bill is too high. What would you do first?

One clear answer

Their system prompt is identical on every call, so most of the input bill is cacheable. Stable first, variable last, session affinity on, and I’ll show the cached-token count in the usage block.

What this means

  • “Their system prompt is identical on every call”: The customer’s fixed opening (instructions, tools, rules) is the same text on every request. In the lesson it is 3,028 tokens, sent with every question.
  • “so most of the input bill is cacheable”: Most of what they pay for the model to read can be reused and billed at the cached price. In the video’s 20-call session, 92% of the input tokens are repeats.
  • “Stable first, variable last”: Put what never changes at the top and what changes (the time, the question) at the bottom, so the opening matches every time. With the time at the top, calls 2 to 12 took 712 ms instead of 36 ms in the lesson’s sample.
  • “session affinity on”: Send a session id (the user field) so each conversation keeps going to the same copy of the model, the one holding its saved notes.
  • “and I’ll show the cached-token count in the usage block”: I prove it with the counts Fireworks returns with each answer: cached_tokens next to prompt_tokens. Most of the prompt should be cached.

Your numbersSaved on this device and collected in the Day 5 wrap-up.

Hint: From python prefix_bench.py --target llamacpp: the first_ms and rest_p50_ms columns. Write both, for example “740 / 36”.

Hint: The speedup_x column of the --order variable-first run. Expect about 1.0.

Hint: From python prefix_bench.py --target fireworks --usage: the first_ms and rest_p50_ms columns. Write both.

Hint: From the usage: line near the end, just above the Cost check line. Write both numbers, cached first.

Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.

You have three things. On llama.cpp: the wait for the first word with the opening first and with the time first. On Fireworks: the opening-first wait, and a usage line where cached_tokens covers most of prompt_tokens.

Connection refused.
That server is not running. make check (from the 03-labs folder) shows what is up. Start the practice server with make mock &, or start llama.cpp as in the run notes and wait until it says it is listening.
Opening first on the practice server shows a speed-up near 1, and call 1 is already fast.
The practice server still remembers the opening from an earlier run: this command run twice, or make smoke while it was up. Restart it: pkill -f felab.mock_server, then make mock & from the 03-labs folder, and run again.
No such file or directory when you start llama.cpp.
The command’s path starts from this lesson’s folder. From the 03-labs folder run cd lessons/06-prefix-caching, then start it again.
llama.cpp returns an error about the context size, or its opening-first run shows no speed-up.
Each of its 4 seats holds only 2,048 tokens by default, less than the 3,000-token opening. Stop it (Ctrl+C) and, from this lesson’s folder, start it with bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 4 (4,096 tokens per seat).
FIREWORKS_API_KEY is not set.
Put your key in the .env file in 03-labs, as on Day 1, or export it in this terminal.
cached_tokens=0 or cached_tokens=None.
Run the Fireworks command again right away: caches expire when idle (the Prefix Caching video). If it still says None, this model does not report it; pass another with --model.
A Fireworks model id is not found.
Model ids change. Copy a current one from the Fireworks model library and pass it with --model.
mlx_lm.generate: command not found.
The Python environment from Day 1 is not active. From the 03-labs folder run source .venv/bin/activate, then try again.
mlx_prompt_cache.sh Bash · 25 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 06 · MLX variant — the same idea, done by hand.
#
# mlx_lm.cache_prompt runs prefill ONCE on the preamble and saves the KV cache
# to disk. mlx_lm.generate then loads that file and only prefills the new question.
# This is literally what a server-side prefix cache does, made visible.
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
cd "$(dirname "$0")"
MODEL=${MLX_MODEL:-mlx-community/Meta-Llama-3.1-8B-Instruct-4bit}
# Dump the same preamble prefix_bench.py uses
python -c "from prefix_bench import PREAMBLE; print(PREAMBLE)" > preamble.txt
echo "▸ 1. cold: prefill the ~3k-token preamble + question (watch 'Prompt: … tokens-per-sec')"
time mlx_lm.generate --model "$MODEL" --max-tokens 40 \
--prompt "$(cat preamble.txt) Question: How do I rotate my API key?"
echo "▸ 2. cache the preamble once"
mlx_lm.cache_prompt --model "$MODEL" --prompt "$(cat preamble.txt)" --prompt-cache-file preamble.safetensors
echo "▸ 3. warm: only the question is prefilled"
time mlx_lm.generate --prompt-cache-file preamble.safetensors --max-tokens 40 \
--prompt " Question: How do I rotate my API key?"
prefix_bench.py Python · 88 lines
"""
Lesson 06 — prefix caching on an agent-shaped workload.
Agents resend the same big preamble (instructions + tool definitions + docs) on
every call and change only the last few lines. If the server remembers the KV
cache for that preamble, it skips re-reading it: TTFT collapses and, on Fireworks,
cached input tokens are billed at a steep discount.
This script sends 12 different questions two ways:
--order stable-first [ 3k-token preamble ][ question ] ← cache-friendly
--order variable-first [ question + timestamp ][ preamble ] ← one changed token early
invalidates everything after it
python prefix_bench.py --target mock
python prefix_bench.py --target mock --order variable-first
python prefix_bench.py --target llamacpp # llama-server reuses slot caches
python prefix_bench.py --target fireworks --usage # ~$0.03; prints cached_tokens
"""
import argparse
import datetime as dt
from felab import add_target_args, banner, client, percentile, record, resolve, stream_once, table
TOOLS = "\n".join(
f"- tool `{name}`: {desc}. Arguments: JSON object with fields id (string), limit (int), filters (object)."
for name, desc in [("search_tickets", "full-text search over support tickets"),
("get_invoice", "fetch an invoice by id"), ("refund", "issue a refund up to the limit"),
("status_page", "read current incident status"), ("escalate", "page the on-call engineer")])
POLICY = ("Refunds over $500 need manager approval. Outages affecting more than 5% of users are SEV-2. "
"Never reveal internal ticket ids. Answer in under 80 words. ") * 75
PREAMBLE = f"You are the support agent for Acme Cloud.\n\n## Tools\n{TOOLS}\n\n## Policy\n{POLICY}"
QUESTIONS = [
"A customer was charged twice this month. What do I do?", "Is there an outage in eu-west right now?",
"How do I rotate my API key?", "Customer wants a $900 refund — can I approve it?",
"What counts as a SEV-2?", "Can I share the internal ticket id with the customer?",
"Which tool finds old tickets about SSO?", "Latency spiked at 09:00 — who do I page?",
"The customer asks for their March invoice.", "How long should my answer be?",
"Card declined on renewal, launch is tomorrow.", "Summarise the refund policy in one line.",
]
def messages(q: str, order: str) -> list[dict]:
if order == "stable-first":
return [{"role": "system", "content": PREAMBLE}, {"role": "user", "content": q}]
# anti-pattern: something that changes every call goes FIRST
stamp = dt.datetime.now().isoformat(timespec="microseconds")
return [{"role": "system", "content": f"Request time {stamp}. User asks: {q}\n\n{PREAMBLE}"},
{"role": "user", "content": q}]
def main() -> None:
ap = add_target_args(argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter))
ap.add_argument("--order", choices=["stable-first", "variable-first"], default="stable-first")
ap.add_argument("--usage", action="store_true", help="one extra non-streamed call to print cached_tokens")
a = ap.parse_args()
t = resolve(a)
banner(t)
cli = client(t)
# `user` = session affinity on Fireworks: keeps you on the replica that holds your cache.
extra = {"user": "lesson-06-session"} if t.is_paid else {}
print(f"preamble ≈ {len(PREAMBLE) // 4:,} tokens · order = {a.order}\n")
ttft = []
for i, q in enumerate(QUESTIONS):
s = stream_once(cli, t.model, messages(q, a.order), max_tokens=48, **extra)
ttft.append(s.ttft_ms)
print(f" call {i + 1:2d} TTFT {s.ttft_ms:7.0f} ms {'(cold)' if i == 0 else ''}")
row = {"target": t.name, "order": a.order, "first_ms": ttft[0],
"rest_p50_ms": percentile(ttft[1:], 50), "speedup_x": ttft[0] / max(1e-6, percentile(ttft[1:], 50))}
print("\n" + table([row]))
record("06-prefix", row)
if a.usage:
r = cli.chat.completions.create(model=t.model, max_tokens=8,
messages=messages("One more question.", a.order), **extra)
u = r.usage
cached = getattr(getattr(u, "prompt_tokens_details", None), "cached_tokens", None)
print(f"\nusage: prompt_tokens={u.prompt_tokens} cached_tokens={cached}")
print("Cost check: cached input is billed at a fraction of normal input — see the model's pricing page.")
if __name__ == "__main__":
main()