Skip to content

Day 5 wrap-up and drill

  1. Overview
  2. Step 1
  3. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

Your notes show how long the first word takes on your Mac and on Fireworks, with the opening reused and without. They also hold Fireworks’ cached-token count.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 1 · Prefix caching

How long the first word takes when the whole opening must be read, and when it is reused.

Example: 740 / 36 in the lesson’s example printout: about 20 times sooner.

How much reuse is left when one changing line sits at the top: none.

Example: 1.0 in the lesson’s sample (744 ÷ 712 = 1.04).

The same test on Fireworks, where caching is on by default.

Example: No lesson sample for Fireworks. For scale, the video’s session fell from 1,900 to 350 ms, but that is the time 95 of 100 calls stay under, not call 1 and the middle.

How much of the prompt Fireworks reused and billed at the cached price.

Example: cached_tokens should cover most of prompt_tokens; the opening alone is about 3,000 of them by the script’s estimate.

Every question from today’s steps. Say your answer out loud, then show the answer and rate yourself. 3 cards

QuestionStep 1 · Prefix caching

A customer’s agent puts the current time (Current time: …) at the very top of its system prompt. The system prompt is the fixed instructions sent at the start of every request. What do you tell them?

Show answerHide answer

In plain words

Move the time to the end. Put everything that never changes first and anything that changes last, so every request starts with the same text the server can reuse.

Picture it

A bookmark only helps if every page before it is the same as last time. Write the current time on page 1 and the bookmark is useless: you start the book again on every visit.

With real numbersthe lesson’s script and example printout, lesson 06

  • Opening first, nothing that changes above it (the lesson’s stable-first run, question at the end): call 1 takes 740 ms, the rest 36 ms (the middle value). 740 ÷ 36 = about 20 times sooner.
  • A line with the time and the question added above the opening (variable first): call 1 takes 744 ms, calls 2 to 12 about 712 ms. 744 ÷ 712 = 1.04, so almost no gain.
  • The only difference between the two runs: the time-first run adds that one line above the opening.
  • The fix changes how the prompt is built. It needs no new server and no new setting.

Words to know

System prompt
The fixed instructions at the start of every request. Example: the script’s 3,028-token support-agent opening.
Stable first, variable last
Build the prompt with the parts that never change at the top. Example: rules and tools first, the time and the question last.
Prefix
The start of a prompt, up to the first token that differs from last time. Example: a new timestamp at the top leaves no shared prefix.
Go deeper: the engineer version

The kit's question

The customer’s agent puts Current time: … at the top of the system prompt. What do you tell them?

The kit's answer

Move it to the end. Stable content first, variable content last.

More detail: Prefix caches match token by token from the start. llama.cpp and SGLang reuse the longest identical prefix; vLLM and the practice server hash fixed-size blocks, each key covering everything before it (64-token blocks in mock_server.py), so one changed token invalidates every block after it. Per-request values such as the time or the user’s name belong after the stable part, ideally in the last user message.

How did you do?

QuestionStep 1 · Prefix caching

A customer runs several identical copies of the model (replicas) behind one address. Why does it matter to send a user value (a session name) with each request, so each session keeps going to the same copy (session affinity)?

Show answerHide answer

In plain words

Each copy keeps its own saved notes, and the copies do not share them. A request that lands on a copy that has not seen this conversation starts cold: that copy must read the conversation again.

Picture it

A chain of cafes where each barista remembers your usual order. Go back to the same barista and your coffee comes fast. Walk up to a new one and you explain it all again.

With real numberslesson 06’s example printout and Day 3’s sizing question

  • Warm, in the lesson’s sample (the copy has already read the text): 36 ms to the first word.
  • Cold (the copy must read it all): 740 ms, about 20 times longer.
  • Day 3’s example service had 18 copies. Sent at random, a conversation’s next request would reach the copy holding its history only 1 time in 18.
  • The script sends user = lesson-06-session on every Fireworks call, so they all reach the same copy.

Words to know

Replica
One complete running copy of the server and model. Example: 18 copies for 120 people on Day 3.
Session affinity
Sending every request of one session to the same copy, so it finds its saved notes. Example: the user field.
Cold and warm
Cold: nothing saved yet, so everything is read; warm: the saved notes are reused. Example: 740 ms cold, 36 ms warm.
Go deeper: the engineer version

The kit's question

Why does user / session affinity matter on a multi-replica deployment?

The kit's answer

Each replica holds its own cache, so a request routed to a different replica starts cold.

More detail: Each replica holds its own KV cache in its own GPU memory; nothing is shared between replicas. A shared system prompt warms on each replica after its first request there, but a session’s own history (earlier turns, tool results) is cached only where it was served. Without affinity those turns land on different replicas, each pays full prefill, and the hit rate falls. Caches also expire, so the first call after an idle spell is slow (the Prefix Caching video).

How did you do?

QuestionStep 1 · Prefix caching

What is the business case for prefix caching: the argument, in money and speed, you would give the customer?

Show answerHide answer

In plain words

Most of an agent’s input is the same opening again, so most of the input bill moves to the much cheaper cached price. The reply also starts several times sooner.

Picture it

A print shop charges full price to set up a page and a small fee for every extra copy. An agent keeps ordering the same page; with caching it pays for the setup once and copy prices after that.

With real numbersthe Prefix Caching video’s example prices, lesson 06’s answer (90% repeated) and the Day 10 memo script

  • An agent whose input is 90% repeated opening, per 1 million input tokens:
  • Without caching: 1,000,000 × $0.30 per million = $0.30.
  • With caching: 100,000 new tokens × $0.30 per million = $0.03, plus 900,000 cached × $0.006 per million = $0.0054.
  • Total $0.0354 instead of $0.30: about 12% of the bill, so 88% saved.
  • At 50% off instead (what the Day 10 memo assumes until you check): $0.03 + 900,000 × $0.15 per million = $0.165, so 45% saved.
  • Lesson example printout: 740 ms becomes 36 ms (the middle call), about 20 times sooner.
  • The video’s session: 1.9 s becomes 0.35 s (the time 95 of 100 calls stay under), about 5 times sooner.

Words to know

Business case
The argument, in money and results, for making a change. Example: 88% off the input bill at the video’s prices.
Input tokens
The tokens you send to the model; providers bill them per million. Example: $0.30 per million in the video’s example.
Cached tokens
Input tokens the server reused; Fireworks bills them at a discount. Example: $0.006 instead of $0.30 per million.
Order of magnitude
About 10 times. Example: 740 ms down to 36 ms is more than one order of magnitude.
Go deeper: the engineer version

The kit's question

What is the business case?

The kit's answer

For an agent that is 90% repeated context, most of the input bill becomes cached-token pricing, and TTFT drops by an order of magnitude.

More detail: Cached-token pricing is set per model. The video’s $0.30 and $0.006 per million match the lab book’s figures for DeepSeek V4.1 Flash on Fireworks (98% off); the Day 10 memo script assumes only 50% off until you check the price page. Your own run uses the script’s default model, gpt-oss-120b, which has its own prices; read them off the price page before quoting a saving. The TTFT gain depends on the workload: 20x (call 1 against the median of calls 2 to 12) in the lesson’s sample, 5.4x at p95 in the video’s measured session (1.9 s to 0.35 s). Quote the customer’s own before-and-after with the cache hit rate; a cache that is never hit only takes memory away from concurrency (the video).

How did you do?

Under 20 seconds each. Record yourself once and listen back.

Step 1 · Prefix caching

Question Our agent bill is too high. What would you do first?

One clear answer

Their system prompt is identical on every call, so most of the input bill is cacheable. Stable first, variable last, session affinity on, and I’ll show the cached-token count in the usage block.

What this means
  • “Their system prompt is identical on every call”: The customer’s fixed opening (instructions, tools, rules) is the same text on every request. In the lesson it is 3,028 tokens, sent with every question.
  • “so most of the input bill is cacheable”: Most of what they pay for the model to read can be reused and billed at the cached price. In the video’s 20-call session, 92% of the input tokens are repeats.
  • “Stable first, variable last”: Put what never changes at the top and what changes (the time, the question) at the bottom, so the opening matches every time. With the time at the top, calls 2 to 12 took 712 ms instead of 36 ms in the lesson’s sample.
  • “session affinity on”: Send a session id (the user field) so each conversation keeps going to the same copy of the model, the one holding its saved notes.
  • “and I’ll show the cached-token count in the usage block”: I prove it with the counts Fireworks returns with each answer: cached_tokens next to prompt_tokens. Most of the prompt should be cached.

2 worked examples from today’s videos. Work each on paper, then reveal.

From the video: Prefix Caching 2:37

Whiteboard 1 of 2

An agent session has 20 calls. Each call sends the same 3,500-token opening plus a new question of about 100 tokens. How many tokens are sent, how many are genuinely new, and what share is a repeat?

Given, in plain words

A token is a chunk of text, about three quarters of a word, so the opening is about 2,600 words. The server has to read the opening once. After that, only the questions are new.

Reveal the answerHide the answer

Answer · in plain words

72,000 tokens are sent, but only 5,500 are new. About 92 of every 100 tokens are repeats the server has already read.

Picture it

A teacher reads the same long instructions aloud before each of 20 questions. The class sits through them 20 times, but only the first reading told them anything new.

Careful

The video says 67,000 repeated, 5,000 new and 93%: it rounds 66,500 up to 67,000 and counts the rest as new. The exact figures are 66,500, 5,500 and 92%; the conclusion is the same: over 90% is repeats.

Worked answer, step by step

  1. Tokens per call: 3,500 (opening) + 100 (question) = 3,600.
  2. Tokens sent in the session: 20 calls × 3,600 = 72,000.
  3. Repeated: the opening is new only on call 1, so 19 × 3,500 = 66,500 tokens are repeats.
  4. Genuinely new: 3,500 (the first opening) + 20 × 100 (the questions) = 5,500. Check: 66,500 + 5,500 = 72,000.
  5. Share that is a repeat: 66,500 ÷ 72,000 = 92%.
  6. With the opening reused, the server reads 5,500 tokens instead of 72,000: 72,000 ÷ 5,500 = about 13 times less reading.
Go deeper: the engineer version

The kit's question

Worked example · 20-turn agent session · 3,500-token prefix · ~100-token questions

The kit's answer

tokens sent 72,000 · genuinely new 5,000 · repeated prefix 67,000 · share that is repeat 93% · Caching turns 93% of the input bill into a rounding error

More detail: The video’s 67,000 is 19 × 3,500 = 66,500 rounded up, and its 5,000 is 72,000 - 67,000. A real multi-turn session also resends its growing history on every turn, which raises the repeated share further (the field guide lists multi-turn chat history among the prefixes worth caching).

Words to know

Turn
One call to the model within a session. Example: 20 turns in the video’s session.
Prefix
The start of a prompt, up to the first token that differs from last time. Example: the 3,500-token opening.
Token
A chunk of text, about three quarters of a word. Example: 3,500 tokens is about 2,600 words.

From the video: Prefix Caching 2:37

Whiteboard 2 of 2

The same 20-call session. Input normally costs $0.30 per million tokens, and cached input $0.006 per million. What does one session’s input cost without caching and with it? What happens to the wait for the first word?

Given, in plain words

Cached input is the repeated opening the server reuses: 66,500 of the 72,000 tokens. The other 5,500 are paid at full price. The video measured the wait that 95 of 100 calls stay under (p95): 1.9 seconds cold, 0.35 seconds with the opening cached.

Reveal the answerHide the answer

Answer · in plain words

The session’s input cost falls from about 2 cents to about a fifth of a cent, about 90% less. And 95 of 100 calls get their first word within 0.35 seconds instead of 1.9, about 5 times sooner.

Picture it

A coffee shop charges full price for the first cup and 2 cents on the dollar for each refill; only a splash of new milk in each cup is full price. Twenty cups cost about as much as two.

Careful

Prices differ by model. Check the price page before you quote a saving; the Day 10 memo script assumes only 50% off until you do.

Worked answer, step by step

  1. Without caching, all 72,000 tokens at full price: 72,000 × $0.30 ÷ 1,000,000 = $0.0216, about 2 cents.
  2. With caching, the 5,500 new tokens at full price: 5,500 × $0.30 ÷ 1,000,000 = $0.00165.
  3. Plus the 66,500 repeated tokens at the cached price: 66,500 × $0.006 ÷ 1,000,000 = $0.0004.
  4. Total with caching: $0.00165 + $0.0004 = $0.00205, about a fifth of a cent.
  5. Saving: $0.00205 ÷ $0.0216 = 9.5% of the old cost, so about 90% less.
  6. The cached price is $0.006 ÷ $0.30 = 2% of the full price: a 98% discount.
  7. Wait for the first word (p95): 1.9 s cold, 0.35 s cached. 1.9 ÷ 0.35 = about 5 times sooner.
Go deeper: the engineer version

The kit's question

The same session, cached · p95 TTFT · input billed at full rate · input cost at $0.30 / $0.006

The kit's answer

cold · cached prefix · p95 TTFT: 1.9 s · 0.35 s · input billed at full rate: 72,000 tok · 5,000 tok · input cost at $0.30 / $0.006: $0.0216 · $0.0019 · Requires prompts ordered stable-first — a code change, not a config flag.

More detail: The video’s $0.0019 uses its rounded 5,000 new and 67,000 repeated tokens: $0.0015 + $0.000402 = $0.0019. The exact 5,500 and 66,500 give $0.00205. Both are more than 90% below $0.0216 (91% and 90.5%), as the video says. Its TTFT figures are p95 for the video’s measured session, a 5.4x gain, against 20x in the lesson’s sample (call 1 against the median of calls 2 to 12). The saving needs prompts ordered stable first: a code change, not a setting.

Words to know

p95
The value 95 out of 100 requests stay under. Example: 95 of 100 calls see their first word within 0.35 s, cached.
Cached tokens
Input tokens the server reused; Fireworks bills them at a discount. Example: $0.006 instead of $0.30 per million.
Input tokens
The tokens you send to the model; providers bill them per million. Example: 72,000 in the video’s session.

How you will use this

Day 10’s sizing memo quotes your speed-up. It uses the last opening-first result you saved, which is your Fireworks run if you follow the steps in order.