Day 6 · Structured output and batch
Part 1 · Day 6 of 12
0 of 12 days done
About 1 h 05 About $0.03 Mac and Fireworks
Today: Make a model’s replies come back in a fixed shape that code can read, every time. Then measure how often they are also right. Finally, send the same test to Fireworks as one file through its Batch API (its bulk service), which answers later at about half the usual price.
Example
Imagine software that sends support tickets to the right team. It needs every model reply in a shape it can read; otherwise the ticket gets stuck. In the practice test, making the server enforce that shape fixes all the unreadable replies, but some tickets still go to the wrong team. Today you will measure those two errors separately: a readable answer is not necessarily a correct one.
By the end you’ll have
The share of well-formed replies you would quote to a customer, measured on 50 test tickets the model never trained on, and a 100-ticket test scored through the Batch API.
Expected: About 100% well-formed with the format enforced, but fewer right answers (78% on the practice server). From the Batch API: an answer for each of the 100 test tickets, matched to its ticket.
The 6 numbers you write down
Well-formed replies on your Mac: prompt only / schema enforced (%)
How often the real model on your Mac follows the schema when only asked, and when the server enforces it. Only the second is guaranteed.
Example: 86 / 100 on the practice server
Right category on your Mac, schema enforced (%)
How often a well-formed reply is also right. Format and correctness are two separate numbers.
Example: 78 on the practice server
Fireworks, schema enforced: well-formed / right category (%)
The figure you would quote to a customer: the share of replies their code can rely on, measured on 50 test tickets the model never trained on, with how often they are right beside it. Day 10’s memo reads your latest run from
results/07-structured.csv.Example: 100 / 78 on the practice server
Practice batch: answered (of 100), and right category (%)
Proves the loop works end to end before you pay: file in, answers matched back by label, scored.
Example: 100 of 100, 78%
Your Fireworks batch job name
Lets you check or finish the job later with
firectl batch-inference-job get <job name>.Example:
triage-batch-plus month, day, hour and minuteFireworks batch: answered (of 100) / well-formed (%) / right category (%)
A real model’s score on the 100 test tickets, which it never trained on, bought at about half the usual price.
Example: 100 / 100 / 78 (the practice run’s figures; yours come from a real model)
How you will use this: On Day 7, a test rig (the eval harness) grades the same tickets more ways: a second model acting as judge, speed, and cost per 1,000 tickets. On Day 10, your sizing memo (a written hardware and cost recommendation) quotes your latest two well-formed shares, asked in the prompt and enforced, from results/07-structured.csv. It names batch as the way to halve the cost of tests and bulk re-processing.
Before you start
The practice server is running.
Both steps rehearse on the kit’s practice server first: free, offline, with made-up numbers. Start it from the
03-labsfolder with the Python environment on (source .venv/bin/activate).make mock &It prints
mock LLM on http://localhost:9000/v1.llama.cpp from Day 2, for step 1.
Step 1 also asks a real model, run on your Mac by llama.cpp (the program that runs models locally). Nothing new downloads: it reuses Day 2’s Llama 3.1 8B, stored at 4 bits per number. In a second terminal, from the
03-labsfolder:bash lessons/02-three-local-servers/serve_llamacpp.shWait for the line saying it is listening (ready for requests). Leave it running for step 1, then stop it with Ctrl+C.
Your Fireworks key and firectl from Day 1.
Both steps end on Fireworks, about $0.03 in all. The scripts read your key from the
.envfile in03-labs. Step 2 also uses firectl, Fireworks’ command-line tool. Check it is still signed in:firectl whoamiIt shows your account. If it asks you to sign in, run
firectl signin.
Warm-up from Day 5
Section titled “Warm-up from Day 5”About 2 minutes. Say your answer out loud, then tap to check it. From Day 5 · Agent workloads.
A customer’s agent puts the current time (Current time: …) at the very top of its system prompt. The system prompt is the fixed instructions sent at the start of every request. What do you tell them?Show answerHide
In plain words
Move the time to the end. Put everything that never changes first and anything that changes last, so every request starts with the same text the server can reuse.
Picture it
A bookmark only helps if every page before it is the same as last time. Write the current time on page 1 and the bookmark is useless: you start the book again on every visit.
With real numbersthe lesson’s script and example printout, lesson 06
- Opening first, nothing that changes above it (the lesson’s stable-first run, question at the end): call 1 takes 740 ms, the rest 36 ms (the middle value). 740 ÷ 36 = about 20 times sooner.
- A line with the time and the question added above the opening (variable first): call 1 takes 744 ms, calls 2 to 12 about 712 ms. 744 ÷ 712 = 1.04, so almost no gain.
- The only difference between the two runs: the time-first run adds that one line above the opening.
- The fix changes how the prompt is built. It needs no new server and no new setting.
Words to know
- System prompt
- The fixed instructions at the start of every request. Example: the script’s 3,028-token support-agent opening.
- Stable first, variable last
- Build the prompt with the parts that never change at the top. Example: rules and tools first, the time and the question last.
- Prefix
- The start of a prompt, up to the first token that differs from last time. Example: a new timestamp at the top leaves no shared prefix.
Go deeper: the engineer version
The kit's question
The customer’s agent puts Current time: … at the top of the system prompt. What do you tell them?
The kit's answer
Move it to the end. Stable content first, variable content last.
More detail: Prefix caches match token by token from the start. llama.cpp and SGLang reuse the longest identical prefix; vLLM and the practice server hash fixed-size blocks, each key covering everything before it (64-token blocks in mock_server.py), so one changed token invalidates every block after it. Per-request values such as the time or the user’s name belong after the stable part, ideally in the last user message.
A customer runs several identical copies of the model (replicas) behind one address. Why does it matter to send a user value (a session name) with each request, so each session keeps going to the same copy (session affinity)?Show answerHide
In plain words
Each copy keeps its own saved notes, and the copies do not share them. A request that lands on a copy that has not seen this conversation starts cold: that copy must read the conversation again.
Picture it
A chain of cafes where each barista remembers your usual order. Go back to the same barista and your coffee comes fast. Walk up to a new one and you explain it all again.
With real numberslesson 06’s example printout and Day 3’s sizing question
- Warm, in the lesson’s sample (the copy has already read the text): 36 ms to the first word.
- Cold (the copy must read it all): 740 ms, about 20 times longer.
- Day 3’s example service had 18 copies. Sent at random, a conversation’s next request would reach the copy holding its history only 1 time in 18.
- The script sends
user=lesson-06-sessionon every Fireworks call, so they all reach the same copy.
Words to know
- Replica
- One complete running copy of the server and model. Example: 18 copies for 120 people on Day 3.
- Session affinity
- Sending every request of one session to the same copy, so it finds its saved notes. Example: the
userfield. - Cold and warm
- Cold: nothing saved yet, so everything is read; warm: the saved notes are reused. Example: 740 ms cold, 36 ms warm.
Go deeper: the engineer version
The kit's question
Why does user / session affinity matter on a multi-replica deployment?
The kit's answer
Each replica holds its own cache, so a request routed to a different replica starts cold.
More detail: Each replica holds its own KV cache in its own GPU memory; nothing is shared between replicas. A shared system prompt warms on each replica after its first request there, but a session’s own history (earlier turns, tool results) is cached only where it was served. Without affinity those turns land on different replicas, each pays full prefill, and the hit rate falls. Caches also expire, so the first call after an idle spell is slow (the Prefix Caching video).
What is the business case for prefix caching: the argument, in money and speed, you would give the customer?Show answerHide
In plain words
Most of an agent’s input is the same opening again, so most of the input bill moves to the much cheaper cached price. The reply also starts several times sooner.
Picture it
A print shop charges full price to set up a page and a small fee for every extra copy. An agent keeps ordering the same page; with caching it pays for the setup once and copy prices after that.
With real numbersthe Prefix Caching video’s example prices, lesson 06’s answer (90% repeated) and the Day 10 memo script
- An agent whose input is 90% repeated opening, per 1 million input tokens:
- Without caching: 1,000,000 × $0.30 per million = $0.30.
- With caching: 100,000 new tokens × $0.30 per million = $0.03, plus 900,000 cached × $0.006 per million = $0.0054.
- Total $0.0354 instead of $0.30: about 12% of the bill, so 88% saved.
- At 50% off instead (what the Day 10 memo assumes until you check): $0.03 + 900,000 × $0.15 per million = $0.165, so 45% saved.
- Lesson example printout: 740 ms becomes 36 ms (the middle call), about 20 times sooner.
- The video’s session: 1.9 s becomes 0.35 s (the time 95 of 100 calls stay under), about 5 times sooner.
Words to know
- Business case
- The argument, in money and results, for making a change. Example: 88% off the input bill at the video’s prices.
- Input tokens
- The tokens you send to the model; providers bill them per million. Example: $0.30 per million in the video’s example.
- Cached tokens
- Input tokens the server reused; Fireworks bills them at a discount. Example: $0.006 instead of $0.30 per million.
- Order of magnitude
- About 10 times. Example: 740 ms down to 36 ms is more than one order of magnitude.
Go deeper: the engineer version
The kit's question
What is the business case?
The kit's answer
For an agent that is 90% repeated context, most of the input bill becomes cached-token pricing, and TTFT drops by an order of magnitude.
More detail: Cached-token pricing is set per model. The video’s $0.30 and $0.006 per million match the lab book’s figures for DeepSeek V4.1 Flash on Fireworks (98% off); the Day 10 memo script assumes only 50% off until you check the price page. Your own run uses the script’s default model, gpt-oss-120b, which has its own prices; read them off the price page before quoting a saving. The TTFT gain depends on the workload: 20x (call 1 against the median of calls 2 to 12) in the lesson’s sample, 5.4x at p95 in the video’s measured session (1.9 s to 0.35 s). Quote the customer’s own before-and-after with the cache hit rate; a cache that is never hit only takes memory away from concurrency (the video).
Today’s steps
Section titled “Today’s steps”2 steps, then the wrap-up.
- Structured output Business software reads a model’s replies with code, and one badly formed reply can crash it. You try three ways of asking for the format, measure how often replies come back well-formed, and see that well-formed is not the same as right.
-
Batch API
Jobs where nobody waits for the answer can go to Fireworks’ Batch API: one file in, answers later, at about half the usual price. You rehearse the round trip offline, then run it for real on 100 test tickets.
Video: The Platform Map 2:51
- Wrap-up and drill Put your numbers in one table, answer 5 questions out loud, explain 2 results and work 2 examples.
Your progress is saved on this device.