Day 5 · Agent workloads
Part 1 · Day 5 of 12
0 of 12 days done
About 45 min About $0.03 Mac and Fireworks
Today: An agent (a program that calls a model again and again) sends the same long opening, its instructions and rules, with every request. Today you measure how much sooner the reply starts when the server reuses that opening. You do it on your Mac, then on Fireworks, which also bills the reused part at a steep discount.
Example
Imagine a support bot that receives the same company handbook before every customer question. Reading that handbook again each time wastes work. In the lesson’s test, the first answer starts after about three quarters of a second; later answers start almost immediately because the server reuses its work on the handbook. Today you will measure that difference and see when the reuse fails.
By the end you’ll have
Your wait for the first word of a reply, with the opening reused and without, on your Mac and on Fireworks, plus how many tokens Fireworks reused.
Expected: With the opening at the top of the prompt, calls 2 to 12 get their first word about 20 times sooner than call 1 in the lesson’s sample; with the time of the request at the top instead, no sooner.
The 4 numbers you write down
llama.cpp, opening first: call 1 / middle of calls 2 to 12 (ms)
How long the first word takes when the whole opening must be read, and when it is reused.
Example: 740 / 36 in the lesson’s example printout: about 20 times sooner.
llama.cpp, time first: speed-up
How much reuse is left when one changing line sits at the top: none.
Example: 1.0 in the lesson’s sample (744 ÷ 712 = 1.04).
Fireworks, opening first: call 1 / middle of calls 2 to 12 (ms)
The same test on Fireworks, where caching is on by default.
Example: No lesson sample for Fireworks. For scale, the video’s session fell from 1,900 to 350 ms, but that is the time 95 of 100 calls stay under, not call 1 and the middle.
Fireworks: cached tokens / prompt tokens
How much of the prompt Fireworks reused and billed at the cached price.
Example: cached_tokens should cover most of prompt_tokens; the opening alone is about 3,000 of them by the script’s estimate.
How you will use this: Day 10’s sizing memo quotes your speed-up. It uses the last opening-first result you saved, which is your Fireworks run if you follow the steps in order.
Before you start
The practice server is running.
The first two runs use the kit’s practice server: free, offline, with made-up speeds. Start it from the
03-labsfolder. If it is still running from an earlier day, restart it: it may already remember today’s opening, and then call 1 is fast too.make mock &It prints
mock LLM on http://localhost:9000/v1. To restart it, first runpkill -f felab.mock_server.llama.cpp and MLX (Apple’s own model software) from Day 2.
Both run the 4-bit Llama 3.1 8B model you downloaded on Day 2 (each keeps its own copy), so there is nothing new to download today. You start llama.cpp in a second terminal during the step, with room for a longer prompt than its default.
Your Fireworks key and spend cap from Day 1.
The last run calls Fireworks and costs about $0.03. The script reads your key from the
.envfile in03-labs, set up on Day 1.
Warm-up from Day 4
Section titled “Warm-up from Day 4”About 2 minutes. Say your answer out loud, then tap to check it. From Day 4 · Memory levers.
Shrinking the stored model from 8 bits to about 4 nearly doubles how fast it writes (decode), but barely changes how fast it reads your prompt (prefill). Why?Show answerHide
In plain words
Writing waits on memory: each new token needs the whole model fetched once, so a smaller model arrives faster. Reading your prompt waits on math instead, and fewer bits do not make a Mac’s math faster.
Picture it
Before every dish, the whole recipe book comes to the cook on a conveyor belt that moves a fixed number of pages per second. Halve the pages and the cook turns out nearly twice as many dishes a minute. Reading the prompt is one delivery followed by lots of cooking, so a thinner book barely helps.
With real numberslesson 04 sample figures, M4 Max
- Model size: 8.5 GB at 8-bit, 4.9 GB at 4-bit (58% of the bytes).
- Writing: 52 tokens/s becomes 86, 1.65 times faster.
- Reading the prompt: 1,150 tokens/s becomes 1,050, about the same (9% slower).
- Speed limits (M4 Max memory speed, 546 GB/s, ÷ model size): 546 ÷ 8.5 = 64 and 546 ÷ 4.9 = 111 tokens/s.
Words to know
- Decode
- The writing phase, one token at a time. Example: the
tg128column (write 128 tokens) measures it. - Prefill
- The reading phase, the whole prompt at once. Example:
pp512measures it. - Bandwidth-bound
- Slowed by how fast memory delivers data (bandwidth, in GB per second). Example: decode; an M4 Max delivers 546 GB/s.
- Compute-bound
- Slowed by calculation speed rather than memory. Example: prefill.
Go deeper: the engineer version
The kit's question
Why does halving the bytes nearly double decode but not prefill?
The kit's answer
Decode is bandwidth-bound; prefill is compute-bound.
More detail: A batch-1 decode step does about two floating-point operations per parameter while streaming every weight (the GPU Bandwidth video), so bytes set the pace. Prefill (pp512) reuses each weight read across 512 tokens, so matrix-math throughput sets the pace. llama.cpp’s smaller formats must be unpacked before the math, which is why prefill even dips slightly.
On an NVIDIA H100 (a GPU: the kind of chip that runs AI models in data centers), models are often stored as FP8, an 8-bit number format. Where does that fit in with shrinking a model to fewer bits (quantization) on a Mac?Show answerHide
In plain words
It is the same shrinking idea, but the H100 can also do its math directly on 8-bit numbers. So both fetching the model and doing the math get faster, not only the fetching.
Picture it
On a Mac, a thinner recipe book only makes each delivery faster. On an H100 the cook has also learned the book’s shorthand, so the cooking speeds up too.
With real numbersthe GPU Bandwidth video (Day 3) and the Quantization video
- A 70B model: 70 billion x 2 bytes = 140 GB at 16-bit; 70 GB at 8-bit; 35 GB at 4-bit.
- H100 memory speed: 3,350 GB/s, about 6 times an M4 Max’s 546.
- Writing-speed limit: 3,350 ÷ 70 = 48 tokens/s at 8-bit; 3,350 ÷ 35 = about 96 at 4-bit.
- The extra on H100: built-in 8-bit math, roughly double the math speed, so prompt reading speeds up too.
- The next generation, Blackwell, adds built-in 4-bit math (NVFP4).
Words to know
- FP8
- An 8-bit number format, 1 byte per number. Example: a 70B model is 70 GB in FP8.
- Tensor cores
- The parts of an NVIDIA GPU built for fast matrix math. Example: the H100’s tensor cores do math directly on FP8 numbers.
- NVFP4
- NVIDIA’s 4-bit number format with built-in math on Blackwell GPUs.
- Blackwell
- The NVIDIA GPU generation after H100. Example: the B200 chip, whose memory moves 8,000 GB/s.
Go deeper: the engineer version
The kit's question
Where does FP8 on an H100 fit in?
The kit's answer
Same idea in hardware. H100/H200 have FP8 tensor cores, and Blackwell adds NVFP4, so the compute also speeds up rather than just memory.
More detail: On Apple silicon the GGUF formats only cut bytes; the math still runs at higher precision, which is why prefill stayed flat in the previous question. On H100, FP8 weights halve decode bytes and FP8 tensor cores roughly double matrix-math throughput, which helps prefill and large batches. The Quantization video: “FP8 on H100+ is close to free; four-bit needs measurement.”
The customer worries that a smaller, faster version of the model will give worse answers. What do you do?Show answerHide
In plain words
Test it instead of arguing. Run 50 of their own real prompts at each size and show quality, speed and cost side by side.
Picture it
A bakery wants to switch to a cheaper flour. You bake their best-sellers both ways and let their own customers taste, instead of quoting a lab report.
With real numberslesson 04’s quality check and the Quantization video’s sample table
- Lesson 04’s quick alarm: 12 questions per size. If the 3-bit version gets 7 right where the 8-bit gets 11, quality has fallen off a cliff.
- For a customer: 50 of their own prompts at each size (lesson 09, on Day 7, builds the test).
- The Quantization video’s example table (made-up numbers to show the shape): at 8-bit instead of 16-bit, accuracy slipped from 89% to 88%, total output rose from 24 to 46 tokens per second, conversations per GPU from 3 to 7.
- The video’s rule: if 2 fewer replies in 100 come back in the exact format asked for, but the server writes twice as many tokens per second, that is usually a good trade; 20 fewer in 100 never is.
Words to know
- Eval
- A repeatable test of a model on a fixed set of prompts. Example: 50 of the customer’s prompts.
- Precision
- How many bits each stored number gets. Example: 16-bit, 8-bit, 4-bit.
- Quality cliff
- The size below which answers suddenly get worse. Example: the 3-bit version (Q3) slips on multi-step and format questions first.
- Throughput
- Total output across everyone at once. Example: 24 against 46 tokens/s in the Quantization video’s sample.
Go deeper: the engineer version
The kit's question
The customer is worried about quality. What do you do?
The kit's answer
Run their own 50-prompt eval at each precision and show the quality, latency and cost table from lesson 09.
More detail: Lesson 09’s table has valid, accuracy, judge, p95 and $/1k tasks columns, and says never to show one column alone. Pick the precision per workload, on held-out prompts. The Quantization video states its rule in schema validity and throughput.
Today’s steps
Section titled “Today’s steps”1 step, then the wrap-up.
Your progress is saved on this device.
