About 10 minutes, free, and it works on your phone. Lock in today before you move on.
Done when
Your numbers show two things. First, how many 32,768-token conversations (about 25,000 words each) fit in 20 GB before and after storing the model’s notes in 8 bits: 4, then 8. (The notes are the KV cache: what the model has noted about each conversation so far, stored in the same chip memory as the model.) Second, how fast each of the three model sizes writes.
Fields you filled in on the step pages are already here. Nothing leaves this browser.
Step 1 · Quantization
How fast each size writes its reply. Your speed limit is your Mac’s memory speed ÷ the model’s size: an M4 Pro moves 273 GB/s, so its limit for the 4.9 GB model is 273 ÷ 4.9 = about 56 tokens per second. Find your chip under Apple menu, About This Mac; ceiling_check.py looks up its memory speed for you.
Example: 52 / 86 / 98 (the lesson’s sample figures for an M4 Max, 546 GB/s)
tok/s
Your writing speed ÷ the speed limit (memory speed ÷ model size). It shows the formula holds on your own machine.
How many of 12 quick questions each size got right: a smoke alarm for the quality cliff, not a full test. A customer gets 50 of their own prompts instead.
Example: The lesson’s alarm example: 11 at 8-bit against 7 at 3-bit would be a cliff.
of 12
Step 2 · KV cache and context length
How much memory the notes for one 32,768-token conversation take, stored at 16 bits and at 8 bits. It is 8 times the 4,096-token figure (8 times the tokens), and the 8-bit figure is about half the 16-bit one. MiB is the unit llama.cpp prints, about a million bytes; 1,000 MiB is about 1 GB.
Example: 4,096 / 2,176 MiB (4.29 / 2.28 GB)
MiB
How much memory the notes take for the longest conversation the model allows, at 16 bits per number. It is more than three times the 4.9 GB model.
Example: 16,384 MiB (17.18 GB), or “failed” on a 16 to 24 GB Mac
MiB
The same notes stored with 8 bits per number: the fix. About half the memory, so about twice as many conversations fit.
Example: 8,704 MiB (9.13 GB)
MiB
How many people one machine serves at once with 20 GB of memory left over after the model loads, before and after the fix.
Shrinking the stored model from 8 bits to about 4 nearly doubles how fast it writes (decode), but barely changes how fast it reads your prompt (prefill). Why?
Show answerHide answer
In plain words
Writing waits on memory: each new token needs the whole model fetched once, so a smaller model arrives faster. Reading your prompt waits on math instead, and fewer bits do not make a Mac’s math faster.
Picture it
Before every dish, the whole recipe book comes to the cook on a conveyor belt that moves a fixed number of pages per second. Halve the pages and the cook turns out nearly twice as many dishes a minute. Reading the prompt is one delivery followed by lots of cooking, so a thinner book barely helps.
With real numberslesson 04 sample figures, M4 Max
Model size: 8.5 GB at 8-bit, 4.9 GB at 4-bit (58% of the bytes).
Writing: 52 tokens/s becomes 86, 1.65 times faster.
Reading the prompt: 1,150 tokens/s becomes 1,050, about the same (9% slower).
Speed limits (M4 Max memory speed, 546 GB/s, ÷ model size): 546 ÷ 8.5 = 64 and 546 ÷ 4.9 = 111 tokens/s.
Words to know
Decode
The writing phase, one token at a time. Example: the tg128 column (write 128 tokens) measures it.
Prefill
The reading phase, the whole prompt at once. Example: pp512 measures it.
Bandwidth-bound
Slowed by how fast memory delivers data (bandwidth, in GB per second). Example: decode; an M4 Max delivers 546 GB/s.
Compute-bound
Slowed by calculation speed rather than memory. Example: prefill.
Go deeper: the engineer version
The kit's question
Why does halving the bytes nearly double decode but not prefill?
The kit's answer
Decode is bandwidth-bound; prefill is compute-bound.
More detail: A batch-1 decode step does about two floating-point operations per parameter while streaming every weight (the GPU Bandwidth video), so bytes set the pace. Prefill (pp512) reuses each weight read across 512 tokens, so matrix-math throughput sets the pace. llama.cpp’s smaller formats must be unpacked before the math, which is why prefill even dips slightly.
On an NVIDIA H100 (a GPU: the kind of chip that runs AI models in data centers), models are often stored as FP8, an 8-bit number format. Where does that fit in with shrinking a model to fewer bits (quantization) on a Mac?
Show answerHide answer
In plain words
It is the same shrinking idea, but the H100 can also do its math directly on 8-bit numbers. So both fetching the model and doing the math get faster, not only the fetching.
Picture it
On a Mac, a thinner recipe book only makes each delivery faster. On an H100 the cook has also learned the book’s shorthand, so the cooking speeds up too.
With real numbersthe GPU Bandwidth video (Day 3) and the Quantization video
A 70B model: 70 billion x 2 bytes = 140 GB at 16-bit; 70 GB at 8-bit; 35 GB at 4-bit.
H100 memory speed: 3,350 GB/s, about 6 times an M4 Max’s 546.
Writing-speed limit: 3,350 ÷ 70 = 48 tokens/s at 8-bit; 3,350 ÷ 35 = about 96 at 4-bit.
The extra on H100: built-in 8-bit math, roughly double the math speed, so prompt reading speeds up too.
The next generation, Blackwell, adds built-in 4-bit math (NVFP4).
Words to know
FP8
An 8-bit number format, 1 byte per number. Example: a 70B model is 70 GB in FP8.
Tensor cores
The parts of an NVIDIA GPU built for fast matrix math. Example: the H100’s tensor cores do math directly on FP8 numbers.
NVFP4
NVIDIA’s 4-bit number format with built-in math on Blackwell GPUs.
Blackwell
The NVIDIA GPU generation after H100. Example: the B200 chip, whose memory moves 8,000 GB/s.
Go deeper: the engineer version
The kit's question
Where does FP8 on an H100 fit in?
The kit's answer
Same idea in hardware. H100/H200 have FP8 tensor cores, and Blackwell adds NVFP4, so the compute also speeds up rather than just memory.
More detail: On Apple silicon the GGUF formats only cut bytes; the math still runs at higher precision, which is why prefill stayed flat in the previous question. On H100, FP8 weights halve decode bytes and FP8 tensor cores roughly double matrix-math throughput, which helps prefill and large batches. The Quantization video: “FP8 on H100+ is close to free; four-bit needs measurement.”
The customer worries that a smaller, faster version of the model will give worse answers. What do you do?
Show answerHide answer
In plain words
Test it instead of arguing. Run 50 of their own real prompts at each size and show quality, speed and cost side by side.
Picture it
A bakery wants to switch to a cheaper flour. You bake their best-sellers both ways and let their own customers taste, instead of quoting a lab report.
With real numberslesson 04’s quality check and the Quantization video’s sample table
Lesson 04’s quick alarm: 12 questions per size. If the 3-bit version gets 7 right where the 8-bit gets 11, quality has fallen off a cliff.
For a customer: 50 of their own prompts at each size (lesson 09, on Day 7, builds the test).
The Quantization video’s example table (made-up numbers to show the shape): at 8-bit instead of 16-bit, accuracy slipped from 89% to 88%, total output rose from 24 to 46 tokens per second, conversations per GPU from 3 to 7.
The video’s rule: if 2 fewer replies in 100 come back in the exact format asked for, but the server writes twice as many tokens per second, that is usually a good trade; 20 fewer in 100 never is.
Words to know
Eval
A repeatable test of a model on a fixed set of prompts. Example: 50 of the customer’s prompts.
Precision
How many bits each stored number gets. Example: 16-bit, 8-bit, 4-bit.
Quality cliff
The size below which answers suddenly get worse. Example: the 3-bit version (Q3) slips on multi-step and format questions first.
Throughput
Total output across everyone at once. Example: 24 against 46 tokens/s in the Quantization video’s sample.
Go deeper: the engineer version
The kit's question
The customer is worried about quality. What do you do?
The kit's answer
Run their own 50-prompt eval at each precision and show the quality, latency and cost table from lesson 09.
More detail: Lesson 09’s table has valid, accuracy, judge, p95 and $/1k tasks columns, and says never to show one column alone. Pick the precision per workload, on held-out prompts. The Quantization video states its rule in schema validity and throughput.
Inside each layer of Llama 3.1 8B, 32 parts look back over the conversation. A design called GQA (grouped-query attention) makes them share 8 sets of notes instead of keeping 32. Why does that matter so much for memory?
Show answerHide answer
In plain words
Sharing makes the notes 4 times smaller. That one design choice is what makes long conversations affordable.
Picture it
32 reporters cover the same meeting. Instead of each writing a full transcript, they work in 8 teams of 4 and each team shares one transcript. Same meeting, a quarter of the paper.
With real numbersthe lesson 05 formula
Only one number in the notes formula changes: 8 note sets per layer instead of 32.
With 8: 128 KB per token. With 32: 512 KB, 4 times more.
One 131,072-token conversation (about 98,000 words): 17.18 GB with sharing, 68.72 GB without.
68.72 GB is 14 times the 4.9 GB model, and most of an 80 GB H100 (NVIDIA’s data-center AI chip).
Words to know
GQA (grouped-query attention)
Groups of the look-back parts (attention heads) share one set of notes. Example: 8 KV heads instead of 32.
KV head (note set)
One set of notes kept in each layer. Example: Llama 3.1 8B keeps 8 per layer.
Layer
One of the stacked processing stages each token passes through; each keeps its own notes. Example: 32 in Llama 3.1 8B.
H100
NVIDIA’s data-center GPU with 80 GB of memory. Example: the GPU in the KV Cache, GPU Bandwidth and Quantization videos.
Go deeper: the engineer version
The kit's question
Why do GQA models (8 KV heads instead of 32) matter so much here?
The kit's answer
The cache is 4× smaller. That architecture choice is what makes long context affordable.
More detail: GQA keeps 32 query heads but shares each K/V head across a group of 4 of them, so attention still looks at the text 32 ways while only 8 K/V heads are cached. Per token: 2 x 32 x 8 x 128 x 2 bytes = 128 KiB with GQA; 2 x 32 x 32 x 128 x 2 bytes = 512 KiB without. The KV Cache video puts the usual saving at 4 to 8 times; the field guide notes that Llama 3 70B needs only about 0.3 MB per token because of it.
A customer’s prompts are almost always short: 95 out of 100 are under 3,000 tokens (about 2,250 words). But they set the limit to 128K tokens (128 x 1,024 = 131,072 tokens, about 98,000 words) “to be safe”. What does that cost them?
Show answerHide answer
In plain words
On a server that sets aside room for the full limit up front, as llama.cpp does, they pay for memory they never use. The same GPU (the chip that runs the model) then serves far fewer people at once.
Picture it
A restaurant sets a 40-seat table for every party “to be safe”, but almost every party is 2 or 3 people. After a few parties the room is full while most chairs sit empty, and new guests wait at the door. Servers with PagedAttention (vLLM) add chairs as guests arrive, so they waste far less, though a few huge parties can still fill the room.
With real numbersLlama 3.1 8B, from lesson 05’s kv_calc.py
The model keeps notes on every token it has read: 128 KB per token.
Room for 128K tokens: 128 KB x 131,072 = about 17 GB per person. With 20 GB free, that is 1 person.
Room for 8K tokens (8,192, still more than double the 3,000 they use): 128 KB x 8,192 = 1.07 GB per person. 20 ÷ 1.07 = 18.7, so 18 people.
The fix: set the limit near real use, about 8K here.
Words to know
Token
A chunk of text, about three quarters of a word. 1,000 tokens is roughly 750 words.
KV cache
The model’s notes on the conversation so far, so it does not reread everything for each new word. Each person has their own, and it grows with every token.
p95
The size that 95 out of 100 requests stay under.
People served at once (concurrency)
How many conversations the GPU runs at the same moment.
Go deeper: the engineer version
The kit's question
The customer’s p95 prompt is 3k tokens but they configured 128k “to be safe”. What does that cost?
The kit's answer
Engines that reserve the maximum per slot strand memory. Capping at about 8k frees it for more concurrent users.
More detail: llama.cpp sets aside the whole -c cache at startup, split evenly across its --parallel slots (see lessons/02-three-local-servers/serve_llamacpp.sh). vLLM’s PagedAttention hands out cache in small pages as requests grow, so it wastes far less; a high maximum still lets a few very long requests take most of the memory. Check the sums: python kv_calc.py --ctx 131072 --free-gb 20 prints 17.18 GB and 1 conversation; python kv_calc.py --ctx 8192 --free-gb 20 prints 1.07 GB and 18.
Storing the conversation notes with 8 bits per number instead of 16 (an FP8 KV cache) halves their memory. What does it cost you?
Show answerHide answer
In plain words
A small risk that answers get a little worse on some tasks. You measure it on the customer’s own examples before you switch it on.
Picture it
Saving a photo at lower quality roughly halves the file, and usually nobody can tell. Sometimes the fine print blurs, so you compare both versions before sending it to a client.
With real numbersLlama 3.1 8B, lesson 05’s kv_calc.py
16-bit: 2 bytes per number. 8-bit (q8_0): 1 byte per number plus a small shared scaling number stored with each block of numbers, 1.0625 bytes in all.
One 131,072-token conversation (about 98,000 words): 16,384 MiB (17.18 GB) becomes 8,704 MiB (9.13 GB): 53%, not 50%, because of that extra.
20 GB of spare memory at 32,768 tokens: 4 conversations become 8.
The price: a small risk to answer quality on some tasks. Check it by running lesson 09’s test (Day 7) on 50 of the customer’s prompts, with 16-bit and with 8-bit notes, and compare.
Words to know
FP8 and q8_0
Two 8-bit formats; FP8 on NVIDIA data-center GPUs, q8_0 in llama.cpp on your Mac.
Bit
A single 0 or 1; 8 bits make a byte. Example: a 16-bit number takes 2 bytes.
Byte
8 bits. Example: one 16-bit number takes 2 bytes.
Eval harness
A repeatable test that scores a model on a fixed set of prompts. Example: lesson 09 (Day 7).
Go deeper: the engineer version
The kit's question
What does FP8 KV cost you?
The kit's answer
A small quality risk on some tasks. Check it with the eval harness in lesson 09.
More detail: q8_0 stores 8-bit integers in blocks that share a scale, which is where 1.0625 bytes per value comes from (kv_calc.py line 25); FP8 is the equivalent on H100-class GPUs and is exactly half of FP16. Run the same eval with the 16-bit and the 8-bit cache and compare before you promise the doubled capacity.
Under 20 seconds each. Record yourself once and listen back.
Step 1 · Quantization
QuestionShould we run this model in FP8 or 4-bit?
One clear answer
Decode speed is bandwidth over bytes per token. Quantization changes the denominator. I never recommend a precision without running the customer’s own eval at that precision.
What this means
“Decode speed”: How fast the model writes its reply, one token at a time.
“is bandwidth over bytes per token”: The top speed is how many GB per second memory delivers, divided by how many GB must be fetched for each token (the whole model). M4 Max: 546 ÷ 4.9 = about 111 tokens/s.
“Quantization changes the denominator”: Storing the model in fewer bits shrinks the bottom number, so the result goes up. 8.5 GB gives about 64 tokens/s; 4.9 GB gives about 111.
“I never recommend a precision without running the customer’s own eval at that precision”: Precision means how many bits each number gets. Before saying “use 4-bit”, I test each option on the customer’s own prompts, because fewer bits can cost accuracy. For this question: FP8 on an H100 usually costs little quality; the Quantization video calls it “close to free”. 4-bit buys more speed and room but carries more risk, so it needs that test most.
Step 2 · KV cache and context length
QuestionThe model fits on the GPU, but we still run out of memory under load. Why?
One clear answer
Weights fit; the cache doesn’t. I size KV per token from the config, cap context to the real p95, and quantize the cache. That usually doubles concurrency on the same GPU.
What this means
“Weights fit”: The model itself, its learned numbers (the weights), is 4.9 GB and fits easily.
“the cache doesn’t”: The per-person notes (the KV cache) run out; one 131,072-token conversation (about 98,000 words) needs 17.18 GB.
“I size KV per token from the config”: I work out how much memory each token of notes takes from the model’s settings file. For Llama 3.1 8B it is 128 KB.
“cap context to the real p95”: I set the maximum conversation length to what 95 of 100 real requests need, not the model’s maximum. For a customer whose requests are almost all under 3,000 tokens, a limit of 8,192 still leaves room to spare: in 20 GB it serves 18 people instead of 1 at 131,072.
“quantize the cache”: I store the notes with 8 bits per number instead of 16, like saving a photo at lower quality. It takes about half the space: 9.13 GB instead of 17.18 GB.
“doubles concurrency on the same GPU”: About twice as many people served at once, with no new hardware. In 20 GB at 32,768 tokens: 4 people become 8.
2 worked examples from today’s videos. Work each on paper, then reveal.
From the video:The KV Cache3:02
Whiteboard 1 of 2
An 8-billion-number model (Llama 3 8B), stored at 8 bits, runs on one H100 GPU with 80 GB of memory. Each conversation can hold up to 8K tokens (8 x 1,024 = 8,192 tokens, about 6,000 words). How many conversations fit at once? Which two cheap changes give 4 times as many?
Given, in plain words
The model is a stack of 32 layers (processing stages), and each keeps its own notes on every token. Per token, each layer stores 8 note sets (KV heads) of 128 numbers each (head dim), in two kinds: keys and values. Each note number takes 16 bits (2 bytes), even though the model itself is stored at 8 bits. Plan for the worst case: every conversation fills its full 8K.
Reveal the answerHide the answer
Answer · in plain words
About 60 conversations fit. After the model (8 GB) and some memory kept back, the video leaves about 64 GB for notes, and each conversation’s notes take about 1 GB. Store the notes in 8 bits and cap conversations at about 4,000 tokens, and about 240 fit on the same GPU.
Picture it
The spare memory is a wall of lockers with room for about 60 full-size ones. Each of the two changes halves the locker size, so 4 small lockers fit in the space of 1 big one.
Before you cap
Ask whether the 8K limit is a real product need or a default nobody questioned (the KV Cache video).
Worked answer, step by step
Notes per token: 2 (keys and values) x 32 layers x 8 KV heads x 128 numbers x 2 bytes = 131,072 bytes = 128 KB.
Notes per conversation: 131,072 bytes x about 8,000 tokens = 1.05 GB, about 1 GB (the video’s figure; exactly 8,192 tokens gives 1.07 GB, and the answer is the same).
The model: 8 billion numbers x 1 byte (8-bit) = 8 GB.
Memory left: 80 - 8 = 72 GB. The video plans on about 64 GB of that for notes. It does not say why; most likely the other 8 GB is a safety margin for the server program itself.
Conversations at once: 64 ÷ 1.05 = 61, so about 60.
Cheap change 1, store the notes in 8 bits: 1.05 ÷ 2 = 0.525 GB each, about half a GB. 64 ÷ 0.525 = 122, about 120 (twice 60).
Cheap change 2, cap conversations at 4,096 tokens: half again, 0.525 ÷ 2 = 0.26 GB each.
64 ÷ 0.26 = 246, about 240: 4 times 60. Same GPU, no new hardware.
Go deeper: the engineer version
The kit's question
Worked example · Llama 3 8B · 32 layers · 8 KV heads · head dim 128 · FP16. Here is the calculation that decides how many users fit. Two cheap changes move that number a long way.
The kit's answer
KV per token: 2 × 32 × 8 × 128 × 2 = 128 KB. 8K context: 1.05 GB per request. Weights (FP8): 8 GB. H100 memory: 80 GB. Memory left for KV: ≈ 64 GB. Capacity: About 60 concurrent 8K sessions — then you are full. Two changes, four times the users. FP16 KV · 8K: KV per request: 1.05 GB. Concurrent sessions: ≈ 60. FP8 KV · 4K: KV per request: 0.26 GB. Concurrent sessions: ≈ 240. Extra hardware: none. Quantising the cache to eight bits halves it. Capping context at four thousand tokens halves it again.
More detail: 1.05 GB equals 8,000 tokens x 131,072 bytes; with exactly 8,192 tokens kv_calc.py gives 1.07 GB, and 64 ÷ 1.07 = 59.6, still about 60. The video does not derive the 64 GB: it is 80 - 8 = 72 GB less about 8 GB, plausibly runtime working memory, which a server such as vLLM sets with --gpu-memory-utilization (the Serving With vLLM video, Day 2). FP8 KV is exactly half; llama.cpp’s q8_0 is 53%.
Words to know
Notes (KV cache)
The model’s memory of each conversation so far. Example: 128 KB per token here.
Layer
One of the stacked processing stages each token passes through; each keeps its own notes. Example: 32 in Llama 3 8B.
FP8
An 8-bit number format, 1 byte per number. Example: 8B model = 8 GB.
H100
NVIDIA’s data-center GPU, 80 GB of memory.
From the video:Quantization2:43
Whiteboard 2 of 2
You must run one 70-billion-number model on one H100 GPU (80 GB of memory), and each conversation can hold 8K tokens (8,192, about 6,000 words). Should you store the model at 16 bits, 8 bits or 4 bits?
Given, in plain words
Each conversation’s notes take about 2.6 GB, whatever size the model is stored at. Plan for the worst case: every conversation fills its full 8K.
Reveal the answerHide the answer
Answer · in plain words
Pick 4-bit if it passes the customer’s own quality test, because it fits the most people. At 16 bits the model needs 2 GPUs; at 8 bits about 3 conversations fit, at 4 bits about 14 to 16.
Picture it
The 80 GB GPU is a small apartment, and the model is a sofa sold in three sizes. The biggest will not fit through the door. The middle one leaves floor space for about 3 guests; the smallest, for about 14 to 16.
Worked answer, step by step
Model size = 70 billion numbers x bytes per number: 16-bit 70 x 2 = 140 GB; 8-bit 70 x 1 = 70 GB; 4-bit 70 x 0.5 = 35 GB.
16-bit: 140 GB is more than 80 GB, so it needs 2 GPUs before anyone connects (140 ÷ 80 = 1.75, round up to 2).
Keep about 1.5 GB for the server software itself (the lab book’s figure).
4-bit: 80 - 35 - 1.5 = 43.5 GB spare. 43.5 ÷ 2.6 = 16.7, so 16. The video plays safe and says about 14.
Either way, 4-bit holds about 5 times as many conversations as 8-bit (14 to 16, against 3).
Answer: pick the smallest size that passes the customer’s own quality test. How many bits you store decides how many people one GPU can serve.
Go deeper: the engineer version
The kit's question
Worked example · one 70B model, three precisions. H100 80GB · 8K context per session.
The kit's answer
FP16: 140 GB → 2 GPUs. FP8: 70 GB → 10 GB spare → ~3 sessions. 4-bit: 35 GB → 45 GB spare → ~14 sessions. KV per session: ≈ 2.6 GB at 8K. The real reason: Quantization buys concurrency, not just speed.
More detail: 2.6 GB is 2 x 80 layers x 8 KV heads x 128 x 2 bytes = 320 KiB per token, x 8,000 tokens (at 8,192, kv_calc.py --model llama-3.1-70b --ctx 8192 gives 2.68 GB). The video’s rows mix assumptions: the FP8 row reserves nothing (80 - 70 = 10 GB; 10 ÷ 2.6 = 3.8, “about 3”), while “about 14” at 4-bit implies about 8 GB reserved (45 - 8 = 37; 37 ÷ 2.6 = 14.2); with no reserve 4-bit gives 17 (45 ÷ 2.6 = 17.3). The lab book’s 1.5 GB framework overhead gives 3 and 16, as in the steps. Real 4-bit files carry per-group scales: the GPU Bandwidth video (Day 3) uses 38 GB for a 70B at 4-bit, giving (80 - 38) ÷ 2.6 = 16. An FP8 KV cache would halve the 2.6 GB and roughly double every session count.
Words to know
Parameters (70B, 8B)
The model’s learned numbers, also called weights. Example: 70B = 70 billion.
16-bit, 8-bit, 4-bit
How many bits store each number: 2 bytes, 1 byte, half a byte. Example: 70B is 140 GB at 16-bit, 35 GB at 4-bit.
Notes (KV cache)
The model’s memory of each conversation so far. Example: about 2.6 GB per 8K conversation here.
H100
NVIDIA’s data-center GPU, 80 GB of memory.
How you will use this
On Day 10 you answer “Can we run a 70B (a 70-billion-number model) on one GPU?” with today’s sums: the model’s size at 16, 8 or 4 bits, plus the notes each conversation needs. Your sizing memo (Day 10’s written hardware recommendation) also names today’s two fixes for when prompts get longer: limit how long a conversation may get, and store the notes in 8 bits.