Day 4 · Memory levers
Part 1 · Day 4 of 12
0 of 12 days done
About 1 h 50 17 GB model download · See Step 1 Free Mac
Today: Shrink a model (billions of learned numbers) by storing each number with fewer bits (the 0s and 1s it is kept in). See what that buys (faster replies, more free memory) and what it can cost (answer quality). Then measure how much memory long conversations use, and fix a server that runs out.
Example
Imagine a support bot reading a very long case history. It keeps working notes so it does not have to reread the whole conversation for every new word it writes. Those notes share the chip’s memory with the model. In this lesson’s extreme long-conversation example, the notes take more than three times as much space as the model itself. That is why a model can fit in memory while the conversation still runs out of room.
By the end you’ll have
A count of how many long conversations fit in 20 GB of spare memory, before and after the fix. The fix stores the notes with 8 bits per number instead of 16, which about halves them.
Expected: 4 conversations of 32,768 tokens (about 25,000 words each) before, 8 after.
The 7 numbers you write down
Writing speed at 8-bit, 4-bit and 3-bit (tokens per second)
How fast each size writes its reply. Your speed limit is your Mac’s memory speed ÷ the model’s size: an M4 Pro moves 273 GB/s, so its limit for the 4.9 GB model is 273 ÷ 4.9 = about 56 tokens per second. Find your chip under Apple menu, About This Mac;
ceiling_check.pylooks up its memory speed for you.Example: 52 / 86 / 98 (the lesson’s sample figures for an M4 Max, 546 GB/s)
Share of the speed limit reached, same three sizes (%)
Your writing speed ÷ the speed limit (memory speed ÷ model size). It shows the formula holds on your own machine.
Example: 81 / 77 / 72 (at 4-bit: 86 ÷ 111.4 = 77%)
Quality check score at 8-bit, 4-bit and 3-bit (out of 12)
How many of 12 quick questions each size got right: a smoke alarm for the quality cliff, not a full test. A customer gets 50 of their own prompts instead.
Example: The lesson’s alarm example: 11 at 8-bit against 7 at 3-bit would be a cliff.
Notes for a 32,768-token conversation (about 25,000 words), 16-bit and 8-bit (MiB)
How much memory the notes for one 32,768-token conversation take, stored at 16 bits and at 8 bits. It is 8 times the 4,096-token figure (8 times the tokens), and the 8-bit figure is about half the 16-bit one. MiB is the unit llama.cpp prints, about a million bytes; 1,000 MiB is about 1 GB.
Example: 4,096 / 2,176 MiB (4.29 / 2.28 GB)
Notes for one 131,072-token conversation (about 98,000 words), 16-bit (MiB)
How much memory the notes take for the longest conversation the model allows, at 16 bits per number. It is more than three times the 4.9 GB model.
Example: 16,384 MiB (17.18 GB), or “failed” on a 16 to 24 GB Mac
Notes for one 131,072-token conversation, 8-bit (MiB)
The same notes stored with 8 bits per number: the fix. About half the memory, so about twice as many conversations fit.
Example: 8,704 MiB (9.13 GB)
Conversations of 32,768 tokens that fit in 20 GB
How many people one machine serves at once with 20 GB of memory left over after the model loads, before and after the fix.
Example: 4 / 8
How you will use this: On Day 10 you answer “Can we run a 70B (a 70-billion-number model) on one GPU?” with today’s sums: the model’s size at 16, 8 or 4 bits, plus the notes each conversation needs. Your sizing memo (Day 10’s written hardware recommendation) also names today’s two fixes for when prompts get longer: limit how long a conversation may get, and store the notes in 8 bits.
Before you start
About 17 GB of free disk space (about 30 GB if you keep every download).
Step 1 downloads the same model (Llama 3.1 8B, made of 8 billion learned numbers) three times, each shrunk to a different size: 8.5 + 4.9 + 4.0 = 17.4 GB. Start that download first; it takes the longest. Later the quality check downloads the 8-bit and 3-bit versions again (8.5 + 4.0 = 12.5 GB) into llama.cpp’s own folder. So once step 1’s speed test (
quant_bench.sh) has printed its table, deletelessons/04-quantization/models/, or have about 30 GB free.The llama.cpp server from Day 2 still starts.
Today’s scripts start llama.cpp (the program that runs the model on your Mac) many times. Quick check, from the
03-labsfolder:bash lessons/02-three-local-servers/serve_llamacpp.shWait for a line saying it is listening (ready for requests). Then stop it with Ctrl+C.
Warm-up from Day 3
Section titled “Warm-up from Day 3”About 2 minutes. Say your answer out loud, then tap to check it. From Day 3 · Benchmark like a field engineer.
When 8 people used the practice server at once instead of 1, its total output rose about 5 times. Yet each person’s reply only slowed from 55 to 39 tokens (word pieces) per second. Why does sharing cost so little?Show answerHide
In plain words
To write each token, the chip must fetch the whole model from memory, and that fetch is the slow part. One fetch serves all 8 people at once, so each extra person adds only a little.
Picture it
A bus takes about the same time to drive its route with 1 rider or with 8. The drive is the slow part; letting a few more people on adds a little time at each stop. This stops working once the seats run out: then extra riders wait for the next bus, and that point is the knee.
With real numberslesson 03’s practice server: it behaves like a real one, but its numbers are made up for teaching
- 1 person alone: 54.1 tokens per second in total; that person sees about 55 (54.9) once the reply is flowing.
- 8 people: 275.1 in total, about 5 times the 54.1 (275.1 ÷ 54.1 = 5.1).
- Each person: 54.9 falls to 38.9 tokens per second, 29% slower.
- So each person keeps 71% of their solo speed while the server does 5 times the work.
Words to know
- Decode
- The writing phase; the model writes one token at a time and fetches the whole model for each. Example: 54.9 tokens per second for one person on the practice server.
- Memory-bound
- Slowed by how fast memory delivers data, not by how fast the chip calculates. Example: decode.
- Batching
- Serving several people with one pass over the model. Example: 8 people share each fetch.
- Tokens per second (tok/s)
- How many tokens are written each second. Example: 38.9 per person with 8 people.
Go deeper: the engineer version
The kit's question
Throughput went up 5×, but per-user speed only dropped from 55 to 39 tok/s. Why so cheap?
The kit's answer
Decode is memory-bound, so sharing one weight read across 8 users costs little extra.
More detail: Each decode step streams all the weights once, whatever the batch size. Each extra user still adds its own KV-cache reads and a little compute, which is why per-user speed falls somewhat (54.9 to 38.9). The practice server builds this in on purpose: one step takes 18 ms alone and grows 6% per extra user, so 18 x 1.42 = 25.6 ms with 8. That is 1,000 ÷ 18 = 55.6 tokens per second alone and 1,000 ÷ 25.56 = 39.1 with 8 (felab/mock_server.py, _itl). In roofline terms, one user’s decode does about 1 to 2 calculations per byte fetched, while an H100 needs about 300 to keep its math busy; batching raises that ratio. The 1-user total (54.1) sits a little under the per-user figure (54.9) because the per-user figure is the median speed once text starts flowing, while the total divides all tokens by wall-clock time, which includes the wait before the first token.
At the busiest moment, 120 people use the service at once. One copy of the server (a replica) handles 8 people while still keeping its speed promise (the SLO). How many copies do you need?Show answerHide
In plain words
120 ÷ 8 = 15 copies covers the busiest moment exactly. Add spare copies for surprises, so plan for about 18.
Picture it
A ride seats 8 per car, and 120 people want to ride at the same moment. You need 15 cars, plus a few spare in case one breaks down or a bigger crowd shows up.
With real numberslesson 03, and the rule in the Day 10 sizing memo
- Busiest moment: 120 people at once.
- One copy keeps its promise up to 8 people (the practice server’s limit).
- 120 ÷ 8 = 15 copies.
- The lesson’s answer, 18, adds 3 spare copies: 20% extra (15 x 1.2 = 18). The Day 10 sizing memo adds the same 20%.
Words to know
- Replica
- One complete running copy of the server and model. Example: 18 replicas for 120 people.
- SLO (service level objective)
- The speed promise. Example: 95 of 100 requests get their first token within 500 ms (half a second).
- Peak traffic
- The most people using the service at the same moment. Example: 120.
- Headroom
- Spare capacity kept for surprises. Example: plan 18 where 15 are needed.
Go deeper: the engineer version
The kit's question
Peak traffic is 120 concurrent users and one replica handles 8 within SLO. How many replicas?
The kit's answer
15, plus headroom. Say 18.
More detail: Replicas = peak concurrency ÷ per-replica capacity at the SLO, rounded up, plus headroom for traffic spikes, a replica restarting and forecast error. Lesson 15’s build_memo.py computes exactly this, math.ceil(a.peak_concurrency / per_replica * 1.2): 120 ÷ 8 x 1.2 = 18. Measure the per-replica capacity with the customer’s real prompt and reply lengths, because longer contexts move the knee.
On the sweep chart, the knee is the point where all the server’s seats are taken and new people start waiting in line. What moves that point?Show answerHide
In plain words
The knee moves to more people when more fit at once: more seats, more memory set aside for the notes (the KV cache), a smaller model, or shorter conversations. It moves to fewer people when one long prompt hogs the chip while it is read.
Picture it
A restaurant is full at a certain number of guests. More tables, more floor space or shorter meals seat more people; one huge order hogging the kitchen slows every other table.
With real numberslesson 03’s practice server and llama.cpp settings, and lessons 04 and 05 (Day 4)
- Practice server: 8 seats, so the knee comes right after 8 people.
- At 8 people, 95 of 100 got their first token within 41.4 ms (thousandths of a second); at 16, within 3,357 ms (over 3 seconds): about 80 times longer.
- Seats share memory: llama.cpp splits 16,384 tokens of conversation room across 8 seats, 2,048 tokens (about 1,500 words) each. More seats means less room for each conversation.
- Memory for notes sets the seats. Day 4’s calculator: 20 GB of spare memory holds 4 conversations of 32,768 tokens (about 25,000 words) with notes at 16 bits per number, and 8 at 8 bits.
- A smaller model frees memory: Llama 3.1 8B at 4 bits is 4.9 GB against 8.5 GB at 8 bits, 3.6 GB more for notes (Day 4).
Words to know
- Knee
- The point where seats run out and waiting time shoots up. Example: after 8 people on the practice server.
- Slot
- One seat on the server, a conversation it can serve at the same time. Example: the practice server has 8.
- Notes (KV cache)
- The model’s memory of each conversation so far, kept in the same memory as the model. Example: 4 conversations of 32,768 tokens fit in 20 GB.
- Prefill
- The reading phase; the model reads the whole prompt at once before writing. Example: a 100,000-token prompt takes a long prefill.
Go deeper: the engineer version
The kit's question
What changes the knee?
The kit's answer
More slots or memory for KV cache, a smaller or quantized model, shorter contexts, and prefill/decode interference.
More detail: In llama.cpp the -c context is split across the --parallel slots (the comment in serve_llamacpp.sh says so), so slots and context length trade off directly. A quantized model frees memory for KV cache and also makes each decode step faster. Long prefills stall the decode steps of everyone batched with them; chunked prefill (Batching & Scheduling video) is the usual fix. The 41.4 and 3,357 ms figures are TTFT p95.
Today’s steps
Section titled “Today’s steps”2 steps, then the wrap-up.
-
Quantization
Quantization stores a model’s numbers with fewer bits (the 0s and 1s each number is kept in), so the model gets smaller. You store one model at three sizes, see that the smaller one writes faster, then check it still answers well.
Video: Quantization 2:43
-
KV cache and context length
You watch the model’s notes on one long conversation (the KV cache) outgrow the model itself, then halve them by storing them in 8 bits. Running out of memory for these notes is the most common failure in real AI services.
Video: The KV Cache 3:02
- Wrap-up and drill Put your numbers in one table, answer 6 questions out loud, explain 2 results and work 2 examples.
Your progress is saved on this device.