Day 3 · Benchmark like a field engineer
Part 1 · Day 3 of 12
0 of 12 days done
About 1 h Free Mac
Today: Find out how many people one server can serve at once before newcomers start waiting. Test the kit’s practice server first, then the real model on your Mac. Then check your Mac against one formula: the fastest it can write for one person = memory speed (GB per second) ÷ model size (GB).
Example
Imagine eight customers asking a support bot for help at once. The practice server has eight places, so all eight can start. When sixteen customers arrive together, half must wait for a place; some wait over three seconds before a reply even starts. Today you will find the crowd size where that wait becomes unacceptable, so you know when one server is no longer enough.
By the end you’ll have
A chart of total output and first-token wait at each crowd size, for both servers, labelled with the test conditions.
Expected: The practice server keeps its promise to start most replies within half a second up to 8 people at once. At 16, some people have to wait for a free place. Your Mac test shows where the same crowding begins on a real model.
The 5 numbers you write down
Test conditions (traffic profile) to write above your chart
The conditions your curve was measured under. Without them a speed number means nothing, because longer prompts or replies move the knee. The sweep fixes how many people are served at once instead of a rate of requests per second, so write the crowd sizes. The server also wraps each prompt in a few tokens of its own.
Example: about 17 in / up to 128 out · 1 to 16 at once · p95 first token within 1 s
People one copy of the server holds within the promise (your Mac)
One server’s capacity. Divide the busiest moment’s crowd by it to get how many copies to run.
Example: Practice server: 8 people within 500 ms, so 120 people need 120 ÷ 8 = 15 copies, 18 with spares.
Total output: 1 person, and at your capacity (your Mac)
How much the server writes for everyone together. Sharing each pass over the model should multiply it several times before the seats fill.
Example: Practice server: 54.1 alone, 275.1 at 8 people, 5 times as much.
Speed each person sees: alone, and at your capacity (your Mac)
How fast one reply appears. It falls only a little as people share the server, because fetching the model is the slow part.
Example: Practice server: 54.9 alone, 38.9 at 8 people, 29% slower.
Your speed alone as a share of your Mac’s speed limit
The speed limit is memory speed ÷ model size (4.9 GB today). Landing at 60 to 85% of it means your measurement makes sense.
Example: An M4 Max Mac (546 GB/s) has a limit of 546 ÷ 4.9 = 111 tok/s. Day 4’s sample figures reach 86: 86 ÷ 111 = 0.77, so 77%.
How you will use this: A customer deciding how many servers to run needs this crowd-size result, with your test conditions written down. On Day 10 your sizing memo uses it to estimate server count and adds room for traffic spikes.
Before you start
The practice server is running.
Today’s first test talks to the kit’s practice server: a stand-in that behaves like a real server, with 8 seats (it serves 8 conversations at a time) and made-up numbers. From the
03-labsfolder, with the Python environment on (source .venv/bin/activate), start it in the background (it keeps running while you type other commands):make mock &It prints a line with
slots=8, its 8 seats. If it says the address is already in use, it is still running from an earlier day: that is fine.Day 2’s llama.cpp server is stopped.
Today’s step restarts llama.cpp with 8 seats. It uses the model file Day 2 downloaded (Llama 3.1 8B, a model of 8 billion numbers, stored at 4 bits: 4.9 GB), so nothing new downloads today. If Day 2’s server is still running, press Ctrl+C in its terminal. The new one needs the same port (the numbered door a server listens on), 8080.
Your Mac’s speed limit, written down.
The kit’s check script prints your Mac’s memory speed and the top writing speed it allows for today’s 4.9 GB model. You compare your measured speed with it at the end of the step.
From the
03-labsfolder:make checkLook for the
bandwidthline: for example~273 GB/s(memory speed, GB per second) and~56 tok/s(the top writing speed, tokens per second) on an M4 Pro Mac. Themockline should say UP.
Warm-up from Day 2
Section titled “Warm-up from Day 2”About 2 minutes. Say your answer out loud, then tap to check it. From Day 2 · Local stack on the Mac.
You start llama.cpp with -c 8192 (room for 8,192 tokens, about 6,000 words) and --parallel 4 (4 seats, so 4 people at once). Why is that a trap?Show answerHide
In plain words
The 8,192 tokens are shared out, not given to each person. Each of the 4 seats gets only 2,048 tokens (about 1,500 words), so a longer conversation does not fit.
Picture it
One large pizza ordered ‘for 4 people’. The box says large, but each person gets a quarter, and anyone hungrier than a quarter goes without.
With real numbersLlama 3.1 8B on llama.cpp with Day 2’s settings, lesson 02
- Room for notes:
-c 8192= 8,192 tokens, about 6,000 words. - Seats:
--parallel 4= 4 conversations at once. - Each seat: 8,192 ÷ 4 = 2,048 tokens, about 1,500 words.
- Day 1’s long test prompt, about 8,000 tokens, is almost 4 times one seat’s room (8,000 ÷ 2,048 = 3.9), so it does not fit.
- The fix: set
-cto seats x tokens each. For 4 seats of 8,192 tokens: 4 x 8,192 = 32,768, which takes 4.29 GB of notes for this model (Day 4 works this out).
Words to know
- Context length (
-c) - The most tokens the server keeps notes for. Example:
-c 8192= 8,192 tokens. - Slot (
--parallel) - One seat on the server: one conversation it serves at the same time. Example:
--parallel 4= 4 seats. - Token
- A chunk of text, about three quarters of a word. Example: 2,048 tokens is about 1,500 words.
- KV cache (notes)
- The model’s notes on each conversation so far, kept in memory. Example: 32,768 tokens of notes take 4.29 GB for Llama 3.1 8B.
Go deeper: the engineer version
The kit's question
Why is --parallel 4 -c 8192 a trap?
The kit's answer
llama.cpp splits the context across slots, so each user gets only 2,048 tokens.
More detail: llama-server sets aside one KV cache of -c tokens at startup and splits it evenly across the --parallel slots (the comment in serve_llamacpp.sh says so), so each slot’s context is c ÷ parallel. Size -c as slots x the longest context one user needs: 4 x 8,192 = 32,768, which is 4,096 MiB of f16 cache for Llama 3.1 8B at 128 KiB per token (lesson 05). llama-server’s startup log lists the context each slot gets, so confirm it there on your build.
A customer runs Ollama, the easy one-command server, for their live service with real users (production). What would you ask them about?Show answerHide
In plain words
Ask how many people use it at the same moment, and what happens when many arrive at once. Ask how they watch its health. Then ask whether a server built for many users, such as vLLM or SGLang (servers for NVIDIA chips), would serve them better.
Picture it
A home espresso machine makes great coffee, one cup at a time. If a busy café runs on one, you ask how many customers come at the morning rush, how long the queue gets, and who notices when it breaks. At some point they need a café machine built to pour hundreds of cups an hour.
With real numbersthe Serving With vLLM video (Day 2) and lesson 02
- Day 2’s llama.cpp script shows its seats:
--parallel 4, so 4 people at once with 2,048 tokens each. Ollama chooses settings like these for you, out of sight. - In the Serving With vLLM video, vLLM (a server built for many users) prints its room for notes at startup: 8,234 blocks of 16 tokens each. A block is a fixed-size piece of note memory.
- 8,234 x 16 = 131,744 tokens of notes. At 8,192 tokens per conversation: 131,744 ÷ 8,192 = 16.08, so 16 people at once if everyone uses the full length.
- In a separate example in the same video, three settings took one vLLM machine from 16 people at once to 50: 50 ÷ 16 = about 3 times as many.
- So ask: how many people at the busiest moment, and do numbers like these cover it?
Words to know
- Production
- A model running as a live service for real users. Example: the customer’s support chat.
- People served at once (concurrency)
- How many conversations the server runs at the same moment. Example: 16 before tuning, 50 after, in the Serving With vLLM video.
- Batching
- Serving several people with one pass over the model. Example: 4 seats share each pass in Day 2’s llama.cpp server.
- Observability
- Seeing what a live service is doing, such as speed, errors and memory, while it runs. Example: llama.cpp’s
--metricsflag.
Go deeper: the engineer version
The kit's question
A customer runs Ollama in production. What would you ask about?
The kit's answer
Concurrency, batching limits, observability, and whether vLLM or SGLang would suit multi-user traffic better.
More detail: Turn each topic into a number. Concurrency: peak simultaneous users. Batching limits: how many requests the engine batches (its parallel slots) and the context each slot gets, which Ollama sets by defaults the operator may never have read. Observability: TTFT and ITL at p50 and p95, errors and memory, live (llama.cpp’s --metrics flag publishes counters at /metrics). Then run the same harness against vLLM or SGLang, which add PagedAttention and the tuning flags from the Serving With vLLM video; Day 9’s optional Spark step does exactly that.
Why report p95 (the time 95 out of 100 requests stay under) and not the average?Show answerHide
In plain words
An average hides the few very slow requests, and those are the ones users remember and complain about.
Picture it
A bus that is on time 19 days out of 20 and an hour late on the 20th has an average delay of only 3 minutes. The rider remembers the day they missed a meeting.
With real numbersDay 1’s stopwatch script (felab/measure.py) and its sample run
- The stopwatch script times 10 requests. p50 is the middle one; with only 10, p95 is the slowest one.
- Lesson sample, short prompt: p50 is 145 ms, p95 is 190 ms.
- So the worst request waited 31% longer than the typical one (190 ÷ 145 = 1.31).
- Inference 101: a fine p50 with an ugly p95 almost always means waiting in line or servers waking up (cold starts), not the model.
Words to know
- p50 (median)
- The middle value: half the requests are faster, half slower. Example: 145 ms.
- p95
- The value 95 out of 100 requests stay under. Example: 190 ms.
- Tail
- The slowest few requests, the ones p95 describes.
- Cold start
- A slow request while a server that was idle or newly started loads the model.
Go deeper: the engineer version
The kit's question
Why p95 and not the average?
The kit's answer
Averages hide the tail, and the tail is what users complain about.
More detail: felab.percentile uses the nearest-rank method, so with 10 samples p95 is the maximum: deliberately pessimistic, because tails are what customers complain about. Agree the percentile and the traffic profile with the customer before you benchmark (Inference 101 recap).
Today’s steps
Section titled “Today’s steps”1 step, then the wrap-up.
-
Concurrency sweep
You measure how many people one server can serve at once before the wait for a reply to start passes the promised limit. That number, not a single speed figure, decides how many servers a customer needs.
Videos: Benchmarking Properly 2:22 · Batching & Scheduling 2:58 · GPU Bandwidth 3:10
- Wrap-up and drill Put your numbers in one table, answer 3 questions out loud, explain 1 result and work 2 examples.
Your progress is saved on this device.