Skip to content

Day 3 · Benchmark like a field engineer

Part 1 · Day 3 of 12

0 of 12 days done

About 1 h Free Mac

Today: Find out how many people one server can serve at once before newcomers start waiting. Test the kit’s practice server first, then the real model on your Mac. Then check your Mac against one formula: the fastest it can write for one person = memory speed (GB per second) ÷ model size (GB).

Example

Imagine eight customers asking a support bot for help at once. The practice server has eight places, so all eight can start. When sixteen customers arrive together, half must wait for a place; some wait over three seconds before a reply even starts. Today you will find the crowd size where that wait becomes unacceptable, so you know when one server is no longer enough.

By the end you’ll have

A chart of total output and first-token wait at each crowd size, for both servers, labelled with the test conditions.

Expected: The practice server keeps its promise to start most replies within half a second up to 8 people at once. At 16, some people have to wait for a free place. Your Mac test shows where the same crowding begins on a real model.

The 5 numbers you write down

  1. Test conditions (traffic profile) to write above your chart

    The conditions your curve was measured under. Without them a speed number means nothing, because longer prompts or replies move the knee. The sweep fixes how many people are served at once instead of a rate of requests per second, so write the crowd sizes. The server also wraps each prompt in a few tokens of its own.

    Example: about 17 in / up to 128 out · 1 to 16 at once · p95 first token within 1 s

  2. People one copy of the server holds within the promise (your Mac)

    One server’s capacity. Divide the busiest moment’s crowd by it to get how many copies to run.

    Example: Practice server: 8 people within 500 ms, so 120 people need 120 ÷ 8 = 15 copies, 18 with spares.

  3. Total output: 1 person, and at your capacity (your Mac)

    How much the server writes for everyone together. Sharing each pass over the model should multiply it several times before the seats fill.

    Example: Practice server: 54.1 alone, 275.1 at 8 people, 5 times as much.

  4. Speed each person sees: alone, and at your capacity (your Mac)

    How fast one reply appears. It falls only a little as people share the server, because fetching the model is the slow part.

    Example: Practice server: 54.9 alone, 38.9 at 8 people, 29% slower.

  5. Your speed alone as a share of your Mac’s speed limit

    The speed limit is memory speed ÷ model size (4.9 GB today). Landing at 60 to 85% of it means your measurement makes sense.

    Example: An M4 Max Mac (546 GB/s) has a limit of 546 ÷ 4.9 = 111 tok/s. Day 4’s sample figures reach 86: 86 ÷ 111 = 0.77, so 77%.

How you will use this: A customer deciding how many servers to run needs this crowd-size result, with your test conditions written down. On Day 10 your sizing memo uses it to estimate server count and adds room for traffic spikes.

Before you start

  1. The practice server is running.

    Today’s first test talks to the kit’s practice server: a stand-in that behaves like a real server, with 8 seats (it serves 8 conversations at a time) and made-up numbers. From the 03-labs folder, with the Python environment on (source .venv/bin/activate), start it in the background (it keeps running while you type other commands):

    make mock &

    It prints a line with slots=8, its 8 seats. If it says the address is already in use, it is still running from an earlier day: that is fine.

  2. Day 2’s llama.cpp server is stopped.

    Today’s step restarts llama.cpp with 8 seats. It uses the model file Day 2 downloaded (Llama 3.1 8B, a model of 8 billion numbers, stored at 4 bits: 4.9 GB), so nothing new downloads today. If Day 2’s server is still running, press Ctrl+C in its terminal. The new one needs the same port (the numbered door a server listens on), 8080.

  3. Your Mac’s speed limit, written down.

    The kit’s check script prints your Mac’s memory speed and the top writing speed it allows for today’s 4.9 GB model. You compare your measured speed with it at the end of the step.

    From the 03-labs folder:

    make check

    Look for the bandwidth line: for example ~273 GB/s (memory speed, GB per second) and ~56 tok/s (the top writing speed, tokens per second) on an M4 Pro Mac. The mock line should say UP.

About 2 minutes. Say your answer out loud, then tap to check it. From Day 2 · Local stack on the Mac.

You start llama.cpp with -c 8192 (room for 8,192 tokens, about 6,000 words) and --parallel 4 (4 seats, so 4 people at once). Why is that a trap?Show answerHide

In plain words

The 8,192 tokens are shared out, not given to each person. Each of the 4 seats gets only 2,048 tokens (about 1,500 words), so a longer conversation does not fit.

Picture it

One large pizza ordered ‘for 4 people’. The box says large, but each person gets a quarter, and anyone hungrier than a quarter goes without.

With real numbersLlama 3.1 8B on llama.cpp with Day 2’s settings, lesson 02

  • Room for notes: -c 8192 = 8,192 tokens, about 6,000 words.
  • Seats: --parallel 4 = 4 conversations at once.
  • Each seat: 8,192 ÷ 4 = 2,048 tokens, about 1,500 words.
  • Day 1’s long test prompt, about 8,000 tokens, is almost 4 times one seat’s room (8,000 ÷ 2,048 = 3.9), so it does not fit.
  • The fix: set -c to seats x tokens each. For 4 seats of 8,192 tokens: 4 x 8,192 = 32,768, which takes 4.29 GB of notes for this model (Day 4 works this out).

Words to know

Context length (-c)
The most tokens the server keeps notes for. Example: -c 8192 = 8,192 tokens.
Slot (--parallel)
One seat on the server: one conversation it serves at the same time. Example: --parallel 4 = 4 seats.
Token
A chunk of text, about three quarters of a word. Example: 2,048 tokens is about 1,500 words.
KV cache (notes)
The model’s notes on each conversation so far, kept in memory. Example: 32,768 tokens of notes take 4.29 GB for Llama 3.1 8B.
Go deeper: the engineer version

The kit's question

Why is --parallel 4 -c 8192 a trap?

The kit's answer

llama.cpp splits the context across slots, so each user gets only 2,048 tokens.

More detail: llama-server sets aside one KV cache of -c tokens at startup and splits it evenly across the --parallel slots (the comment in serve_llamacpp.sh says so), so each slot’s context is c ÷ parallel. Size -c as slots x the longest context one user needs: 4 x 8,192 = 32,768, which is 4,096 MiB of f16 cache for Llama 3.1 8B at 128 KiB per token (lesson 05). llama-server’s startup log lists the context each slot gets, so confirm it there on your build.

A customer runs Ollama, the easy one-command server, for their live service with real users (production). What would you ask them about?Show answerHide

In plain words

Ask how many people use it at the same moment, and what happens when many arrive at once. Ask how they watch its health. Then ask whether a server built for many users, such as vLLM or SGLang (servers for NVIDIA chips), would serve them better.

Picture it

A home espresso machine makes great coffee, one cup at a time. If a busy café runs on one, you ask how many customers come at the morning rush, how long the queue gets, and who notices when it breaks. At some point they need a café machine built to pour hundreds of cups an hour.

With real numbersthe Serving With vLLM video (Day 2) and lesson 02

  • Day 2’s llama.cpp script shows its seats: --parallel 4, so 4 people at once with 2,048 tokens each. Ollama chooses settings like these for you, out of sight.
  • In the Serving With vLLM video, vLLM (a server built for many users) prints its room for notes at startup: 8,234 blocks of 16 tokens each. A block is a fixed-size piece of note memory.
  • 8,234 x 16 = 131,744 tokens of notes. At 8,192 tokens per conversation: 131,744 ÷ 8,192 = 16.08, so 16 people at once if everyone uses the full length.
  • In a separate example in the same video, three settings took one vLLM machine from 16 people at once to 50: 50 ÷ 16 = about 3 times as many.
  • So ask: how many people at the busiest moment, and do numbers like these cover it?

Words to know

Production
A model running as a live service for real users. Example: the customer’s support chat.
People served at once (concurrency)
How many conversations the server runs at the same moment. Example: 16 before tuning, 50 after, in the Serving With vLLM video.
Batching
Serving several people with one pass over the model. Example: 4 seats share each pass in Day 2’s llama.cpp server.
Observability
Seeing what a live service is doing, such as speed, errors and memory, while it runs. Example: llama.cpp’s --metrics flag.
Go deeper: the engineer version

The kit's question

A customer runs Ollama in production. What would you ask about?

The kit's answer

Concurrency, batching limits, observability, and whether vLLM or SGLang would suit multi-user traffic better.

More detail: Turn each topic into a number. Concurrency: peak simultaneous users. Batching limits: how many requests the engine batches (its parallel slots) and the context each slot gets, which Ollama sets by defaults the operator may never have read. Observability: TTFT and ITL at p50 and p95, errors and memory, live (llama.cpp’s --metrics flag publishes counters at /metrics). Then run the same harness against vLLM or SGLang, which add PagedAttention and the tuning flags from the Serving With vLLM video; Day 9’s optional Spark step does exactly that.

Why report p95 (the time 95 out of 100 requests stay under) and not the average?Show answerHide

In plain words

An average hides the few very slow requests, and those are the ones users remember and complain about.

Picture it

A bus that is on time 19 days out of 20 and an hour late on the 20th has an average delay of only 3 minutes. The rider remembers the day they missed a meeting.

With real numbersDay 1’s stopwatch script (felab/measure.py) and its sample run

  • The stopwatch script times 10 requests. p50 is the middle one; with only 10, p95 is the slowest one.
  • Lesson sample, short prompt: p50 is 145 ms, p95 is 190 ms.
  • So the worst request waited 31% longer than the typical one (190 ÷ 145 = 1.31).
  • Inference 101: a fine p50 with an ugly p95 almost always means waiting in line or servers waking up (cold starts), not the model.

Words to know

p50 (median)
The middle value: half the requests are faster, half slower. Example: 145 ms.
p95
The value 95 out of 100 requests stay under. Example: 190 ms.
Tail
The slowest few requests, the ones p95 describes.
Cold start
A slow request while a server that was idle or newly started loads the model.
Go deeper: the engineer version

The kit's question

Why p95 and not the average?

The kit's answer

Averages hide the tail, and the tail is what users complain about.

More detail: felab.percentile uses the nearest-rank method, so with 10 samples p95 is the maximum: deliberately pessimistic, because tails are what customers complain about. Agree the percentile and the traffic profile with the customer before you benchmark (Inference 101 recap).

Start step 1: Concurrency sweep (50 min)

Your progress is saved on this device.