Skip to content

Day 3 wrap-up and drill

  1. Overview
  2. Step 1
  3. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

You have a chart of total output and first-token wait at each crowd size, for the practice server and your Mac, with its test conditions written down. You also know how many people one copy of the server holds within the promise.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 1 · Concurrency sweep

The conditions your curve was measured under. Without them a speed number means nothing, because longer prompts or replies move the knee. The sweep fixes how many people are served at once instead of a rate of requests per second, so write the crowd sizes. The server also wraps each prompt in a few tokens of its own.

Example: about 17 in / up to 128 out · 1 to 16 at once · p95 first token within 1 s

One server’s capacity. Divide the busiest moment’s crowd by it to get how many copies to run.

Example: Practice server: 8 people within 500 ms, so 120 people need 120 ÷ 8 = 15 copies, 18 with spares.

How much the server writes for everyone together. Sharing each pass over the model should multiply it several times before the seats fill.

Example: Practice server: 54.1 alone, 275.1 at 8 people, 5 times as much.

How fast one reply appears. It falls only a little as people share the server, because fetching the model is the slow part.

Example: Practice server: 54.9 alone, 38.9 at 8 people, 29% slower.

The speed limit is memory speed ÷ model size (4.9 GB today). Landing at 60 to 85% of it means your measurement makes sense.

Example: An M4 Max Mac (546 GB/s) has a limit of 546 ÷ 4.9 = 111 tok/s. Day 4’s sample figures reach 86: 86 ÷ 111 = 0.77, so 77%.

Every question from today’s steps. Say your answer out loud, then show the answer and rate yourself. 3 cards

QuestionStep 1 · Concurrency sweep

When 8 people used the practice server at once instead of 1, its total output rose about 5 times. Yet each person’s reply only slowed from 55 to 39 tokens (word pieces) per second. Why does sharing cost so little?

Show answerHide answer

In plain words

To write each token, the chip must fetch the whole model from memory, and that fetch is the slow part. One fetch serves all 8 people at once, so each extra person adds only a little.

Picture it

A bus takes about the same time to drive its route with 1 rider or with 8. The drive is the slow part; letting a few more people on adds a little time at each stop. This stops working once the seats run out: then extra riders wait for the next bus, and that point is the knee.

With real numberslesson 03’s practice server: it behaves like a real one, but its numbers are made up for teaching

  • 1 person alone: 54.1 tokens per second in total; that person sees about 55 (54.9) once the reply is flowing.
  • 8 people: 275.1 in total, about 5 times the 54.1 (275.1 ÷ 54.1 = 5.1).
  • Each person: 54.9 falls to 38.9 tokens per second, 29% slower.
  • So each person keeps 71% of their solo speed while the server does 5 times the work.

Words to know

Decode
The writing phase; the model writes one token at a time and fetches the whole model for each. Example: 54.9 tokens per second for one person on the practice server.
Memory-bound
Slowed by how fast memory delivers data, not by how fast the chip calculates. Example: decode.
Batching
Serving several people with one pass over the model. Example: 8 people share each fetch.
Tokens per second (tok/s)
How many tokens are written each second. Example: 38.9 per person with 8 people.
Go deeper: the engineer version

The kit's question

Throughput went up 5×, but per-user speed only dropped from 55 to 39 tok/s. Why so cheap?

The kit's answer

Decode is memory-bound, so sharing one weight read across 8 users costs little extra.

More detail: Each decode step streams all the weights once, whatever the batch size. Each extra user still adds its own KV-cache reads and a little compute, which is why per-user speed falls somewhat (54.9 to 38.9). The practice server builds this in on purpose: one step takes 18 ms alone and grows 6% per extra user, so 18 x 1.42 = 25.6 ms with 8. That is 1,000 ÷ 18 = 55.6 tokens per second alone and 1,000 ÷ 25.56 = 39.1 with 8 (felab/mock_server.py, _itl). In roofline terms, one user’s decode does about 1 to 2 calculations per byte fetched, while an H100 needs about 300 to keep its math busy; batching raises that ratio. The 1-user total (54.1) sits a little under the per-user figure (54.9) because the per-user figure is the median speed once text starts flowing, while the total divides all tokens by wall-clock time, which includes the wait before the first token.

How did you do?

QuestionStep 1 · Concurrency sweep

At the busiest moment, 120 people use the service at once. One copy of the server (a replica) handles 8 people while still keeping its speed promise (the SLO). How many copies do you need?

Show answerHide answer

In plain words

120 ÷ 8 = 15 copies covers the busiest moment exactly. Add spare copies for surprises, so plan for about 18.

Picture it

A ride seats 8 per car, and 120 people want to ride at the same moment. You need 15 cars, plus a few spare in case one breaks down or a bigger crowd shows up.

With real numberslesson 03, and the rule in the Day 10 sizing memo

  • Busiest moment: 120 people at once.
  • One copy keeps its promise up to 8 people (the practice server’s limit).
  • 120 ÷ 8 = 15 copies.
  • The lesson’s answer, 18, adds 3 spare copies: 20% extra (15 x 1.2 = 18). The Day 10 sizing memo adds the same 20%.

Words to know

Replica
One complete running copy of the server and model. Example: 18 replicas for 120 people.
SLO (service level objective)
The speed promise. Example: 95 of 100 requests get their first token within 500 ms (half a second).
Peak traffic
The most people using the service at the same moment. Example: 120.
Headroom
Spare capacity kept for surprises. Example: plan 18 where 15 are needed.
Go deeper: the engineer version

The kit's question

Peak traffic is 120 concurrent users and one replica handles 8 within SLO. How many replicas?

The kit's answer

15, plus headroom. Say 18.

More detail: Replicas = peak concurrency ÷ per-replica capacity at the SLO, rounded up, plus headroom for traffic spikes, a replica restarting and forecast error. Lesson 15’s build_memo.py computes exactly this, math.ceil(a.peak_concurrency / per_replica * 1.2): 120 ÷ 8 x 1.2 = 18. Measure the per-replica capacity with the customer’s real prompt and reply lengths, because longer contexts move the knee.

How did you do?

QuestionStep 1 · Concurrency sweep

On the sweep chart, the knee is the point where all the server’s seats are taken and new people start waiting in line. What moves that point?

Show answerHide answer

In plain words

The knee moves to more people when more fit at once: more seats, more memory set aside for the notes (the KV cache), a smaller model, or shorter conversations. It moves to fewer people when one long prompt hogs the chip while it is read.

Picture it

A restaurant is full at a certain number of guests. More tables, more floor space or shorter meals seat more people; one huge order hogging the kitchen slows every other table.

With real numberslesson 03’s practice server and llama.cpp settings, and lessons 04 and 05 (Day 4)

  • Practice server: 8 seats, so the knee comes right after 8 people.
  • At 8 people, 95 of 100 got their first token within 41.4 ms (thousandths of a second); at 16, within 3,357 ms (over 3 seconds): about 80 times longer.
  • Seats share memory: llama.cpp splits 16,384 tokens of conversation room across 8 seats, 2,048 tokens (about 1,500 words) each. More seats means less room for each conversation.
  • Memory for notes sets the seats. Day 4’s calculator: 20 GB of spare memory holds 4 conversations of 32,768 tokens (about 25,000 words) with notes at 16 bits per number, and 8 at 8 bits.
  • A smaller model frees memory: Llama 3.1 8B at 4 bits is 4.9 GB against 8.5 GB at 8 bits, 3.6 GB more for notes (Day 4).

Words to know

Knee
The point where seats run out and waiting time shoots up. Example: after 8 people on the practice server.
Slot
One seat on the server, a conversation it can serve at the same time. Example: the practice server has 8.
Notes (KV cache)
The model’s memory of each conversation so far, kept in the same memory as the model. Example: 4 conversations of 32,768 tokens fit in 20 GB.
Prefill
The reading phase; the model reads the whole prompt at once before writing. Example: a 100,000-token prompt takes a long prefill.
Go deeper: the engineer version

The kit's question

What changes the knee?

The kit's answer

More slots or memory for KV cache, a smaller or quantized model, shorter contexts, and prefill/decode interference.

More detail: In llama.cpp the -c context is split across the --parallel slots (the comment in serve_llamacpp.sh says so), so slots and context length trade off directly. A quantized model frees memory for KV cache and also makes each decode step faster. Long prefills stall the decode steps of everyone batched with them; chunked prefill (Batching & Scheduling video) is the usual fix. The 41.4 and 3,357 ms figures are TTFT p95.

How did you do?

Under 20 seconds each. Record yourself once and listen back.

Step 1 · Concurrency sweep

Question The vendor’s benchmark says this GPU is fast. How many of them do we need for launch?

One clear answer

Single-stream numbers sell hardware; concurrency curves size deployments. I sweep, find the highest concurrency that holds p95 TTFT under the SLO, and size replicas from peak traffic divided by that.

What this means
  • “Single-stream numbers sell hardware”: a speed measured with one person at a time makes a chip look good in a brochure, but says nothing about a crowd. Example: the practice server’s 54.9 tokens per second for one person.
  • “concurrency curves size deployments”: how many servers to buy comes from a chart of output and waiting time at 1, 2, 4, 8 and 16 people at once.
  • “I sweep”: I run the same test at each crowd size in turn, with the customer’s own prompt and reply lengths.
  • “find the highest concurrency that holds p95 TTFT under the SLO”: I find the biggest crowd at which 95 of 100 people still get their first token within the promised time. Practice server: 8 people, 41.4 ms against a 500 ms promise; at 16 it is 3,357 ms.
  • “and size replicas from peak traffic divided by that”: copies of the server needed = the most people at once ÷ people per copy. 120 ÷ 8 = 15, plus 20% spare = 18.

2 worked examples from today’s videos. Work each on paper, then reveal.

From the video: Benchmarking Properly 2:22

Whiteboard 1 of 2

You tested a server at 1, 4, 8, 16 and 32 people at once. Each request sends 1,024 tokens (about 770 words) and gets 256 back (about 190 words), with 200 requests per crowd size. The promise: 95 of 100 people see the first token within 1 second. How many people can one server take?

Given, in plain words

The results, one row per crowd size: total output in tokens per second, then the p95 wait for the first token. 1 person: 24, 210 ms. 4 people: 78, 260 ms. 8: 135, 380 ms. 16: 228, 720 ms. 32: 264, 1,850 ms.

Reveal the answerHide the answer

Answer · in plain words

16 people per server. At 16, 95 of 100 still see the first token within 720 ms, under the 1-second promise; at 32 it is 1,850 ms, almost twice the promise.

Picture it

A kitchen sends out more plates per hour as tables fill, but the wait for the first course grows. Its capacity is the most tables at which nearly every table is still served on time. It is not the most plates the kitchen can ever cook.

Careful

The narration says the p95 crosses 1 second ‘somewhere between eight and sixteen’. The table says between 16 and 32 (720 ms at 16, 1,850 ms at 32), and the on-screen answer agrees: size at 16.

Worked answer, step by step

  1. The promise: 95 of 100 people see the first token within 1,000 ms (1 second).
  2. Read down the p95 column: 210, 260, 380 and 720 ms all keep the promise.
  3. At 32 people it is 1,850 ms: 1.85 times the promise, so 32 is too many.
  4. The last tested size that keeps it is 16, so size one server at 16 people. The exact crowd where the wait passes 1 second lies somewhere between 16 and 32; it was not tested.
  5. Total output there: 228 tokens per second, or 228 ÷ 16 = about 14 per person, against 24 for one person alone.
  6. The curve is flattening: output rose 69% from 8 to 16 people (135 to 228), but only 16% from 16 to 32 (228 to 264).
Go deeper: the engineer version

The kit's question

Worked example · the table you produce · 1,024 in / 256 out · 200 prompts per point

The kit's answer

Size at concurrency 16: 228 tok/s inside a 1 s p95 budget

More detail: Read the table from the bottom: throughput flattens past 16, and p95 TTFT crosses the 1 s budget between 16 and 32, so 16 is the highest measured concurrency under the SLO. To locate the crossing more precisely, add levels between 16 and 32 with --levels. The same video shows why the percentile matters: from c = 8 to c = 32 the mean TTFT only went from 290 to 410 ms, while the p95 went from 380 to 1,850 ms, so a mean-only report would call both fine. Lesson 03’s sweep.py uses the same 1,000 ms budget by default.

Words to know

p95
The value 95 out of 100 requests stay under. Example: 720 ms at 16 people.
TTFT (time to first token)
How long until the first token of the reply appears. Example: 210 ms (p95) for one person.
Throughput
Total output across everyone at once. Example: 228 tokens per second at 16 people.
Traffic profile
The shape of the load a test uses: tokens in, tokens out, requests per second, and the speed promise. Example: 1,024 in, 256 out, first token within 1 s at p95.

From the video: GPU Bandwidth 3:10

Whiteboard 2 of 2

A customer needs 100 tokens per second (about 75 words a second) for each person, on a model with 70 billion numbers (a 70B model). What hardware does that require?

Given, in plain words

The model is stored at 8 bits (FP8): 1 byte per number, so 70 GB. Writing each token means fetching the whole model from memory once. Memory speeds of three NVIDIA data-center GPUs: H100 3,350 GB per second, H200 4,800, B200 8,000.

Reveal the answerHide the answer

Answer · in plain words

It needs memory that delivers about 7,000 GB a second (100 x 70 GB). At 8 bits only a B200 (8,000) is above that, and only on paper: real speeds land lower. A B200 with the model at 4 bits (limit 229) clears 100 with room; an H100 at 4 bits (96) does not.

Picture it

A tank drains through one pipe. Each bucket is one token, and filling it takes the whole model: 70 litres. To fill 100 buckets every second, the pipe must carry 7,000 litres a second. A pipe that carries 3,350 fills only 48 buckets a second, however fast you work at the other end.

Careful

The video calls B200 and 4-bit on an H100 ‘two ways to say yes’. By its own numbers, 4-bit on an H100 stays just under 100 (about 96), and both figures are upper limits, not measured speeds.

Worked answer, step by step

  1. Model size: 70 billion numbers x 1 byte = 70 GB, fetched once per token.
  2. Memory speed needed: 100 tokens per second x 70 GB = 7,000 GB per second.
  3. H100: 3,350 ÷ 70 = 48 tokens per second, under half the need.
  4. H200: 4,800 ÷ 70 = 69. Still short.
  5. B200: 8,000 ÷ 70 = 114. Clears 100 on paper, by 14%.
  6. Or shrink the model to 4 bits, half a byte per number: 70 x 0.5 = 35 GB. An H100 then reaches 3,350 ÷ 35 = 96: close, but still under 100.
  7. These are upper limits. The same video says real models like this one reach 60 to 85% of them. For a B200 at 8 bits: 114 x 0.6 = 68 and 114 x 0.85 = 97. So expect 68 to 97 tokens per second: under 100.
  8. A B200 with the model at 4 bits: 8,000 ÷ 35 = 229. At 60 to 85% that is 229 x 0.6 = 137 to 229 x 0.85 = 195. That clears 100 with room. Measure before you promise 100.
Go deeper: the engineer version

The kit's question

A customer says they need one hundred tokens per second per user on a seventy billion parameter model. Turn that into a hardware requirement.

The kit's answer

B200 — or quantise to 4-bit and an H100 reaches ~96 tok/s

More detail: Single-stream decode ceiling = memory bandwidth ÷ bytes read per token; the kit’s felab/hardware.py uses the same formula and the same bandwidths (H100 SXM 3,350, H200 4,800, B200 8,000 GB/s). The requirement is per user, so batching does not help here: it raises total throughput while each user’s speed falls a little (lesson 03’s practice server: 54.9 to 38.9 tokens per second). The narration’s own verdict for the H100 is ‘not on an H one hundred, not at that precision’. Its 4-bit figure assumes 35 GB; the same video’s desktop example uses 38 GB for a 70B at 4-bit (real 4-bit files carry shared scales), which gives 3,350 ÷ 38 = 88 tokens per second. All figures are for one GPU; splitting the model across GPUs (the Fitting Big Models video, Day 9) changes the arithmetic.

Words to know

Bandwidth (memory bandwidth)
How many GB per second memory delivers to the chip. Example: H100 3,350 GB/s.
Ceiling (speed limit)
The fastest a model can write for one person: memory speed ÷ model size. Example: 3,350 ÷ 70 = 48 tokens/s.
FP8
An 8-bit number format, 1 byte per number. Example: a 70B model is 70 GB in FP8.
H100, H200, B200
Three generations of NVIDIA’s data-center GPUs, older to newer. Example: memory speeds of 3,350, 4,800 and 8,000 GB/s.

How you will use this

A customer deciding how many servers to run needs this crowd-size result, with your test conditions written down. On Day 10 your sizing memo uses it to estimate server count and adds room for traffic spikes.