QuestionStep 1 · Concurrency sweep
When 8 people used the practice server at once instead of 1, its total output rose about 5 times. Yet each person’s reply only slowed from 55 to 39 tokens (word pieces) per second. Why does sharing cost so little?
Show answerHide answer
In plain words
To write each token, the chip must fetch the whole model from memory, and that fetch is the slow part. One fetch serves all 8 people at once, so each extra person adds only a little.
Picture it
A bus takes about the same time to drive its route with 1 rider or with 8. The drive is the slow part; letting a few more people on adds a little time at each stop. This stops working once the seats run out: then extra riders wait for the next bus, and that point is the knee.
With real numberslesson 03’s practice server: it behaves like a real one, but its numbers are made up for teaching
- 1 person alone: 54.1 tokens per second in total; that person sees about 55 (54.9) once the reply is flowing.
- 8 people: 275.1 in total, about 5 times the 54.1 (275.1 ÷ 54.1 = 5.1).
- Each person: 54.9 falls to 38.9 tokens per second, 29% slower.
- So each person keeps 71% of their solo speed while the server does 5 times the work.
Words to know
- Decode
- The writing phase; the model writes one token at a time and fetches the whole model for each. Example: 54.9 tokens per second for one person on the practice server.
- Memory-bound
- Slowed by how fast memory delivers data, not by how fast the chip calculates. Example: decode.
- Batching
- Serving several people with one pass over the model. Example: 8 people share each fetch.
- Tokens per second (tok/s)
- How many tokens are written each second. Example: 38.9 per person with 8 people.
Go deeper: the engineer version
The kit's question
Throughput went up 5×, but per-user speed only dropped from 55 to 39 tok/s. Why so cheap?
The kit's answer
Decode is memory-bound, so sharing one weight read across 8 users costs little extra.
More detail: Each decode step streams all the weights once, whatever the batch size. Each extra user still adds its own KV-cache reads and a little compute, which is why per-user speed falls somewhat (54.9 to 38.9). The practice server builds this in on purpose: one step takes 18 ms alone and grows 6% per extra user, so 18 x 1.42 = 25.6 ms with 8. That is 1,000 ÷ 18 = 55.6 tokens per second alone and 1,000 ÷ 25.56 = 39.1 with 8 (felab/mock_server.py, _itl). In roofline terms, one user’s decode does about 1 to 2 calculations per byte fetched, while an H100 needs about 300 to keep its math busy; batching raises that ratio. The 1-user total (54.1) sits a little under the per-user figure (54.9) because the per-user figure is the median speed once text starts flowing, while the total divides all tokens by wall-clock time, which includes the wait before the first token.
How did you do?