Batching & Scheduling
file 05 · 2:58 · subtitles burned in
Chapters
In this video Three ways a server keeps its expensive chip busy: continuous batching, chunked prefill and speculative decoding.
What it explains: continuous batching, chunked prefill, speculation
3 key points
Continuous batching gives a free seat to the next request the moment a reply finishes, instead of waiting for the whole group.
The video’s 12 requests on 4 seats: the chip is busy 87% of the time instead of 61%. All the work is done at step 17 instead of 24 (each step writes one token for every busy seat).
Chunked prefill reads one huge prompt in pieces, so it cannot freeze everyone else’s replies.
The video’s example is a 100,000-token prompt (about 75,000 words), read in chunks between other people’s writing steps. That prompt’s first token comes a little later; everyone else’s reply keeps appearing word by word.
Speculative decoding cuts waiting: a small model drafts a few tokens and the big model checks them all in one pass, so the output does not change.
Day 9 comes back to it. The video’s real case stacks three fixes (reusing a repeated prompt, chunked prefill, speculation) on the same 2 GPUs (the chips that run the model). The p95 wait for the first token fell from 2.8 s to 0.40 s, 7 times faster. The p95 gap between tokens fell from 48 to 26 ms, the part speculation speeds up.
Used in the course
- Day 3 · Step 1 · Concurrency sweep
- Day 9 · Step 2 · Speculative decoding (rewatch, 1:01 to 1:23)
The title card in the video says “Video 5 of 18”: that is the file order. The course plays the videos in the order of its days.