Speculation & Scale-Out
file 13 · 2:41 · subtitles burned in
Chapters
In this video How a small draft model speeds up a big one, why the acceptance rate decides whether it pays, and what linking a second machine does and does not buy.
What it explains: speculative decoding, acceptance rate, two boxes
3 key points
A small draft model guesses a few tokens; the big model checks them all in one pass.
Correct guesses are free tokens, and the big model still decides every token, so the output does not change. With 75% of guesses kept and 4 guesses a round, it writes about 3 tokens per big pass instead of 1.
The acceptance rate, the share of guesses kept, decides whether it pays.
At 80% kept with 4 guesses, the video estimates about 2 times faster once the draft model’s own time is paid. At 40% it calls the gain barely worth it, and a mismatched draft model can make writing slower.
A second machine adds memory, not speed.
Two linked 128 GB desktops (256 GB) can run a 235-billion-parameter model at 4 bits, but it writes about 12 tokens a second (11.7). Good for testing that model, not for serving many users.
Used in the course
The title card in the video says “Video 13 of 18”: that is the file order. The course plays the videos in the order of its days.