About 10 minutes, free, and it works on your phone. Lock in today before you move on.
Done when
Your notes hold two results. From step 1: how much of gpt-oss:20b is read per token, in GB and as a share. From step 2: writing speeds with and without the draft model on structured and prose prompts, with acceptance rates, and the Fireworks deployment deleted.
Fields you filled in on the step pages are already here. Nothing leaves this browser.
Step 1 · MoE vs dense
How fast a model writes when it reads all of itself for every token.
Example: 84.0 in the lesson’s sample (a Mac like an M4 Max).
tok/s
How fast a model writes when it reads only the experts it picks.
Example: 118.0 in the sample: faster than the dense model, from a file 2.8 times bigger.
tok/s
Your stopwatch estimate of how much of the model is read for each token.
Example: 3.5 GB and 25.1% in the sample, against 17.1% published: within 2 times, a good result.
GB / %
Step 2 · Speculative decoding
How much speculation speeds up easy-to-guess output on your Mac.
Example: 62.0 / 118.0 in the lesson’s sample: 1.9 times.
tok/s
How much speculation helps when the next words are hard to guess.
Example: 61.5 / 70.0 in the sample: 1.14 times.
tok/s
The number that decides whether speculation pays.
Example: 0.780 / 0.410 in the sample.
share of guesses kept
The same test on a Fireworks deployment with a 1B draft model guessing 4 tokens a round.
Example: The lesson has no sample for this run. Expect the structured row faster than the prose row.
tok/s
Step 3 · The Spark track
How many full-length conversations fit in vLLM’s memory for notes at once.
Example: The lesson prints no sample. The kit’s numbers give at most (0.8 × 128 − 8) ÷ 1.07 = about 88; expect fewer after vLLM’s own working memory. The Serving With vLLM video’s log, from another machine, shows about 16.
x
How many people one Spark serves while 95 of 100 still get their first token within a second.
Example: The lesson prints no sample. For scale, the lab book’s SGLang figure for this model: 368 tokens a second in total at 32 people.
users / tok/s
How much of the prompt SGLang reused instead of reading again.
Example: cached_tokens close to prompt_tokens on the same line: nearly the whole opening (the script estimates about 3,028 tokens) was reused.
A customer asks why their 120B mixture-of-experts model (120 billion parameters, split among many specialists) writes faster than their 70B dense model (70 billion, all used for every token). What do you tell them?
Show answerHide answer
In plain words
Writing speed depends on how much the model reads for each token, not on its total size. The MoE reads about 5 billion parameters per token; the dense model reads all 70 billion.
Picture it
Looking something up in a 1,200-page encyclopedia is quick when the index sends you to 50 pages. Reading a 700-page manual cover to cover is slow, although it is the smaller book. You still need shelf space for all 1,200 pages.
With real numberslesson 12’s published sizes and the lab book’s DGX Spark measurements
gpt-oss 120B: 117 billion parameters in total, 5.1 billion used per token (the published figures in lesson 12’s script).
A 70B dense model uses all 70 billion for every token: 70 ÷ 5.1 = about 14 times as much to read.
Writing speed is capped at memory speed ÷ bytes read per token, so fewer bytes means faster writing.
Measured on a DGX Spark (lab book): gpt-oss 120B writes 55.4 tokens a second; Llama 3.1 70B stored at 8 bits writes 2.7.
55.4 ÷ 2.7 = about 20 times faster, although the MoE holds more parameters (117 billion against 70).
The measured gap is bigger than 14 partly because the lab book’s two runs used different server software (llama.cpp and SGLang) and number formats. Read it as roughly 14 to 20 times.
Words to know
Parameters
The model’s learned numbers. Example: 70B means 70 billion.
Active parameters
The parameters read to write one token. Example: 5.1 billion for gpt-oss 120B.
Decode
The writing phase, one token at a time; its speed is set by the bytes read per token.
DGX Spark
NVIDIA’s desktop AI computer: 128 GB of memory moving 273 GB a second.
Go deeper: the engineer version
The kit's question
A customer asks why their 120B MoE is faster than their 70B dense. What do you say?
The kit's answer
About 5B active parameters against 70B read per token, so decode is bandwidth ÷ active bytes.
More detail: Decode is memory-bound: tok/s is about bandwidth × efficiency ÷ bytes read per token. For an MoE those bytes are the chosen experts plus the parts every token uses (attention, embeddings, router) and the KV cache, each at its own precision. The lab book inverts the Spark figure directly: 273 ÷ 55 ≈ 5 GB per token out of a 63 GB file, under 10% of the model. The two figures also differ in format: the 70B is FP8, 1 byte per parameter; gpt-oss 120B is MXFP4, a 4-bit format. Memory is still sized on the total: all 59 GiB (63.4 GB) must fit.
Why can a mixture-of-experts model be harder, not easier, to serve to many people at once on GPUs?
Show answerHide answer
In plain words
Different people’s tokens need different experts, so they cannot all share one read of the model, as they do with a dense model. Spread over several GPUs, tokens must also travel between chips to reach their experts.
Picture it
A bus is efficient because everyone rides the same route. If every passenger needs a different part of town, you end up running many half-empty minibuses. And if the specialists work in different buildings, patients spend time walking between them.
With real numbersDay 3’s practice server and the field guide’s expert example
Dense model on Day 3’s practice server: 8 people got about 5 times one person’s total output, because they share each read of the model.
The field guide’s MoE example: 4 experts in a layer, 2 used per token.
One person: 2 of the 4 experts are read, half of that layer’s experts.
Two people whose tokens pick different pairs: all 4 experts are read, and each read serves only one person.
With the 4 experts on 4 GPUs, as in the field guide, each token is sent to the GPUs that hold its 2 experts, and the results are sent back.
Words to know
Expert
One specialist block in an MoE; only the few the router picks are read for a token.
Batching
Serving several people with one pass over the model.
Expert parallelism
Placing different experts on different GPUs, so tokens travel to the GPUs that hold their experts.
All-to-all
A network step where every GPU sends data to every other GPU. Example: moving tokens to their experts.
Go deeper: the engineer version
The kit's question
Why can MoE be harder to serve at high concurrency on GPUs?
The kit's answer
Different tokens hit different experts, so batches fragment. Expert parallelism across GPUs and all-to-all traffic also add complexity.
More detail: Batched decode shares one weight read across the batch. In an MoE layer each token’s router picks its own experts, so as the batch grows the set of experts read approaches all of them, while each expert’s share of the batch shrinks: more bytes per step and smaller, less efficient matrix multiplies. Expert parallelism puts experts on different GPUs and needs an all-to-all exchange of token activations in every MoE layer, out to the experts and back. That adds network traffic, and the busiest expert sets the pace.
A model needs 160 GB of memory, and each H100 GPU has 80 GB, so it must be split over 2 GPUs. Do you split every layer (one of the model’s stacked processing stages) across both (TP=2, tensor parallel) or give each GPU half of the layers (PP=2, pipeline parallel)?
Show answerHide answer
In plain words
Split every layer (TP=2) when both GPUs sit in one server with a very fast link between them: each token then finishes sooner. Use the half-the-layers split (PP=2) only when the link between the GPUs is slow.
Picture it
Two cooks can split every dish: each chops half the vegetables, and they combine their halves before every next step. That constant passing only works side by side. Or one cook makes starters and the other mains: one hand-off per order, fine in separate kitchens. But each order waits for both, and a cook stands idle when orders are few.
With real numbersthe check’s numbers, the field guide and lesson 14’s two-box notes
160 GB ÷ 80 GB per GPU = 2 GPUs at the very least, before any room for notes.
TP=2: each GPU holds half of every layer, and the two sync inside every layer, many times per token.
Inside one server, NVLink joins H100s at about 900 GB a second (field guide): fast enough for constant syncing.
Two DGX Sparks are joined by a 200-gigabit cable: 200 ÷ 8 = 25 GB a second. NVLink’s 900 adds both directions together; even halved to 450, it is 18 times the cable (450 ÷ 25).
Over a link like that, PP=2 often does better: one hand-off per token between the two halves, instead of syncs in every layer.
Words to know
Tensor parallelism (TP)
Splitting every layer across GPUs so they work on each token together; they sync inside every layer.
Pipeline parallelism (PP)
Giving each GPU a block of layers; each token passes from one GPU to the next.
NVLink
NVIDIA’s very fast direct link between GPUs inside one server. Example: about 900 GB/s on H100.
Pipeline bubble
Time a GPU sits idle in a pipeline, waiting for work from the stage before it.
Go deeper: the engineer version
The kit's question
The model needs 160 GB and each H100 has 80 GB. TP=2 or PP=2?
The kit's answer
TP=2 inside one NVLink node for latency. PP only if the link between GPUs is slow.
More detail: TP shards each weight matrix and all-reduces activations inside every layer, so it cuts per-token latency but is bound by interconnect latency and bandwidth: keep it inside one NVLink domain. PP places consecutive layer blocks on each GPU and passes activations once per stage boundary. It tolerates slower links but adds per-hop latency, and bubbles unless enough requests keep every stage busy. Across boxes the kit’s advice flips to trying PP first, because TP wants NVLink-class bandwidth (TWO_BOX.md). Note that 2 × 80 GB leaves no headroom for a 160 GB model: the Fitting Big Models video adds about 10% for activations and the runtime.
Why must the small draft model share the big model’s tokenizer? (The tokenizer is the part that cuts text into tokens and gives each token a number.)
Show answerHide answer
In plain words
The big model checks the guesses as token numbers, not as text. If the two models cut and number text differently, the same number means different words, and the check means nothing.
Picture it
Two warehouse clerks check an order by product code. That only works if they use the same catalogue: in one, code 4521 is a lamp; in the other, a sofa.
With real numbersthe lesson’s two models, lesson 13
Big model: Llama 3.1 8B at 4 bits. Draft: Llama 3.2 1B at 4 bits, from the same family, so it uses the same token numbers.
The draft sends up to 8 guesses a round, as token numbers (--draft-max 8).
The big model compares each guessed number with the number it would have picked, and keeps the matching run.
On Fireworks the lesson uses the same pair, Llama 3.1 8B checking a Llama 3.2 1B draft, with 4 guesses a round.
Words to know
Tokenizer
The part of a model that cuts text into tokens and numbers them. Example: Llama 3.1 8B and Llama 3.2 1B share one.
Token ID
The number a tokenizer gives each token; models compare these numbers, not the text.
Vocabulary
The full list of tokens a tokenizer knows, each with its own number.
Go deeper: the engineer version
The kit's question
Why must the draft share the tokenizer?
The kit's answer
The target verifies token ids, not text.
More detail: Verification compares the draft’s proposed token ids with the target’s own next-token choices position by position (with sampling, it accepts against the target’s probabilities), so both models must share one vocabulary and tokenization. A draft from another family proposes ids that mean different strings to the target. N-gram speculation and predicted outputs avoid the issue, because their drafts come from the target’s own tokens.
Why does speculative decoding help less when the server is busy with many people at once (high concurrency)?
Show answerHide answer
In plain words
With one person, the chip spends most of each step waiting for memory, so it has spare math power to check guesses almost for free. With many people, that spare power already goes to the other people’s tokens, so checking guesses now costs real time.
Picture it
A delivery van on a single drop has lots of empty space, so carrying extra parcels on the off chance costs nothing. Once it is full of other customers’ parcels, every extra parcel pushes out a real one.
With real numbersthe lab book’s DGX Spark measurements and lesson 13’s formula
One person: writing reads the whole model for each token, and the math units mostly wait (Day 3: memory-bound).
Llama 3.1 8B at 8 bits on a DGX Spark (lab book): 20.5 tokens a second for 1 person, 368 in total for 32.
368 ÷ 20.5 = 18 times the output from the same memory reads: the spare math is now in use.
Checking 4 guesses means the math for 5 tokens per person instead of 1, and every rejected guess is wasted math.
So speculation pays most for one or a few people who want fast replies.
Words to know
Latency-bound
A workload where what matters is how fast each reply arrives, typical of one or a few people at a time.
Memory-bound
Slowed by how fast memory delivers data, not by math. Example: writing for one person.
Batch size
How many requests share one pass over the model. Example: 1, then 32, on the Spark.
Compute
The chip’s math power. Example: what checking guesses uses.
Go deeper: the engineer version
The kit's question
Why does speculation help less at high concurrency?
The kit's answer
With a full batch the GPU is already busy, so spare compute for verification shrinks. It shines on latency-bound, low-batch workloads.
More detail: At batch 1, decode does about two FLOPs per parameter (1 to 2 per byte fetched at 16 or 8 bits), against an H100 ridge point of about 300 (989 teraflops ÷ 3.35 TB/s, the GPU Bandwidth video), so compute sits idle and verifying k drafted tokens in the same weight read is almost free. As the batch grows, arithmetic intensity rises toward the compute roof; verification adds k + 1 positions per sequence, and rejected positions become wasted compute that displaces other users’ tokens. Speculation shines on latency-bound, low-batch workloads; at high concurrency its gain shrinks and can turn negative.
For editing tasks, where most of the answer repeats the input, what can you use instead of a draft model?
Show answerHide answer
In plain words
Use text you already have as the guess. N-gram speculation finds the last few words it wrote earlier in the prompt and guesses what came next there. Fireworks’ predicted outputs let you send the expected answer, such as the original document, as the draft.
Picture it
Proofreading: instead of an intern drafting from scratch, you hand the editor last week’s version. Most of it passes at a glance, and only the changed lines need real work.
With real numbersfireworks_spec.sh, the lab book and lesson 13’s formula
The script’s comment names the option: --ngram-speculation-length=3, no draft model; it reuses n-grams (short runs of tokens) from the prompt.
The lab book lists four ways on Fireworks: the default draft model, your own draft model, n-gram, or predicted outputs.
Suppose an edit keeps 9 of 10 guesses (α = 0.9, the highest rate in lesson 13’s table; measure it on real edits). With 4 guesses: 1 + 0.9 + 0.81 + 0.73 + 0.66 = 4.1 tokens per big pass.
With no draft model to run, guessing costs almost nothing, so the ideal speed-up approaches those 4.1 times. Real gains come in lower.
Words to know
N-gram
A run of n tokens in a row. Example: a 3-gram is 3 tokens.
N-gram speculation
Guessing the next tokens by copying what followed the same words earlier in the prompt; no draft model.
Predicted outputs
A Fireworks option: you send the expected answer as the draft for the model to check.
Go deeper: the engineer version
The kit's question
What is the zero-model alternative for editing tasks?
The kit's answer
N-gram speculation or Fireworks predicted outputs: the draft comes from the prompt itself.
More detail: N-gram (prompt-lookup) speculation matches the last few generated tokens against earlier text in the context and proposes the tokens that followed; predicted outputs let the client send the expected completion, for example the original file for a code edit. Both need no draft model and suit edits, rewrites and extraction, where the output copies long spans of the input. Measure acceptance on the real traffic, as with any drafter.
vLLM’s startup log says its memory for notes holds 18 full conversations of 8K tokens (8,192, about 6,000 words) at once: maximum concurrency 18x. How do you double that?
Show answerHide answer
In plain words
Make each conversation’s notes smaller, or give the notes more memory. Store them in 8 bits instead of 16, lower the cap on conversation length, or use a smaller or more compressed model.
Picture it
A wardrobe holds 18 coats. Vacuum-pack each coat to half its size, or allow jackets instead of long coats, and 36 fit. Or take out the big suitcase that was using some of the space.
With real numbersDay 4’s calculator for Llama 3.1 8B, and the vLLM settings from the Serving With vLLM video
Notes for one 8,192-token conversation at 16 bits: 1.07 GB (Day 4).
Say the server has 20 GB for notes, as in Day 4’s example: 20 ÷ 1.07 = 18.6, so vLLM reports about 18.6x (18 whole conversations): the check’s 18x.
Or cap conversations at 4,096 tokens (--max-model-len 4096): also 0.54 GB each, also 37, if real requests fit in 4,096.
Or a smaller or more compressed model: its weights shrink, and vLLM gives the freed memory to notes.
Words to know
Maximum concurrency
vLLM’s count of how many full-length conversations fit in its memory for notes.
KV cache dtype
The number format the notes are stored in; 8-bit instead of 16-bit halves their memory.
Max model length (--max-model-len)
vLLM’s cap on tokens per conversation, which caps that conversation’s notes.
FP8
An 8-bit number format, 1 byte per number.
Go deeper: the engineer version
The kit's question
The vLLM log says max concurrency is 18× at 8k context. How do you double it?
The kit's answer
Use an FP8 KV cache, a lower --max-model-len or a smaller or quantized model.
More detail: vLLM’s figure is KV-cache tokens ÷ max_model_len. The Serving With vLLM video’s sample log: 8,234 blocks × 16 tokens ≈ 131,000 tokens, about 16 requests at 8K. FP8 KV halves bytes per token, so it doubles the figure exactly (the video: “roughly doubles concurrency”). Halving --max-model-len also doubles it, but is only safe if the real p95 fits (Day 4). A smaller or quantized model frees memory inside --gpu-memory-utilization for the KV pool, which raises the figure by however much it frees.
When would you pick SGLang over vLLM, the two main open-source servers for NVIDIA GPUs?
Show answerHide answer
In plain words
When many requests share long openings or branch from one conversation, as agents do, or when most output must follow a strict format such as JSON (a common format for structured data). Either way, test both on the customer’s real traffic before choosing.
Picture it
Two delivery firms: one is best at many different parcels, the other at many parcels to the same street. Which is cheaper depends on your parcels, so you give each a trial week.
With real numbersDay 5’s prefix test and this step’s SGLang run
Day 5’s shared opening: about 3,028 tokens by the script’s estimate, sent with each of 12 questions.
Reusing it made the first word 20 times sooner in the lesson’s sample: 740 ms down to 36 ms (thousandths of a second).
SGLang keeps saved notes like a family tree (RadixAttention): one shared opening at the root, a branch for each follow-up, so every branch reuses the common start.
This step’s proof: SGLang’s usage: line shows cached_tokens close to prompt_tokens on the same line: nearly the whole opening was reused.
Words to know
SGLang
An open-source server for NVIDIA GPUs, strong at reusing shared prompt openings.
RadixAttention
SGLang’s store of saved notes, kept as a tree of shared openings.
Agent
A program that calls a model many times, using tools, to finish a task.
Structured generation
Making every reply follow a fixed format, such as JSON.
Go deeper: the engineer version
The kit's question
When would you pick SGLang over vLLM?
The kit's answer
For heavy shared-prefix or branching agent workloads and structured-generation-heavy pipelines. Benchmark both on the real traffic.
More detail: SGLang’s RadixAttention keeps the KV cache in a radix tree keyed by token prefixes and schedules requests to maximise prefix hits, which suits multi-turn and branching agent traffic; its constrained decoding is strong for JSON-heavy pipelines. vLLM’s automatic prefix caching also reuses identical block prefixes, so the gap depends on the traffic: benchmark both on the customer’s prompts, with their cache hit rates.
Why run vLLM and SGLang on the Spark only inside containers, never installed straight onto the machine with pip (Python’s package installer)?
Show answerHide answer
In plain words
The GPU’s driver, NVIDIA’s CUDA software and the server must be exactly matching versions. A container ships a set known to work together; installing with pip on the machine mixes versions and can break them.
Picture it
A meal kit comes with exactly the ingredients the recipe was tested with. Shop for each item yourself and you end up with a different flour, and the cake fails.
With real numbersthe Serving With vLLM video (from Day 2), the lab book’s DGX Spark page and lesson 14’s scripts
The Serving With vLLM video: when versions clash, the usual workaround switches off a GPU speed-up (CUDA graphs) and loses 20 to 30% of total output.
The lab book: the pip route for vLLM on the Spark’s arm64 chip has broken for people (mismatched versions of PyTorch, the Python library vLLM is built on, and CUDA).
NVIDIA ships about 65 DGX Spark setup guides (playbooks) that fix container versions known to work.
1_vllm.sh pins one image tag, nvcr.io/nvidia/vllm:25.09-py3; swap it only for another tag built for the Spark (GB10/arm64).
Words to know
Container
A sealed package holding a program and the exact software versions it needs, so it runs the same anywhere.
CUDA
NVIDIA’s software layer that lets programs run on its GPUs.
Driver
The software that lets the operating system use a piece of hardware such as a GPU.
pip
Python’s tool for installing packages. Example: pip install vllm, which the lesson forbids on the Spark.
Go deeper: the engineer version
The kit's question
Why containers only?
The kit's answer
CUDA, driver and engine versions have to match exactly. Containers pin them, and host pip installs break them.
More detail: On GB10 (arm64 with a Blackwell GPU) the host OS, driver and CUDA move together on NVIDIA’s schedule. Prebuilt wheels can expect a different CUDA than the host has, and the usual workaround, eager mode without CUDA graphs (--enforce-eager), costs 20 to 30%. NGC and vendor containers pin CUDA, PyTorch and the engine together, so a customer can reproduce your numbers with the same tag.
Under 20 seconds each. Record yourself once and listen back.
Step 1 · MoE vs dense
QuestionWhy is the MoE model faster than the smaller dense one?
One clear answer
Memory is set by total parameters; decode speed by active parameters. I can estimate a model’s active footprint from its decode speed on known hardware.
What this means
“Memory is set by total parameters”: The machine must hold every expert, the whole model. gpt-oss:20b needs its full 13.8 GB file in memory, though it reads only part of it per token.
“decode speed by active parameters”: How fast it writes depends only on what it reads per token: 3.6 of its 21 billion parameters. In the sample it writes 118 tokens a second, against 84 for the 4.9 GB dense 8B.
“I can estimate a model’s active footprint from its decode speed on known hardware”: If I know the machine’s memory speed, I time the writing and work backwards: 546 × 0.75 ÷ 118 = 3.5 GB read per token, about a quarter of the file.
Step 2 · Speculative decoding
QuestionWill speculative decoding help us?
One clear answer
Acceptance rate decides whether speculation pays. On structured traffic it can double decode speed; a generic drafter on unusual traffic can make things slower. I measure α on their prompts before recommending it.
What this means
“Acceptance rate decides whether speculation pays”: The share of the small model’s guesses the big model keeps is the number that matters. At 0.8 with 4 guesses it is 2.4 times faster; at 0.3 with 8 guesses, 0.79 times: slower.
“On structured traffic it can double decode speed”: Decode speed is writing speed. Predictable output such as JSON data, code or extraction is easy to guess. In the lesson’s sample it went from 62 to 118 tokens a second, 1.9 times.
“a generic drafter on unusual traffic can make things slower”: A draft model that has not seen this kind of text guesses badly, and the wasted guesses cost more than they save.
“I measure α on their prompts before recommending it”: I run the customer’s own prompts with speculation on and read the acceptance rate (α) before I promise any speed-up.
Step 3 · The Spark track
QuestionWhat do you test models on, and how do you keep the numbers honest?
One clear answer
Mac for iteration, Spark for the server stack. It’s the same harness pointed at both, so the numbers are comparable, and I validate the fabric before I scale across boxes.
What this means
“Mac for iteration”: I try ideas on my Mac first: free, quick and private.
“Spark for the server stack”: I run the production server software, vLLM and SGLang in containers, on an NVIDIA DGX Spark, which a Mac cannot run.
“It’s the same harness pointed at both”: The same test scripts from Days 1, 3 and 5 run against both machines; only the target changes (--target spark).
“so the numbers are comparable”: Any difference comes from the machine and the server software, not from how I measured. The Spark and an M4 Pro Mac both move 273 GB a second, so one person’s writing speed should be close: a gap there is mostly the software. Reading long prompts and serving many people also need math power, where the Spark is far stronger (the lab book: compute-rich, bandwidth-poor).
“and I validate the fabric before I scale across boxes”: Before I split one model over two linked Sparks, I test the link (the fabric) on its own, with a network speed test and NVIDIA’s GPU-to-GPU test (NCCL). A slow link wipes out the gain.
2 worked examples from today’s videos. Work each on paper, then reveal.
From the video:Fitting Big Models3:07
Whiteboard 1 of 2
A customer asks: can we run a 405-billion-parameter model (Llama 3.1 405B) on H100 GPUs, which have 80 GB each and come 8 to a server? 32 people use it at once, each with conversations of up to 8K tokens (8,192, about 6,000 words). At what precision, and on how many GPUs?
Given, in plain words
The model has 405 billion parameters (learned numbers). Each takes 2 bytes at 16 bits, 1 byte at 8 bits and half a byte at 4 bits. The notes (KV cache) come on top. The field guide lists this model as 126 layers (processing stages). Each layer stores 8 note sets of 128 numbers per token, in two kinds (keys and values), at 2 bytes each. Add about 10% for the server’s own working memory (the video’s rule).
Reveal the answerHide the answer
Answer · in plain words
Not at 16 bits: the model alone is bigger than a full 8-GPU server. At 8 bits it fits on one 8-GPU server with room for everyone’s notes. At 4 bits it fits on 4 GPUs only if the notes are stored in 8 bits too, and only if the customer’s quality test passes.
Picture it
The model is furniture sold in three sizes, and each GPU is a van. The full-size set will not fit in the whole fleet of 8. The middle size fits in the fleet with room for everyone’s boxes (the notes). The smallest fits in 3 vans, but the 4th van takes all the boxes only if you pack them tighter.
Careful
The video’s table puts the notes for 32 people at 8K at about 40 GB. Worked from the field guide’s settings for this model they come to about 135 GB, so the steps use 135. The 8-GPU answer holds either way; the 4-GPU answer needs the 8-bit notes.
Worked answer, step by step
Weights at 16 bits: 405 billion × 2 bytes = 810 GB. A server of 8 H100s holds 8 × 80 = 640 GB. 810 is more than 640: no.
Weights at 8 bits (FP8): 405 × 1 = 405 GB, which leaves 640 − 405 = 235 GB on 8 GPUs.
Weights at 4 bits: 405 × 0.5 = 202.5 GB, about 203. That fits on 4 GPUs (320 GB).
Per conversation at 8,192 tokens: 516,096 × 8,192 = 4.23 GB. For 32 people: 32 × 4.23 = 135 GB.
8 bits on 8 GPUs, adding the 10% (× 1.1): (405 + 135) × 1.1 = 594 GB, under 640. Yes.
4 bits on 4 GPUs: (203 + 135) × 1.1 = 372 GB, over 320. No.
4 bits with the notes in 8 bits too (135 ÷ 2 = 68 GB): (203 + 68) × 1.1 = 298 GB, under 320. Yes.
Answer: 8 H100s at 8 bits; or 4 H100s at 4 bits with 8-bit notes, if the customer’s own quality test passes.
Go deeper: the engineer version
The kit's question
Worked example · can we run 405B?
The kit's answer
8×H100 at FP8 — or 4×H100 at 4-bit if quality holds
More detail: The video’s table: FP16 weights 810 GB, no (8×H100 = 640 GB); FP8 405 GB, yes on 8×H100; 4-bit 203 GB, yes on 4×H100; “KV at 8K, 32 users ≈ 40 GB on top”. With the field guide’s config for Llama 3.1 405B (126 layers, 8 KV heads, head dim 128) the KV cache is 504 KiB per token in FP16, 4.23 GB per 8,192-token session and 135 GB for 32 sessions (68 GB in FP8), about 3.4 times the video’s figure. The FP8 plan survives with room to spare; the 4-bit plan on 4×H100 needs FP8 KV or fewer sessions. The 8 GPUs of one server share NVLink, about 900 GB/s on H100, so tensor parallelism across them is the usual split.
Words to know
Parameters
The model’s learned numbers. Example: 405 billion here.
Notes (KV cache)
The model’s memory of each conversation, kept for every token. Example: about 0.52 MB per token for this model.
Precision
How many bits store each number. Example: 16, 8 or 4 here.
H100
NVIDIA’s data-center GPU, with 80 GB of memory.
From the video:Speculation & Scale-Out2:41
Whiteboard 2 of 2
A draft model’s guesses are kept 80% of the time (acceptance rate α = 0.8), and it guesses 4 tokens a round (k = 4). How many tokens does the big model produce per pass, what speed-up is that in practice, and what happens at α = 0.4?
Given, in plain words
Tokens per big pass = (1 − α^(k+1)) ÷ (1 − α), where α^(k+1) means α multiplied by itself k + 1 times. It is the same as 1 + α + α² + ... + α^k: one token of the big model’s own, plus, for each guess, the chance that it and every guess before it are kept. Each round costs one big pass plus 4 guesses by the draft model. Lesson 13’s calculator counts each guess as a tenth of a big pass, a 1-billion-parameter draft for an 8-billion one.
Reveal the answerHide the answer
Answer · in plain words
About 3.4 tokens per big pass instead of 1. After paying for the draft model’s guesses, that is about 2 to 2.4 times faster. At α = 0.4 it drops to about 1.2 times: barely worth it.
Picture it
Autocomplete on a phone: when it guesses 8 of 10 words right, you accept whole runs of words with one tap. When it guesses 4 of 10, you stop to correct so often that it barely saves time. Unlike autocomplete, a wrong guess never changes the text: the big model writes that token itself.
Careful
The video’s “1.6 times at α = 0.4” is the ideal figure, before the draft model’s cost; with lesson 13’s cost rule it is about 1.2 times. The lesson’s sample table (up to 8 guesses a round, so not the same case): α 0.78 on structured prompts gave 1.9 times.
Worked answer, step by step
0.8^5 = 0.8 × 0.8 × 0.8 × 0.8 × 0.8 = 0.328.
Tokens per big pass: (1 − 0.328) ÷ (1 − 0.8) = 0.672 ÷ 0.2 = 3.36, about 3.4 (the video rounds to 3.3).
If guessing were free, that would be 3.4 times faster (the video’s ideal 3.3×).
Guessing costs 4 × 0.1 = 0.4 of a big pass, so each round takes 1.4 passes of time.
Real speed-up: 3.36 ÷ 1.4 = 2.4 times. The video calls it about 2 times in practice.
At α = 0.4: 0.4^5 = 0.01, so (1 − 0.01) ÷ (1 − 0.4) = 0.99 ÷ 0.6 = 1.65 tokens per pass. That is the video’s 1.6 times, before guessing costs.
After guessing costs: 1.65 ÷ 1.4 = 1.18, about 1.2 times: barely worth it.
Go deeper: the engineer version
The kit's question
Worked example · speculation maths
The kit's answer
With an acceptance rate of eighty percent and four drafted tokens, expected tokens per big model pass is one minus zero point eight to the fifth, over one minus zero point eight, which is about three point three. So you get three point three tokens for one verification pass plus four cheap draft passes. Call it a two times speed up in practice.
More detail: E = (1 − α^(k+1)) / (1 − α) is the expected number of tokens per target pass when each drafted token is accepted independently with probability α: the accepted run plus one token from the target. spec_math.py divides by (1 + k·c), with c = 0.1 for a 1B draft on an 8B target: 3.36 ÷ 1.4 = 2.4×. E itself is 3.36, which the lesson rounds to 3.4 and the video to 3.3. The video’s row “at α = 0.4: 1.6×” matches E (1.65), not the cost-adjusted 1.18×.
Words to know
Acceptance rate (α)
The share of guesses the big model keeps. Example: 0.8.
k
Tokens guessed per round. Example: 4.
Draft model
The small model that guesses. Example: a 1-billion-parameter model for an 8-billion one.
Verification pass
One big-model pass that checks all the guesses at once.
How you will use this
Two of Day 10’s 20-second drill questions come from today: “Why is the MoE faster than the smaller dense model?” and “Will speculative decoding help us?”. The sizing memo also quotes your speculation speeds from results/13-spec.csv.