Skip to content

MoE vs dense

course 14 of 22

lesson 12 · lessons/12-moe-vs-dense

40 min Free Mac

What you will do and why

Some of the biggest open models are built as a team of specialists and read only a few of them for each token. You time one on your Mac and work out, from its speed alone, how much of it is read.

Why it matters: A customer sees the larger model file and assumes it will answer more slowly. In the lesson’s sample, that model actually writes 118 word pieces a second against the smaller model’s 84. It reads only the specialist parts it needs for each new piece, so the file’s total size is a poor speed prediction.

You are done when: moe_probe.py printed a row for both models: gpt-oss:20b writes faster than Llama 3.1 8B although its file is about 3 times bigger, and its active_% is well under 100. On a 16 GB Mac, the --demo table counts.

Start this first

About 14 GB. Start it now and watch the video while it downloads. Ollama must be running (ollama serve &).

Terminal window
ollama pull gpt-oss:20b

Fitting Big Models · 3:07

Download mp4 (11.3 MB)

Chapters

In this video How to answer “will it fit?” with a formula, why mixture-of-experts models are big in memory but fast, and the two ways to split a model across GPUs.

The narrator’s “next video” is the file order. Your next step is below.

3 key points

  1. Memory needed = the model’s weights + its notes + about 10% for the server’s working memory.

    The weights are the model’s learned numbers (parameters): count × bytes each. A 405-billion-parameter model at 16 bits (2 bytes each) is 810 GB, more than eight 80 GB H100s (NVIDIA’s data-center chips) hold: 640 GB.

  2. A mixture-of-experts model needs memory for all of itself, but reads only a few experts per token.

    The video’s example on a 128 GB desktop (a DGX Spark): a 120-billion-parameter model stored at 4 bits (half a byte each) is about 60 GB on disk but reads about 5 GB per token. It writes 55 tokens a second, against 23 for a dense 14-billion-parameter one.

  3. If it does not fit: store it in fewer bits, add GPUs, or allow shorter conversations and fewer users. Moving part of the model off the GPU is the last resort.

    Offloading (moving layers to the CPU or disk) is about 10 times slower (the video’s “order of magnitude”), because the model then crosses a much narrower connection than the GPU’s own memory.

Some models are built as a mixture of experts (MoE): many specialist blocks, with a small router that sends each token to only a few of them. The whole model must fit in memory, but its writing speed depends only on the part read for each token. So a big MoE can write faster than a much smaller ordinary (dense) model.

Picture it

A hospital employs dozens of specialists, and every one needs an office: that is the memory. But the receptionist sends each patient to only two to four of them, so a visit is as quick as at a small clinic. A dense model is a clinic where every doctor sees every patient. One difference: the model repeats this choice at each of its layers (processing stages), so one token meets a new small team at every stage.

With real numbersthe lesson’s sample table, lesson 12 (a Mac like an M4 Max)

  • The sample Mac’s memory moves 546 GB a second. The script assumes a model gets 75% of that: 546 × 0.75 = 409.5 GB a second.
  • Dense Llama 3.1 8B: a 4.9 GB file writing 84 tokens a second. Memory delivers 409.5 GB a second, shared by those 84 tokens: 409.5 ÷ 84 = 4.9 GB read per token, the whole file.
  • Mixture-of-experts gpt-oss:20b: a 13.8 GB file, 2.8 times bigger (13.8 ÷ 4.9), yet it writes 118 tokens a second, 1.4 times faster.
  • 409.5 ÷ 118 = 3.5 GB read per token. 3.5 ÷ 13.8 = 25%: about a quarter of the file.
  • The published figure: 3.6 of its 21 billion parameters are used per token, 17%. A stopwatch estimate within 2 times of that is a good result.

Words to know

Mixture of experts (MoE)
A model made of many specialist blocks (experts); for each token a small router picks a few of them. Example: gpt-oss:20b.
Active parameters
The parameters (learned numbers) actually read to write one token. Example: 3.6 of gpt-oss:20b’s 21 billion.
Dense model
A model that uses all of its numbers for every token. Example: Llama 3.1 8B.
Bandwidth (memory bandwidth)
How many GB per second memory delivers to the chip. Example: 546 GB/s on an M4 Max.

From the lesson

lessons/12-moe-vs-dense/README.md

A dense model is one big team where everyone works on every token. A Mixture of Experts (MoE) model is a large company with a receptionist, the router, who sends each token to only 2–4 specialists out of dozens.

memory needed speed depends on
dense 8B all 8B all 8B, read per token
MoE 21B total / 3.6B active all 21B (every expert must be loaded) only ~3.6B, read per token

So MoE gives big-model quality at small-model decode speed, if you have the memory. That is why DeepSeek, gpt-oss, Qwen3-MoE and others are built this way, and why unified memory (Mac, DGX Spark) suits them well.

When the model does not fit, you have three options: offload layers to CPU (slow, because PCIe or system RAM becomes the bottleneck), tensor parallel (split every layer across GPUs; needs fast NVLink; lowers latency) or pipeline parallel (give each GPU a block of layers; tolerates slower links; adds bubbles).

  • moe_probe.py: two lines of physics turn a stopwatch into an architecture estimate.
  • KNOWN holds the published total and active sizes to check your estimate against.
Terminal window
cd lessons/12-moe-vs-dense
python moe_probe.py --demo # see the shape first
ollama pull llama3.1:8b && ollama pull gpt-oss:20b
python moe_probe.py llama3.1:8b gpt-oss:20b
# 64 GB+ Mac? also try: ollama pull qwen3:30b-a3b

What each command does

  1. python moe_probe.py --demo

    Prints the lesson’s sample table without downloading anything: illustrative figures for a Mac like an M4 Max (546 GB a second), not a measured run. Look for gpt-oss:20b writing 118 tokens a second from a 13.8 GB file, faster than the 4.9 GB dense model’s 84.

  2. ollama pull llama3.1:8b && ollama pull gpt-oss:20b

    Downloads both models through Ollama. You have the 8B from Day 1 and started gpt-oss:20b above, so this only finishes what is missing. Look for both downloads to complete without an error.

  3. python moe_probe.py llama3.1:8b gpt-oss:20b

    Times each model writing up to 200 tokens, after a short untimed warm-up, and works backwards to the GB read per token. The first line shows the memory speed it used for your chip. Some chips come in two versions (M3 Max, M4 Max); the script assumes the faster one, so look yours up and add --bw <GB/s>. Look for gpt-oss:20b faster than llama3.1:8b, and its active_% far below 100. Optional on a Mac with 64 GB or more: ollama pull qwen3:30b-a3b (another mixture-of-experts model, 3.3 of its 30.5 billion parameters used per token), then add it to this command.

How to read it

Each row is one model. file_gb is its size, tok_s its writing speed, gb_per_token the GB read per token (memory speed × 75% ÷ tok_s); active_% is that share of the file, published_% the maker’s figure. The dense model reads almost all of itself (99.5%); the MoE about a quarter (25.1%, against 17.1% published), higher because some parts are read for every token and the 75% is an estimate.

model file_gb tok_s gb_per_token active_% published_%
llama3.1:8b 4.9 84.0 4.9 99.5 100.0
gpt-oss:20b 13.8 118.0 3.5 25.1 17.1

The MoE file is about 3× bigger and still faster. Your active_% will come out above the published figure, because attention layers, embeddings and the KV cache are read on every token too. Getting within 2× from a stopwatch is a good result.

3 questions. Say your answer out loud, then tap to check it.

A customer asks why their 120B mixture-of-experts model (120 billion parameters, split among many specialists) writes faster than their 70B dense model (70 billion, all used for every token). What do you tell them?Show answerHide

In plain words

Writing speed depends on how much the model reads for each token, not on its total size. The MoE reads about 5 billion parameters per token; the dense model reads all 70 billion.

Picture it

Looking something up in a 1,200-page encyclopedia is quick when the index sends you to 50 pages. Reading a 700-page manual cover to cover is slow, although it is the smaller book. You still need shelf space for all 1,200 pages.

With real numberslesson 12’s published sizes and the lab book’s DGX Spark measurements

  • gpt-oss 120B: 117 billion parameters in total, 5.1 billion used per token (the published figures in lesson 12’s script).
  • A 70B dense model uses all 70 billion for every token: 70 ÷ 5.1 = about 14 times as much to read.
  • Writing speed is capped at memory speed ÷ bytes read per token, so fewer bytes means faster writing.
  • Measured on a DGX Spark (lab book): gpt-oss 120B writes 55.4 tokens a second; Llama 3.1 70B stored at 8 bits writes 2.7.
  • 55.4 ÷ 2.7 = about 20 times faster, although the MoE holds more parameters (117 billion against 70).
  • The measured gap is bigger than 14 partly because the lab book’s two runs used different server software (llama.cpp and SGLang) and number formats. Read it as roughly 14 to 20 times.

Words to know

Parameters
The model’s learned numbers. Example: 70B means 70 billion.
Active parameters
The parameters read to write one token. Example: 5.1 billion for gpt-oss 120B.
Decode
The writing phase, one token at a time; its speed is set by the bytes read per token.
DGX Spark
NVIDIA’s desktop AI computer: 128 GB of memory moving 273 GB a second.
Go deeper: the engineer version

The kit's question

A customer asks why their 120B MoE is faster than their 70B dense. What do you say?

The kit's answer

About 5B active parameters against 70B read per token, so decode is bandwidth ÷ active bytes.

More detail: Decode is memory-bound: tok/s is about bandwidth × efficiency ÷ bytes read per token. For an MoE those bytes are the chosen experts plus the parts every token uses (attention, embeddings, router) and the KV cache, each at its own precision. The lab book inverts the Spark figure directly: 273 ÷ 55 ≈ 5 GB per token out of a 63 GB file, under 10% of the model. The two figures also differ in format: the 70B is FP8, 1 byte per parameter; gpt-oss 120B is MXFP4, a 4-bit format. Memory is still sized on the total: all 59 GiB (63.4 GB) must fit.

Why can a mixture-of-experts model be harder, not easier, to serve to many people at once on GPUs?Show answerHide

In plain words

Different people’s tokens need different experts, so they cannot all share one read of the model, as they do with a dense model. Spread over several GPUs, tokens must also travel between chips to reach their experts.

Picture it

A bus is efficient because everyone rides the same route. If every passenger needs a different part of town, you end up running many half-empty minibuses. And if the specialists work in different buildings, patients spend time walking between them.

With real numbersDay 3’s practice server and the field guide’s expert example

  • Dense model on Day 3’s practice server: 8 people got about 5 times one person’s total output, because they share each read of the model.
  • The field guide’s MoE example: 4 experts in a layer, 2 used per token.
  • One person: 2 of the 4 experts are read, half of that layer’s experts.
  • Two people whose tokens pick different pairs: all 4 experts are read, and each read serves only one person.
  • With the 4 experts on 4 GPUs, as in the field guide, each token is sent to the GPUs that hold its 2 experts, and the results are sent back.

Words to know

Expert
One specialist block in an MoE; only the few the router picks are read for a token.
Batching
Serving several people with one pass over the model.
Expert parallelism
Placing different experts on different GPUs, so tokens travel to the GPUs that hold their experts.
All-to-all
A network step where every GPU sends data to every other GPU. Example: moving tokens to their experts.
Go deeper: the engineer version

The kit's question

Why can MoE be harder to serve at high concurrency on GPUs?

The kit's answer

Different tokens hit different experts, so batches fragment. Expert parallelism across GPUs and all-to-all traffic also add complexity.

More detail: Batched decode shares one weight read across the batch. In an MoE layer each token’s router picks its own experts, so as the batch grows the set of experts read approaches all of them, while each expert’s share of the batch shrinks: more bytes per step and smaller, less efficient matrix multiplies. Expert parallelism puts experts on different GPUs and needs an all-to-all exchange of token activations in every MoE layer, out to the experts and back. That adds network traffic, and the busiest expert sets the pace.

A model needs 160 GB of memory, and each H100 GPU has 80 GB, so it must be split over 2 GPUs. Do you split every layer (one of the model’s stacked processing stages) across both (TP=2, tensor parallel) or give each GPU half of the layers (PP=2, pipeline parallel)?Show answerHide

In plain words

Split every layer (TP=2) when both GPUs sit in one server with a very fast link between them: each token then finishes sooner. Use the half-the-layers split (PP=2) only when the link between the GPUs is slow.

Picture it

Two cooks can split every dish: each chops half the vegetables, and they combine their halves before every next step. That constant passing only works side by side. Or one cook makes starters and the other mains: one hand-off per order, fine in separate kitchens. But each order waits for both, and a cook stands idle when orders are few.

With real numbersthe check’s numbers, the field guide and lesson 14’s two-box notes

  • 160 GB ÷ 80 GB per GPU = 2 GPUs at the very least, before any room for notes.
  • TP=2: each GPU holds half of every layer, and the two sync inside every layer, many times per token.
  • Inside one server, NVLink joins H100s at about 900 GB a second (field guide): fast enough for constant syncing.
  • Two DGX Sparks are joined by a 200-gigabit cable: 200 ÷ 8 = 25 GB a second. NVLink’s 900 adds both directions together; even halved to 450, it is 18 times the cable (450 ÷ 25).
  • Over a link like that, PP=2 often does better: one hand-off per token between the two halves, instead of syncs in every layer.

Words to know

Tensor parallelism (TP)
Splitting every layer across GPUs so they work on each token together; they sync inside every layer.
Pipeline parallelism (PP)
Giving each GPU a block of layers; each token passes from one GPU to the next.
NVLink
NVIDIA’s very fast direct link between GPUs inside one server. Example: about 900 GB/s on H100.
Pipeline bubble
Time a GPU sits idle in a pipeline, waiting for work from the stage before it.
Go deeper: the engineer version

The kit's question

The model needs 160 GB and each H100 has 80 GB. TP=2 or PP=2?

The kit's answer

TP=2 inside one NVLink node for latency. PP only if the link between GPUs is slow.

More detail: TP shards each weight matrix and all-reduces activations inside every layer, so it cuts per-token latency but is bound by interconnect latency and bandwidth: keep it inside one NVLink domain. PP places consecutive layer blocks on each GPU and passes activations once per stage boundary. It tolerates slower links but adds per-hop latency, and bubbles unless enough requests keep every stage busy. Across boxes the kit’s advice flips to trying PP first, because TP wants NVLink-class bandwidth (TWO_BOX.md). Note that 2 × 80 GB leaves no headroom for a 160 GB model: the Fitting Big Models video adds about 10% for activations and the runtime.

Question Why is the MoE model faster than the smaller dense one?

One clear answer

Memory is set by total parameters; decode speed by active parameters. I can estimate a model’s active footprint from its decode speed on known hardware.

What this means

  • “Memory is set by total parameters”: The machine must hold every expert, the whole model. gpt-oss:20b needs its full 13.8 GB file in memory, though it reads only part of it per token.
  • “decode speed by active parameters”: How fast it writes depends only on what it reads per token: 3.6 of its 21 billion parameters. In the sample it writes 118 tokens a second, against 84 for the 4.9 GB dense 8B.
  • “I can estimate a model’s active footprint from its decode speed on known hardware”: If I know the machine’s memory speed, I time the writing and work backwards: 546 × 0.75 ÷ 118 = 3.5 GB read per token, about a quarter of the file.

Your numbersSaved on this device and collected in the Day 9 wrap-up.

Hint: The tok_s column of the llama3.1:8b row from python moe_probe.py llama3.1:8b gpt-oss:20b.

Hint: The tok_s column of the gpt-oss:20b row. On a 16 GB Mac, use the --demo row.

Hint: The gb_per_token and active_% columns of the gpt-oss:20b row. Write both, for example “3.5 / 25.1”.

moe_probe.py printed a row for both models: gpt-oss:20b writes faster than Llama 3.1 8B although its file is about 3 times bigger, and its active_% is well under 100. On a 16 GB Mac, the --demo table counts.

Connection refused.
Ollama is not running. Start it with ollama serve &, then run the command again.
gpt-oss:20b not pulled — run: ollama pull gpt-oss:20b.
Run that pull. The name must match exactly, including :20b.
pass --bw <GB/s>.
The script does not know your chip’s memory speed. Look it up and add it by hand, for example python moe_probe.py llama3.1:8b gpt-oss:20b --bw 546 for a chip that moves 546 GB a second.
gpt-oss:20b is very slow or will not load.
Your Mac is probably short of memory: the model needs its whole 13.8 GB, and Day 1’s rule keeps models under about 70% of memory. Close other apps, or use the --demo numbers.
moe_probe.py Python · 81 lines
"""
Lesson 12 — infer a model's ACTIVE size from how fast it decodes.
A dense model reads all its weights for every token. A Mixture-of-Experts model has
many "expert" blocks but a router picks only a few per token, so it reads a fraction.
We can see that fraction from the outside, using nothing but a stopwatch:
bytes read per token ≈ bandwidth (GB/s) ÷ measured decode tok/s
active fraction ≈ bytes per token ÷ file size
Uses Ollama's native API because it reports exact eval_count and eval_duration.
ollama pull llama3.1:8b && ollama pull gpt-oss:20b # ~5 GB + ~14 GB
python moe_probe.py llama3.1:8b gpt-oss:20b
python moe_probe.py --demo # sample numbers, no download
"""
import argparse
import json
import urllib.request
from felab import record, table
from felab.hardware import detect
OLLAMA = "http://localhost:11434"
# published totals / active params, for checking your inference (billions)
KNOWN = {"llama3.1:8b": (8.0, 8.0), "gpt-oss:20b": (21.0, 3.6), "qwen3:30b-a3b": (30.5, 3.3),
"qwen3:30b": (30.5, 3.3), "gpt-oss:120b": (117.0, 5.1)}
EFFICIENCY = 0.75 # engines hit ~60–85% of peak bandwidth; we assume 75% when inverting
DEMO = [("llama3.1:8b", 4.9, 84.0), ("gpt-oss:20b", 13.8, 118.0)] # M4 Max-like, illustrative
def _post(path: str, body: dict) -> dict:
req = urllib.request.Request(OLLAMA + path, data=json.dumps(body).encode(),
headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=600) as r:
return json.loads(r.read())
def file_gb(model: str) -> float:
with urllib.request.urlopen(OLLAMA + "/api/tags", timeout=5) as r:
for m in json.loads(r.read())["models"]:
if m["name"] == model or m["name"] == model + ":latest":
return m["size"] / 1e9
raise SystemExit(f"{model} not pulled — run: ollama pull {model}")
def decode_tps(model: str) -> float:
_post("/api/generate", {"model": model, "prompt": "hi", "stream": False, "options": {"num_predict": 4}}) # load
r = _post("/api/generate", {"model": model, "stream": False, "options": {"num_predict": 200, "temperature": 0},
"prompt": "Write 150 words about why caches make computers fast."})
return r["eval_count"] / (r["eval_duration"] / 1e9)
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("models", nargs="*", default=["llama3.1:8b", "gpt-oss:20b"])
ap.add_argument("--bw", type=float, help="GB/s (default: detect)")
ap.add_argument("--demo", action="store_true")
a = ap.parse_args()
bw = a.bw or (546 if a.demo else detect()["bw_gbs"])
if not bw:
raise SystemExit("pass --bw <GB/s>")
data = DEMO if a.demo else [(m, file_gb(m), decode_tps(m)) for m in a.models]
rows = []
for m, size, tps in data:
read = bw * EFFICIENCY / tps # GB actually streamed per token
total, active = KNOWN.get(m, (None, None))
rows.append({"model": m, "file_gb": size, "tok_s": tps, "gb_per_token": read,
"active_%": 100 * min(1.0, read / size),
"published_%": 100 * active / total if total else float("nan")})
record("12-moe", {"bw_gbs": bw, **rows[-1]})
print(f"bandwidth {bw} GB/s × {EFFICIENCY:.0%} efficiency\n")
print(table(rows))
print("\nThe MoE file is ~3× bigger, yet it decodes FASTER: it only reads its active experts.\n"
"Memory must hold ALL experts (file_gb); speed depends on ACTIVE bytes (gb_per_token).")
if __name__ == "__main__":
main()