Skip to content

Speculative decoding

course 15 of 22

lesson 13 · lessons/13-speculative-decoding

50 min 30 to 45 min of one deployment Mac, then Fireworks

What you will do and why

A small draft model guesses the next few tokens and the big model checks them all in one pass. You measure when that speeds up writing and when it slows it down, on your Mac and on a paid Fireworks deployment.

Why it matters: In the lesson’s sample, the draft model lifts predictable output (lists, code, JSON data) from 62 to 118 tokens a second, about 1.9 times. Creative prose only goes from 61.5 to 70.

You are done when: Your Mac’s table has four rows (baseline and speculation, structured and prose) with an acceptance rate on each speculation row. The Fireworks run finished, and its deployment is gone from the list it printed.

Speculation & Scale-Out · 2:41

Download mp4 (10.1 MB)

Chapters

In this video How a small draft model speeds up a big one, why the acceptance rate decides whether it pays, and what linking a second machine does and does not buy.

The narrator’s “next video” is the file order. Your next step is below.

3 key points

  1. A small draft model guesses a few tokens; the big model checks them all in one pass.

    Correct guesses are free tokens, and the big model still decides every token, so the output does not change. With 75% of guesses kept and 4 guesses a round, it writes about 3 tokens per big pass instead of 1.

  2. The acceptance rate, the share of guesses kept, decides whether it pays.

    At 80% kept with 4 guesses, the video estimates about 2 times faster once the draft model’s own time is paid. At 40% it calls the gain barely worth it, and a mismatched draft model can make writing slower.

  3. A second machine adds memory, not speed.

    Two linked 128 GB desktops (256 GB) can run a 235-billion-parameter model at 4 bits, but it writes about 12 tokens a second (11.7). Good for testing that model, not for serving many users.

Rewatch from Day 3: Batching & Scheduling, 1:01 to 1:23ShowHide

Normally the big model writes one token per pass, and each pass reads the whole model. With speculative decoding, a small draft model quickly guesses the next few tokens. The big model checks all the guesses in one pass: it keeps them up to the first one it would not have written, then writes the next token itself. The output is the same, with fewer big passes.

Picture it

An intern drafts the next few words of an email, and a senior engineer approves or corrects them at a glance. With a predictable email, such as a form letter, most words pass and it goes fast. With a creative one, most guesses are wrong, and checking them wastes the senior’s time.

With real numbersthe lesson’s formula and sample table, lesson 13

  • Acceptance rate α (alpha): the share of guesses the big model keeps. Take α = 0.8, with 4 guesses a round (k = 4).
  • The big model always adds 1 token of its own. Guess 1 is kept 80% of the time (0.8). Guess 2 counts only if guess 1 was kept too: 0.8 × 0.8 = 0.64. Guess 3: 0.51. Guess 4: 0.41.
  • Add them: 1 + 0.8 + 0.64 + 0.51 + 0.41 = 3.36 tokens per big pass instead of 1. The lesson’s formula, (1 − 0.8^5) ÷ (1 − 0.8), is a shortcut for this sum.
  • The draft model also runs 4 times a round, at about a tenth of a big pass each: 1 + 4 × 0.1 = 1.4 passes of time. 3.36 ÷ 1.4 = 2.4 times faster.
  • With α = 0.3 and 8 guesses: 1 + 0.3 + 0.09 + 0.03 + ... = 1.43 tokens per pass, but 1 + 8 × 0.1 = 1.8 passes of time. 1.43 ÷ 1.8 = 0.79: slower than no draft at all.
  • The lesson’s sample: predictable prompts 62 to 118 tokens a second (1.9 times, α 0.78); prose 61.5 to 70 (1.14 times, α 0.41).

Words to know

Draft model
A small, fast model that guesses the next few tokens. Example: Llama 3.2 1B guessing for Llama 3.1 8B.
Acceptance rate (α, alpha)
The share of the draft’s guesses the big model keeps. Example: 0.78 on predictable prompts.
k
How many tokens the draft model guesses per round. Example: up to 8 on your Mac, 4 on Fireworks.
Verification pass
One pass of the big model that checks all the guesses at once, the way it reads a prompt.

From the lesson

lessons/13-speculative-decoding/README.md

Decode is slow because the big model writes one token per pass. Speculative decoding has a small, fast draft model guess the next k tokens, and the big model checks all k in one pass. Checking is parallel, like prefill, so it is cheap. Every guess it accepts is a free token. Every miss costs a little wasted drafting.

It is an intern drafting an email while the senior engineer approves or corrects it word by word. It is fast when the intern is predictable, and slower than no intern when they aren’t.

E (tokens per big-model pass) = (1 − α^(k+1)) / (1 − α) α = acceptance rate
α 0.8, k 4 → 3.4 tokens per pass → ~2.4× faster
α 0.3, k 8 → ~0.8× → SLOWER

Structured output (JSON, code, extraction) has high α. Creative prose has low α. Always measure α on the customer’s own traffic.

  • spec_math.py builds the α × k table. Find the ↓ cells.
  • serve_spec_llamacpp.sh needs a draft with the same tokenizer as the target, because the ids are compared directly.
  • spec_bench.py compares the two traffic types and reads timings.draft_n_accepted (llama.cpp) and, on Fireworks, prints one perf_metrics sample (the alpha column stays nan there, so read acceptance in that sample).
  • fireworks_spec.sh sets a trap … EXIT so the deployment is deleted even if you press Ctrl-C.
Terminal window
cd lessons/13-speculative-decoding
python spec_math.py
bash serve_spec_llamacpp.sh off & sleep 30
python spec_bench.py --target llamacpp --label baseline; kill %1
bash serve_spec_llamacpp.sh & sleep 40
python spec_bench.py --target llamacpp --label spec; kill %1
bash fireworks_spec.sh # paid: deploy, measure, then auto-delete

What each command does

  1. python spec_math.py

    Prints the formula as a table: the speed-up for acceptance rates from 0.3 to 0.9 and for 2, 4, 6 or 8 guesses a round, counting each guess as a tenth of a big pass. Look for the two cells marked with a down arrow (slower than no draft): 0.89 and 0.79, at 0.3 with 6 and 8 guesses. The best cell is 3.40, at 0.9 with 8 guesses.

  2. bash serve_spec_llamacpp.sh off & sleep 30

    Starts llama.cpp in the background with Day 2’s 4-bit Llama 3.1 8B and no draft model, then waits 30 seconds for it to load. This is the baseline. Look for baseline (no draft) on :8080.

  3. python spec_bench.py --target llamacpp --label baseline; kill %1

    Sends 5 predictable prompts (lists, code, JSON data) and 5 creative ones, and reports the writing speed for each kind. Then kill %1 stops the server. %1 is this terminal’s first background job, so run jobs first if you started others here. Look for two baseline rows at about the same speed (62.0 and 61.5 in the sample) with alpha nan: no draft, so nothing to accept.

  4. bash serve_spec_llamacpp.sh & sleep 40

    Starts the same model with a draft model, Llama 3.2 1B, which shares its tokenizer (it cuts text into the same numbered tokens) and guesses up to 8 tokens a round. The first start also downloads the draft model. Look for the line speculative: target, followed by both model names.

  5. python spec_bench.py --target llamacpp --label spec; kill %1

    The same 10 prompts with speculation on, then the server stops. Look for the structured row far faster than its baseline (118.0 against 62.0 in the sample, 1.9 times) with a high alpha (0.780), and the prose row only a little faster (70.0) with a low one (0.410).

  6. bash fireworks_spec.sh

    Paid. Creates a Fireworks deployment of Llama 3.1 8B with a 1B draft model guessing 4 tokens a round. It prints a dot every 20 seconds until ready, runs the same 10 prompts, then deletes the deployment, even on Ctrl+C. Look for the fireworks perf_metrics sample: line and two fw-spec rows. Their alpha is nan: read acceptance in the perf_metrics sample instead. Their tok_s includes the trip over the internet, so compare the two rows with each other, not with your Mac’s. There is no Fireworks run without a draft, so no speed-up figure. The last lines list your deployments: yours should be gone.

How to read it

Four rows from your Mac, two per run: label (baseline or spec), traffic (structured or prose), tok_s (writing speed) and alpha (the share of guesses kept, nan with no draft). Compare each spec row with the baseline row of the same traffic. In the sample, structured goes from 62.0 to 118.0 (1.9 times) and prose from 61.5 to 70.0 (1.14 times): the predictable traffic gains most.

label traffic tok_s alpha
baseline structured 62.0 nan
baseline prose 61.5 nan
spec structured 118.0 0.780 ← ~1.9×
spec prose 70.0 0.410 ← small gain; could be a loss with bigger k

3 questions. Say your answer out loud, then tap to check it.

Why must the small draft model share the big model’s tokenizer? (The tokenizer is the part that cuts text into tokens and gives each token a number.)Show answerHide

In plain words

The big model checks the guesses as token numbers, not as text. If the two models cut and number text differently, the same number means different words, and the check means nothing.

Picture it

Two warehouse clerks check an order by product code. That only works if they use the same catalogue: in one, code 4521 is a lamp; in the other, a sofa.

With real numbersthe lesson’s two models, lesson 13

  • Big model: Llama 3.1 8B at 4 bits. Draft: Llama 3.2 1B at 4 bits, from the same family, so it uses the same token numbers.
  • The draft sends up to 8 guesses a round, as token numbers (--draft-max 8).
  • The big model compares each guessed number with the number it would have picked, and keeps the matching run.
  • On Fireworks the lesson uses the same pair, Llama 3.1 8B checking a Llama 3.2 1B draft, with 4 guesses a round.

Words to know

Tokenizer
The part of a model that cuts text into tokens and numbers them. Example: Llama 3.1 8B and Llama 3.2 1B share one.
Token ID
The number a tokenizer gives each token; models compare these numbers, not the text.
Vocabulary
The full list of tokens a tokenizer knows, each with its own number.
Go deeper: the engineer version

The kit's question

Why must the draft share the tokenizer?

The kit's answer

The target verifies token ids, not text.

More detail: Verification compares the draft’s proposed token ids with the target’s own next-token choices position by position (with sampling, it accepts against the target’s probabilities), so both models must share one vocabulary and tokenization. A draft from another family proposes ids that mean different strings to the target. N-gram speculation and predicted outputs avoid the issue, because their drafts come from the target’s own tokens.

Why does speculative decoding help less when the server is busy with many people at once (high concurrency)?Show answerHide

In plain words

With one person, the chip spends most of each step waiting for memory, so it has spare math power to check guesses almost for free. With many people, that spare power already goes to the other people’s tokens, so checking guesses now costs real time.

Picture it

A delivery van on a single drop has lots of empty space, so carrying extra parcels on the off chance costs nothing. Once it is full of other customers’ parcels, every extra parcel pushes out a real one.

With real numbersthe lab book’s DGX Spark measurements and lesson 13’s formula

  • One person: writing reads the whole model for each token, and the math units mostly wait (Day 3: memory-bound).
  • Llama 3.1 8B at 8 bits on a DGX Spark (lab book): 20.5 tokens a second for 1 person, 368 in total for 32.
  • 368 ÷ 20.5 = 18 times the output from the same memory reads: the spare math is now in use.
  • Checking 4 guesses means the math for 5 tokens per person instead of 1, and every rejected guess is wasted math.
  • So speculation pays most for one or a few people who want fast replies.

Words to know

Latency-bound
A workload where what matters is how fast each reply arrives, typical of one or a few people at a time.
Memory-bound
Slowed by how fast memory delivers data, not by math. Example: writing for one person.
Batch size
How many requests share one pass over the model. Example: 1, then 32, on the Spark.
Compute
The chip’s math power. Example: what checking guesses uses.
Go deeper: the engineer version

The kit's question

Why does speculation help less at high concurrency?

The kit's answer

With a full batch the GPU is already busy, so spare compute for verification shrinks. It shines on latency-bound, low-batch workloads.

More detail: At batch 1, decode does about two FLOPs per parameter (1 to 2 per byte fetched at 16 or 8 bits), against an H100 ridge point of about 300 (989 teraflops ÷ 3.35 TB/s, the GPU Bandwidth video), so compute sits idle and verifying k drafted tokens in the same weight read is almost free. As the batch grows, arithmetic intensity rises toward the compute roof; verification adds k + 1 positions per sequence, and rejected positions become wasted compute that displaces other users’ tokens. Speculation shines on latency-bound, low-batch workloads; at high concurrency its gain shrinks and can turn negative.

For editing tasks, where most of the answer repeats the input, what can you use instead of a draft model?Show answerHide

In plain words

Use text you already have as the guess. N-gram speculation finds the last few words it wrote earlier in the prompt and guesses what came next there. Fireworks’ predicted outputs let you send the expected answer, such as the original document, as the draft.

Picture it

Proofreading: instead of an intern drafting from scratch, you hand the editor last week’s version. Most of it passes at a glance, and only the changed lines need real work.

With real numbersfireworks_spec.sh, the lab book and lesson 13’s formula

  • The script’s comment names the option: --ngram-speculation-length=3, no draft model; it reuses n-grams (short runs of tokens) from the prompt.
  • The lab book lists four ways on Fireworks: the default draft model, your own draft model, n-gram, or predicted outputs.
  • Suppose an edit keeps 9 of 10 guesses (α = 0.9, the highest rate in lesson 13’s table; measure it on real edits). With 4 guesses: 1 + 0.9 + 0.81 + 0.73 + 0.66 = 4.1 tokens per big pass.
  • With no draft model to run, guessing costs almost nothing, so the ideal speed-up approaches those 4.1 times. Real gains come in lower.

Words to know

N-gram
A run of n tokens in a row. Example: a 3-gram is 3 tokens.
N-gram speculation
Guessing the next tokens by copying what followed the same words earlier in the prompt; no draft model.
Predicted outputs
A Fireworks option: you send the expected answer as the draft for the model to check.
Go deeper: the engineer version

The kit's question

What is the zero-model alternative for editing tasks?

The kit's answer

N-gram speculation or Fireworks predicted outputs: the draft comes from the prompt itself.

More detail: N-gram (prompt-lookup) speculation matches the last few generated tokens against earlier text in the context and proposes the tokens that followed; predicted outputs let the client send the expected completion, for example the original file for a code edit. Both need no draft model and suit edits, rewrites and extraction, where the output copies long spans of the input. Measure acceptance on the real traffic, as with any drafter.

Question Will speculative decoding help us?

One clear answer

Acceptance rate decides whether speculation pays. On structured traffic it can double decode speed; a generic drafter on unusual traffic can make things slower. I measure α on their prompts before recommending it.

What this means

  • “Acceptance rate decides whether speculation pays”: The share of the small model’s guesses the big model keeps is the number that matters. At 0.8 with 4 guesses it is 2.4 times faster; at 0.3 with 8 guesses, 0.79 times: slower.
  • “On structured traffic it can double decode speed”: Decode speed is writing speed. Predictable output such as JSON data, code or extraction is easy to guess. In the lesson’s sample it went from 62 to 118 tokens a second, 1.9 times.
  • “a generic drafter on unusual traffic can make things slower”: A draft model that has not seen this kind of text guesses badly, and the wasted guesses cost more than they save.
  • “I measure α on their prompts before recommending it”: I run the customer’s own prompts with speculation on and read the acceptance rate (α) before I promise any speed-up.

Your numbersSaved on this device and collected in the Day 9 wrap-up.

Hint: The tok_s of the two structured rows: baseline, then spec. Write both, for example “62.0 / 118.0”.

Hint: The tok_s of the two prose rows: baseline, then spec.

Hint: The alpha column of the two spec rows.

Hint: The two fw-spec rows from bash fireworks_spec.sh. Their alpha is nan; add the acceptance figure from the perf_metrics line if you find one.

Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.

Your Mac’s table has four rows (baseline and speculation, structured and prose) with an acceptance rate on each speculation row. The Fireworks run finished, and its deployment is gone from the list it printed.

spec_bench.py cannot connect right after a server start.
The server was still loading. The first speculative start also downloads the draft model, which can take longer than the 40-second wait. The ; kill %1 on that line then stops the server too: start it again (bash serve_spec_llamacpp.sh &), wait until it says it is listening, then run the spec_bench.py line again.
The server will not start because port 8080 is in use.
Another llama.cpp server from Day 2, 4 or 5 is still running. Stop it with Ctrl+C in its terminal, then start again.
kill %1 says there is no such job, or stops the wrong program.
%1 is the first background job of this terminal. Run jobs to list them and use the right number, for example kill %2.
The server refuses the draft model.
The draft must share the big model’s tokenizer. If you changed DRAFT_REPO or LLAMACPP_REPO, set them back, or pick a draft from the same model family.
fireworks_spec.sh stops with an error from firectl.
Check that you are signed in (firectl whoami, Day 1) and that your quota allows a deployment (firectl quota list). Then run make fw-check: it should list nothing.
fireworks_spec.sh never says ready.
Press Ctrl+C: the script deletes the deployment on the way out. Run make fw-check to confirm nothing is left, and try again later.
fireworks_spec.sh Bash · 28 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 13 · step 4 — speculation on a Fireworks dedicated deployment. ⏱ PAID
#
# Deploy → measure → DELETE, in one script, with a trap so Ctrl-C still deletes.
# Budget: ~30–45 min of one deployment. Most supported models already ship a
# default drafter; here we set one explicitly so you can see the knobs:
# --draft-model a small model with the same tokenizer
# --draft-token-count k (start at 4, per the docs)
# (alternative: --ngram-speculation-length=3 — no draft model; reuses n-grams from the prompt)
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
cd "$(dirname "$0")"
TARGET_MODEL=${TARGET_MODEL:-accounts/fireworks/models/llama-v3p1-8b-instruct}
DRAFT_MODEL=${DRAFT_MODEL:-accounts/fireworks/models/llama-v3p2-1b-instruct}
out=$(firectl deployment create "$TARGET_MODEL" --deployment-shape default \
--draft-model="$DRAFT_MODEL" --draft-token-count=4)
echo "$out"
DEP=$(echo "$out" | grep -oE 'deployments/[A-Za-z0-9-]+' | head -1 | cut -d/ -f2)
ACCOUNT=$(echo "$out" | grep -oE 'accounts/[a-z0-9-]+/deployments' | head -1 | cut -d/ -f2)
trap 'echo "▸ deleting $DEP"; firectl deployment delete "$DEP"; firectl deployment list' EXIT
echo "▸ $DEP created $(date +%H:%M) — ⏱ clock running"
until firectl deployment get "$DEP" | grep -qiE "state.*READY"; do sleep 20; echo -n "."; done; echo " ready"
python spec_bench.py --target fireworks --model "$TARGET_MODEL#accounts/$ACCOUNT/deployments/$DEP" --label fw-spec
# the trap deletes the deployment on exit
serve_spec_llamacpp.sh Bash · 26 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 13 · step 2 — llama.cpp with a draft model, on your Mac.
#
# Target: Llama-3.1-8B Q4. Draft: Llama-3.2-1B Q4 (same tokenizer family — REQUIRED:
# the target verifies the draft's token ids, so both must share a vocabulary).
#
# -hfd REPO:QUANT draft model from Hugging Face (--hf-repo-draft)
# --draft-max 8 up to k=8 guesses per step
# --draft-min 1
# -ngld 99 draft on the GPU too
#
# Usage: bash serve_spec_llamacpp.sh # with speculation → :8080
# bash serve_spec_llamacpp.sh off # same target, no draft (the baseline)
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
TARGET=${LLAMACPP_REPO:-bartowski/Meta-Llama-3.1-8B-Instruct-GGUF}:Q4_K_M
DRAFT=${DRAFT_REPO:-bartowski/Llama-3.2-1B-Instruct-GGUF}:Q4_K_M
if [[ "${1:-on}" == "off" ]]; then
echo "▸ baseline (no draft) on :8080"
exec llama-server -hf "$TARGET" -c 8192 -ngl 99 --port 8080
fi
echo "▸ speculative: target $TARGET + draft $DRAFT on :8080"
exec llama-server -hf "$TARGET" -hfd "$DRAFT" -c 8192 -ngl 99 -ngld 99 \
--draft-max 8 --draft-min 1 --port 8080
spec_bench.py Python · 73 lines
"""
Lesson 13 · step 3 — measure speculation on two kinds of traffic.
Runs 5 "structured" prompts (JSON, lists, code — predictable) and 5 "prose" prompts
(creative — unpredictable) and reports decode tok/s for each. Run it twice: against
the baseline server and the speculative one, with --label so the rows line up.
llama-server also returns `timings.draft_n` / `draft_n_accepted` on each response;
if present we print the acceptance rate α directly.
bash serve_spec_llamacpp.sh off & python spec_bench.py --target llamacpp --label baseline
bash serve_spec_llamacpp.sh & python spec_bench.py --target llamacpp --label spec
python spec_bench.py --target fireworks --model "<model>#<deployment>" --label fw-spec
"""
import argparse
import time
from felab import add_target_args, banner, client, record, resolve, table
PROMPTS = {
"structured": [
"Return a JSON array of the 12 months, each as {\"n\": <number>, \"name\": \"<name>\"}.",
"Write a Python function that returns the first 30 Fibonacci numbers, with a docstring.",
"List the numbers from 1 to 60 separated by commas.",
"Convert to JSON: name Ada, role engineer, city London, skills python sql go.",
"Write an HTML table with 8 rows: country and capital for European countries.",
],
"prose": [
"Write a surreal short story about a lighthouse that collects lost umbrellas.",
"Invent a new board game and describe its strangest rule.",
"Write a poem about latency in the voice of a tired sea captain.",
"Describe an alien market using only smells and sounds.",
"Pitch a movie where the villain is a very polite fog.",
],
}
def main() -> None:
ap = add_target_args(argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter))
ap.add_argument("--label", default="run")
ap.add_argument("--max-tokens", type=int, default=256)
a = ap.parse_args()
t = resolve(a)
banner(t)
cli = client(t)
extra = {"extra_body": {"perf_metrics_in_response": True}} if t.is_paid else {}
rows = []
for kind, prompts in PROMPTS.items():
toks = secs = drafted = accepted = 0
for p in prompts:
t0 = time.perf_counter()
r = cli.chat.completions.create(model=t.model, temperature=0, max_tokens=a.max_tokens,
messages=[{"role": "user", "content": p}], **extra)
secs += time.perf_counter() - t0
toks += r.usage.completion_tokens if r.usage else 0
timings = (r.model_extra or {}).get("timings") or {} # llama.cpp
drafted += timings.get("draft_n", 0)
accepted += timings.get("draft_n_accepted", 0)
pm = (r.model_extra or {}).get("perf_metrics") # Fireworks
if pm and kind == "structured" and p == prompts[0]:
print(f" fireworks perf_metrics sample: {pm}")
rows.append({"label": a.label, "traffic": kind, "tok_s": toks / secs if secs else 0.0,
"alpha": accepted / drafted if drafted else float("nan")})
record("13-spec", {"target": t.name, **rows[-1]})
print("\n" + table(rows))
print("\nCompare tok_s between --label baseline and --label spec for each traffic type.")
if __name__ == "__main__":
main()
spec_math.py Python · 35 lines
"""
Lesson 13 · step 1 — when does speculation pay? The maths, before the GPU.
A small DRAFT model guesses k tokens; the big TARGET model checks all k in ONE pass
(checking is parallel, like prefill). Accepted guesses are free tokens.
expected tokens per target pass E = (1 − α^(k+1)) / (1 − α)
speedup ≈ E / (1 + k·c)
α acceptance rate — how often the draft guesses what the target would say
k tokens drafted per step
c cost of one draft token relative to one target token (1B vs 8B ≈ 0.1)
python spec_math.py
python spec_math.py --c 0.2 # a bigger/slower drafter
"""
import argparse
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--c", type=float, default=0.1, help="draft/target cost ratio")
a = ap.parse_args()
alphas = [0.3, 0.5, 0.6, 0.7, 0.8, 0.9]
ks = [2, 4, 6, 8]
print(f"speedup vs normal decoding (draft cost ratio c = {a.c})\n")
print(" α \\ k " + "".join(f"{k:>8}" for k in ks))
for al in alphas:
cells = []
for k in ks:
E = (1 - al ** (k + 1)) / (1 - al)
s = E / (1 + k * a.c)
cells.append(f"{s:7.2f}{'×' if s >= 1 else '↓'}")
print(f" {al:4.1f} " + "".join(cells))
print("\n↓ = slower than no speculation. Structured/repetitive output (JSON, code, extraction)\n"
"has high α; creative prose has low α. So: measure α on the customer's traffic first.")