Quantization
course 6 of 22
lesson 04 · lessons/04-quantization
50 min Free Mac
What you will do and why
Quantization stores a model’s numbers with fewer bits (the 0s and 1s each number is kept in), so the model gets smaller. You store one model at three sizes, see that the smaller one writes faster, then check it still answers well.
Why it matters: Lesson example: the 4-bit version (4.9 GB) writes 86 tokens (word pieces) per second; the 8-bit version (8.5 GB) writes 52. That is 1.65 times faster.
You are done when: The benchmark table shows all three sizes with an efficiency column (expect 60 to 85%), and the quality check has printed a score out of 12 for each size.
Start this first
First, in the 03-labs folder, switch on the lab’s Python environment from Day 1: source .venv/bin/activate. The command downloads the same model at three sizes (about 17 GB), times each one and prints a table. Leave it running while you watch the video below. For the commands further down, open a second terminal window, go to 03-labs, run source .venv/bin/activate, then cd lessons/04-quantization.
cd lessons/04-quantization && bash quant_bench.shQuantization · 2:43
Download mp4 (9.7 MB)Watch while the 17 GB download runs.
Chapters
In this video Why storing a model’s numbers in fewer bits makes it faster and fits more people on one chip, and how to prove the answers stay good.
The narrator’s “next video” is the file order. Your next step is below.
3 key points
Storing a model at 8 bits instead of 16 halves its size; on NVIDIA’s H100 (a data-center AI chip) and newer chips it also about doubles the math speed.
A 70-billion-number model is 140 GB at 16 bits (2 bytes per number) and 70 GB at 8 bits. On one 80 GB H100 that leaves 10 GB spare. Each 8K-token conversation (about 6,000 words) needs about 2.6 GB of notes, so only 3 whole ones fit (10 ÷ 2.6 = 3.8).
Storing it at 4 bits cuts it to a quarter of the 16-bit size, but prove quality on the customer’s own tests first.
The same model is 35 GB at 4 bits, leaving 45 GB spare. At 2.6 GB of notes per conversation that is at most 17 (45 ÷ 2.6 = 17.3); the video plays safe and says about 14. Its rule: if 2 fewer replies in 100 come back in the exact format asked for, but the server writes twice as many tokens per second, that is usually a good trade; 20 fewer in 100 never is.
Storing the conversation notes (the KV cache) in 8 bits is often a bigger win than shrinking the model.
It doubles how many people fit in the same memory. You measure it in step 2: 20 GB of spare memory holds 4 conversations of 32,768 tokens with 16-bit notes, and 8 with 8-bit notes.
Rewatch from Day 3: GPU Bandwidth, 0:32 to 0:54ShowHide
In plain words
Section titled “In plain words”Quantization stores each of the model’s billions of learned numbers (its weights) with fewer bits, so the model gets smaller. Smaller writes faster, because each new token (word piece) means fetching the whole model from memory. Shrink it too far and the answers start to slip.
Picture it
Before every dish, the whole recipe book travels to the cook on a conveyor belt that moves a fixed number of pages per second. Print the book in smaller type and it arrives sooner, so dishes come out faster. Print it too small and the cook starts misreading recipes.
With real numbersLlama 3.1 8B on an M4 Max Mac, lesson 04’s sample figures
- Three sizes of the same model: 8.5 GB at 8-bit, 4.9 GB at 4-bit, 4.0 GB at 3-bit. At 16-bit it would be about 16 GB (8 billion numbers x 2 bytes). (The real formats use a little more than their names: about 8.5, 4.8 and 3.9 bits per number.)
- Each new token needs one full fetch of the model. An M4 Max (a high-end Apple chip) moves 546 GB per second, so it can fetch the 8.5 GB model 546 ÷ 8.5 = 64 times a second: at most 64 tokens per second. At 4.9 GB: 546 ÷ 4.9 = 111.
- Sample writing speed: 52 tokens per second at 8-bit, 86 at 4-bit. That is 58% of the bytes (4.9 ÷ 8.5) for 1.65 times the speed (86 ÷ 52).
- Share of the limit reached at 4-bit: 86 ÷ 111.4 = 77%. Real servers reach 60 to 85% of their limit.
- Reading the prompt barely moves: 1,150 tokens per second at 8-bit, 1,050 at 4-bit.
- Quality alarm: 12 quick questions per size. If 8-bit gets 11 right and 3-bit only 7, you have found the quality cliff (the size below which answers suddenly get worse).
Words to know
- Decode
- The writing phase: the model writes one token at a time and fetches the whole model for each. Example: 86 tokens per second at 4-bit (the
tg128column). - Prefill
- The reading phase: the model reads the whole prompt at once before writing. Example: 1,050 tokens per second at 4-bit (the
pp512column). - Memory bandwidth
- How many GB per second memory delivers to the chip. Example: an M4 Max moves 546 GB/s.
- Ceiling (speed limit)
- The fastest a model can write: memory bandwidth ÷ model size. Example: 546 ÷ 4.9 = 111 tokens per second.
From the lesson
lessons/04-quantization/README.md
What and why
Section titled “What and why”Quantization stores each weight in fewer bits: 16 → 8 → 4. Think of it as a photo saved at lower JPEG quality. The file is smaller and loads faster, and past a certain point you notice the blur.
Decode reads every weight once per token, so fewer bytes means proportionally more tokens:
decode ceiling (tok/s) ≈ memory bandwidth (GB/s) ÷ model size (GB)
M4 Max 546 GB/s ÷ 8B @ Q8 (8.5 GB) ≈ 64 tok/sM4 Max 546 GB/s ÷ 8B @ Q4 (4.9 GB) ≈ 111 tok/s ← same model, ~1.7× fasterReal engines reach 60–85% of that ceiling. Prefill barely changes with quantization because it is compute-bound. That contrast is the whole lesson.
Read the code first
Section titled “Read the code first”quant_bench.sh:llama-benchruns the engine with no server, so you measure the engine alone.ceiling_check.py: the formula above in six lines, printed next to what you measured.quality_check.py: 12 questions as a smoke alarm for the quality cliff.
cd lessons/04-quantizationpython ceiling_check.py --demo # see the shape first, no download
bash quant_bench.sh # downloads Q8_0, Q4_K_M, Q3_K_M and benches them
for q in Q8_0 Q4_K_M Q3_K_M; do # quality per quant bash ../02-three-local-servers/serve_llamacpp.sh $q & sleep 20 python quality_check.py --target llamacpp --label $q kill %1doneWhat each command does
python ceiling_check.py --demoPrints the lesson’s sample table for an M4 Max (546 GB per second) without downloading anything. Look for writing speed (
tg128) rising as the model shrinks, 52 to 86 to 98 tokens per second, while reading speed (pp512) stays near 1,000. It also saves these sample rows in03-labs/results/04-quant.csv, so do not quote them as your own.bash quant_bench.shDownloads the model at three sizes (about 17 GB, the long part), times each with
llama-bench(llama.cpp’s timing tool, run without a server), then prints the same table for your Mac. If you started it before the video, let that run finish instead of starting a second one. Look for everyefficiency_%(the share of the speed limit reached) between 60 and 85.for q in Q8_0 Q4_K_M Q3_K_M; doStarts the server with each size in turn, asks it 12 quick questions, then stops it. Run it after
quant_bench.shfinishes. Look for a line such asQ8_0: 11/12per size; a clear drop at 3-bit is the quality cliff. The first run downloads the 8-bit and 3-bit versions again (12.5 GB) and can outlast the loop’s 20-second wait: see Stuck? if a size prints a connection error.
What you should see
Section titled “What you should see”How to read it
Each row is one size of the same model. size_gb is its size; pp512 how fast it reads a prompt; tg128 how fast it writes (both in tokens per second). ceiling is its speed limit (memory speed ÷ size); efficiency_% the share of that limit it reached. From Q8_0 (8-bit) to Q4_K_M (4-bit), writing rises from 52 to 86 (1.65 times) while reading stays about the same. These are the lesson’s sample figures for an M4 Max; your rows follow your own Mac’s memory speed and print smallest model first.
quant size_gb pp512 tg128 ceiling efficiency_%Q8_0 8.5 1,150 52.0 64.2 81.0Q4_K_M 4.9 1,050 86.0 111.4 77.2 ← decode ×1.65, prefill ~sameQ3_K_M 4.0 980 98.0 136.5 71.8Quality usually holds at Q8 and Q4_K_M, and Q3 starts dropping the multi-step and format questions first.
Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
Shrinking the stored model from 8 bits to about 4 nearly doubles how fast it writes (decode), but barely changes how fast it reads your prompt (prefill). Why?Show answerHide
In plain words
Writing waits on memory: each new token needs the whole model fetched once, so a smaller model arrives faster. Reading your prompt waits on math instead, and fewer bits do not make a Mac’s math faster.
Picture it
Before every dish, the whole recipe book comes to the cook on a conveyor belt that moves a fixed number of pages per second. Halve the pages and the cook turns out nearly twice as many dishes a minute. Reading the prompt is one delivery followed by lots of cooking, so a thinner book barely helps.
With real numberslesson 04 sample figures, M4 Max
- Model size: 8.5 GB at 8-bit, 4.9 GB at 4-bit (58% of the bytes).
- Writing: 52 tokens/s becomes 86, 1.65 times faster.
- Reading the prompt: 1,150 tokens/s becomes 1,050, about the same (9% slower).
- Speed limits (M4 Max memory speed, 546 GB/s, ÷ model size): 546 ÷ 8.5 = 64 and 546 ÷ 4.9 = 111 tokens/s.
Words to know
- Decode
- The writing phase, one token at a time. Example: the
tg128column (write 128 tokens) measures it. - Prefill
- The reading phase, the whole prompt at once. Example:
pp512measures it. - Bandwidth-bound
- Slowed by how fast memory delivers data (bandwidth, in GB per second). Example: decode; an M4 Max delivers 546 GB/s.
- Compute-bound
- Slowed by calculation speed rather than memory. Example: prefill.
Go deeper: the engineer version
The kit's question
Why does halving the bytes nearly double decode but not prefill?
The kit's answer
Decode is bandwidth-bound; prefill is compute-bound.
More detail: A batch-1 decode step does about two floating-point operations per parameter while streaming every weight (the GPU Bandwidth video), so bytes set the pace. Prefill (pp512) reuses each weight read across 512 tokens, so matrix-math throughput sets the pace. llama.cpp’s smaller formats must be unpacked before the math, which is why prefill even dips slightly.
On an NVIDIA H100 (a GPU: the kind of chip that runs AI models in data centers), models are often stored as FP8, an 8-bit number format. Where does that fit in with shrinking a model to fewer bits (quantization) on a Mac?Show answerHide
In plain words
It is the same shrinking idea, but the H100 can also do its math directly on 8-bit numbers. So both fetching the model and doing the math get faster, not only the fetching.
Picture it
On a Mac, a thinner recipe book only makes each delivery faster. On an H100 the cook has also learned the book’s shorthand, so the cooking speeds up too.
With real numbersthe GPU Bandwidth video (Day 3) and the Quantization video
- A 70B model: 70 billion x 2 bytes = 140 GB at 16-bit; 70 GB at 8-bit; 35 GB at 4-bit.
- H100 memory speed: 3,350 GB/s, about 6 times an M4 Max’s 546.
- Writing-speed limit: 3,350 ÷ 70 = 48 tokens/s at 8-bit; 3,350 ÷ 35 = about 96 at 4-bit.
- The extra on H100: built-in 8-bit math, roughly double the math speed, so prompt reading speeds up too.
- The next generation, Blackwell, adds built-in 4-bit math (NVFP4).
Words to know
- FP8
- An 8-bit number format, 1 byte per number. Example: a 70B model is 70 GB in FP8.
- Tensor cores
- The parts of an NVIDIA GPU built for fast matrix math. Example: the H100’s tensor cores do math directly on FP8 numbers.
- NVFP4
- NVIDIA’s 4-bit number format with built-in math on Blackwell GPUs.
- Blackwell
- The NVIDIA GPU generation after H100. Example: the B200 chip, whose memory moves 8,000 GB/s.
Go deeper: the engineer version
The kit's question
Where does FP8 on an H100 fit in?
The kit's answer
Same idea in hardware. H100/H200 have FP8 tensor cores, and Blackwell adds NVFP4, so the compute also speeds up rather than just memory.
More detail: On Apple silicon the GGUF formats only cut bytes; the math still runs at higher precision, which is why prefill stayed flat in the previous question. On H100, FP8 weights halve decode bytes and FP8 tensor cores roughly double matrix-math throughput, which helps prefill and large batches. The Quantization video: “FP8 on H100+ is close to free; four-bit needs measurement.”
The customer worries that a smaller, faster version of the model will give worse answers. What do you do?Show answerHide
In plain words
Test it instead of arguing. Run 50 of their own real prompts at each size and show quality, speed and cost side by side.
Picture it
A bakery wants to switch to a cheaper flour. You bake their best-sellers both ways and let their own customers taste, instead of quoting a lab report.
With real numberslesson 04’s quality check and the Quantization video’s sample table
- Lesson 04’s quick alarm: 12 questions per size. If the 3-bit version gets 7 right where the 8-bit gets 11, quality has fallen off a cliff.
- For a customer: 50 of their own prompts at each size (lesson 09, on Day 7, builds the test).
- The Quantization video’s example table (made-up numbers to show the shape): at 8-bit instead of 16-bit, accuracy slipped from 89% to 88%, total output rose from 24 to 46 tokens per second, conversations per GPU from 3 to 7.
- The video’s rule: if 2 fewer replies in 100 come back in the exact format asked for, but the server writes twice as many tokens per second, that is usually a good trade; 20 fewer in 100 never is.
Words to know
- Eval
- A repeatable test of a model on a fixed set of prompts. Example: 50 of the customer’s prompts.
- Precision
- How many bits each stored number gets. Example: 16-bit, 8-bit, 4-bit.
- Quality cliff
- The size below which answers suddenly get worse. Example: the 3-bit version (Q3) slips on multi-step and format questions first.
- Throughput
- Total output across everyone at once. Example: 24 against 46 tokens/s in the Quantization video’s sample.
Go deeper: the engineer version
The kit's question
The customer is worried about quality. What do you do?
The kit's answer
Run their own 50-prompt eval at each precision and show the quality, latency and cost table from lesson 09.
More detail: Lesson 09’s table has valid, accuracy, judge, p95 and $/1k tasks columns, and says never to show one column alone. Pick the precision per workload, on held-out prompts. The Quantization video states its rule in schema validity and throughput.
Explain what you learned
Section titled “Explain what you learned”Question Should we run this model in FP8 or 4-bit?
One clear answer
Decode speed is bandwidth over bytes per token. Quantization changes the denominator. I never recommend a precision without running the customer’s own eval at that precision.
What this means
- “Decode speed”: How fast the model writes its reply, one token at a time.
- “is bandwidth over bytes per token”: The top speed is how many GB per second memory delivers, divided by how many GB must be fetched for each token (the whole model). M4 Max: 546 ÷ 4.9 = about 111 tokens/s.
- “Quantization changes the denominator”: Storing the model in fewer bits shrinks the bottom number, so the result goes up. 8.5 GB gives about 64 tokens/s; 4.9 GB gives about 111.
- “I never recommend a precision without running the customer’s own eval at that precision”: Precision means how many bits each number gets. Before saying “use 4-bit”, I test each option on the customer’s own prompts, because fewer bits can cost accuracy. For this question: FP8 on an H100 usually costs little quality; the Quantization video calls it “close to free”. 4-bit buys more speed and room but carries more risk, so it needs that test most.
Your numbersSaved on this device and collected in the Day 4 wrap-up.
Hint: The tg128 column of your quant_bench.sh table. Your table lists the 3-bit row (Q3_K_M) first, so read it bottom to top: 8-bit, 4-bit, 3-bit, for example “52 / 86 / 98”.
Hint: The efficiency_% column, in the same order: bottom row first. 60 to 85% is normal. Well below 60% means something else is slowing the Mac: step 1’s Stuck? section lists the usual causes.
Hint: The line such as Q8_0: 11/12 that each run prints after its 12 questions (server log lines may surround it). The scores are also saved in 03-labs/results/04-quality.csv (column score).
Done when
Section titled “Done when”The benchmark table shows all three sizes with an efficiency column (expect 60 to 85%), and the quality check has printed a score out of 12 for each size.
Stuck?
Section titled “Stuck?”- The download stops because the disk is full.
- The three sizes need about 17 GB, and the quality loop later stores the 8-bit and 3-bit versions again (about 12.5 GB). Free some space and run
bash quant_bench.shagain: sizes already downloaded are skipped. Once its table has printed, you can delete themodelsfolder insidelessons/04-quantization. If the disk filled during the quality loop instead, delete that folder now (the speed test no longer needs it) and run the loop again. - It stops with
Unknown bandwidth — pass --bw. - The script’s list of Apple chips does not include yours. The error lists the chips it knows with their speeds; your Mac’s tech specs page on Apple’s website lists its memory bandwidth. Then run
python ceiling_check.py --bw 273, with your figure in place of 273. The timings are already saved, so nothing downloads again. - Efficiency is well below 60%.
- Something else is slowing the Mac: another app using memory heavily, a hot Mac slowing itself down, or part of the model running on the Mac’s main processor instead of its graphics cores. Close big apps (video calls, browser tabs playing video), plug in, and run
bash quant_bench.shagain; downloads are skipped if you have not deletedmodels/. Some chips, such as the M3 Max and M4 Max, come in two versions, and the script assumes the faster one (felab/hardware.py). If yours is slower, runpython ceiling_check.py --bwfollowed by your figure in GB per second. - The quality loop prints a connection error (for example
Connection refused) for one size. - The server was still downloading that size when the questions started: the loop waits only 20 seconds. Start that size alone, for example
bash ../02-three-local-servers/serve_llamacpp.sh Q8_0, wait for the line saying it is listening, stop it with Ctrl+C, then run the loop again. llama-bench: command not found- llama.cpp is missing. Day 1’s setup installs it; to add it again, run
brew install llama.cpp, thenbash quant_bench.sh.
Code in this step
Section titled “Code in this step”ceiling_check.py Python · 73 lines
"""Lesson 04 · step 2 — does the physics hold?
For each llama-bench result in bench/*.json:
ceiling tok/s = memory bandwidth (GB/s) ÷ model size (GB) efficiency = measured tg128 ÷ ceiling (expect 60–85%)
If efficiency is way below 60%, something else is the bottleneck: layers on the CPU(-ngl too low), thermal throttling, or another app hogging memory bandwidth.
python ceiling_check.py # auto-detects your Mac's bandwidth python ceiling_check.py --bw 546 # or tell it (GB/s) python ceiling_check.py --demo # no downloads: sample numbers, see the shape"""import argparseimport jsonfrom pathlib import Path
from felab import record, tablefrom felab.hardware import APPLE_BW, ceiling, detect
HERE = Path(__file__).parent
# Illustrative figures for an 8B model on an M4 Max (546 GB/s), used by --demo only.DEMO = [ {"quant": "Q8_0", "size_gb": 8.5, "pp512": 1150.0, "tg128": 52.0}, {"quant": "Q4_K_M", "size_gb": 4.9, "pp512": 1050.0, "tg128": 86.0}, {"quant": "Q3_K_M", "size_gb": 4.0, "pp512": 980.0, "tg128": 98.0},]
def load_bench() -> list[dict]: """Parse llama-bench -o json files into {quant, size_gb, pp512, tg128}.""" rows = [] for f in sorted((HERE / "bench").glob("*.json")): runs = json.loads(f.read_text()) pp = next((r["avg_ts"] for r in runs if r.get("n_prompt", 0) > 0 and r.get("n_gen", 0) == 0), 0) tg = next((r["avg_ts"] for r in runs if r.get("n_gen", 0) > 0 and r.get("n_prompt", 0) == 0), 0) size = runs[0].get("model_size", 0) / 1e9 rows.append({"quant": f.stem, "size_gb": size, "pp512": pp, "tg128": tg}) return rows
def main() -> None: ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--bw", type=float, help="memory bandwidth GB/s (default: detect)") ap.add_argument("--demo", action="store_true", help="use built-in sample numbers") a = ap.parse_args()
bw = a.bw or (546 if a.demo else detect()["bw_gbs"]) if not bw: raise SystemExit(f"Unknown bandwidth — pass --bw. Apple chips: {APPLE_BW}") rows = DEMO if a.demo else load_bench() if not rows: raise SystemExit("No bench/*.json yet — run quant_bench.sh (or try --demo).")
out = [] for r in rows: ceil = ceiling(bw, r["size_gb"]) out.append({**r, "ceiling": ceil, "efficiency_%": 100 * r["tg128"] / ceil}) record("04-quant", {"bw_gbs": bw, **out[-1]})
print(f"memory bandwidth: {bw} GB/s\n") print(table(out)) print("\nRead it like this:") print(" • tg128 (decode) scales with 1/size — halve the bytes, nearly double the speed.") print(" • pp512 (prefill) barely moves — prefill is compute-bound, not bandwidth-bound.") print(" • Now run quality_check.py per quant: speed without a quality check is meaningless.")
if __name__ == "__main__": main()quality_check.py Python · 57 lines
"""Lesson 04 · step 3 — a 12-question sanity check per quantization level.
Not a benchmark: a smoke alarm. If Q3 gets 7/12 where Q8 gets 11/12, you have foundthe cliff. For a customer you would swap these questions for 50 of THEIR prompts(lesson 09 builds that harness properly).
Serve one quant at a time, then run: bash ../02-three-local-servers/serve_llamacpp.sh Q3_K_M python quality_check.py --target llamacpp --label Q3_K_M"""import argparse
from felab import add_target_args, banner, client, record, resolve
# (question, substring that must appear in a correct answer). Mixed: arithmetic,# facts, multi-step reasoning, format-following — low-bit quants fail the last two first.QA = [ ("What is 17 * 23? Answer with the number only.", "391"), ("What is the capital of Australia? One word.", "Canberra"), ("If a train leaves at 14:40 and the trip takes 95 minutes, when does it arrive? HH:MM only.", "16:15"), ("Spell 'bandwidth' backwards. Letters only.", "htdiwdnab"), ("How many bytes are in one FP16 number? Digit only.", "2"), ("Return exactly this JSON and nothing else: {\"ok\": true}", "\"ok\": true"), ("What is 2 to the power of 10? Number only.", "1024"), ("Which planet is known as the Red Planet? One word.", "Mars"), ("Alice is taller than Bob. Bob is taller than Carol. Who is shortest? One word.", "Carol"), ("Convert 3.5 GB to MB using 1 GB = 1000 MB. Number only.", "3500"), ("What is the chemical symbol for sodium? Symbol only.", "Na"), ("List the first three prime numbers separated by commas, nothing else.", "2, 3, 5"),]
def main() -> None: ap = add_target_args(argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)) ap.add_argument("--label", default="?", help="what is being served, e.g. Q4_K_M") a = ap.parse_args() t = resolve(a) banner(t) cli = client(t)
score = 0 for q, want in QA: r = cli.chat.completions.create(model=t.model, temperature=0, max_tokens=40, messages=[{"role": "user", "content": q}]) got = (r.choices[0].message.content or "").strip() ok = want.replace(" ", "").lower() in got.replace(" ", "").lower() score += ok print(f" {'✓' if ok else '✗'} {q[:60]:60s} → {got[:40]!r}")
print(f"\n{a.label}: {score}/{len(QA)}") record("04-quality", {"target": t.name, "label": a.label, "score": score, "of": len(QA)})
if __name__ == "__main__": main()quant_bench.sh Bash · 32 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 04 · step 1 — the same model at three precisions, benchmarked raw.## Q8_0 ~8.5 bits/weight ≈ 8.5 GB for 8B near-lossless# Q4_K_M ~4.8 bits/weight ≈ 4.9 GB for 8B the default everybody ships# Q3_K_M ~3.9 bits/weight ≈ 4.0 GB for 8B where quality starts to slip## llama-bench runs WITHOUT a server: pure engine speed.# -p 512 prefill 512 tokens → "pp512" tok/s (compute-bound)# -n 128 generate 128 tokens → "tg128" tok/s (bandwidth-bound) ← the one we check## Downloads ~17 GB in total. Delete models/ afterwards if disk is tight.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"REPO=${LLAMACPP_REPO:-bartowski/Meta-Llama-3.1-8B-Instruct-GGUF}QUANTS=(${QUANTS:-Q8_0 Q4_K_M Q3_K_M})mkdir -p models bench
command -v hf >/dev/null || pip install -q "huggingface_hub[cli]"
for q in "${QUANTS[@]}"; do echo "▸ $q — download (skipped if present)" hf download "$REPO" --include "*${q}.gguf" --local-dir models >/dev/null f=$(ls models/*"${q}".gguf | head -1) echo "▸ $q — llama-bench on $(du -h "$f" | cut -f1)" llama-bench -m "$f" -p 512 -n 128 -ngl 99 -o json > "bench/${q}.json"done
echopython ceiling_check.py