Skip to content

Lessons and code

All 19 lessons, in order. Each one is a step of the course: open it from the Day and Step column, or follow the course from the start.

# day · step folder what you do runs on
00 Day 1 · Step 2 00-setup install Ollama, llama.cpp, MLX; Fireworks guardrails Mac
01 Day 1 · Step 3 01-latency-harness measure TTFT and ITL; the harness every later lesson reuses Mac → FW
02 Day 2 · Step 1 02-three-local-servers one model served by Ollama, llama.cpp and MLX Mac
03 Day 3 · Step 1 03-concurrency-sweep the throughput vs TTFT curve that sizes a deployment Mac
04 Day 4 · Step 1 04-quantization Q8 vs Q4 vs Q3: speed, ceiling maths, quality check Mac
05 Day 4 · Step 2 05-kv-cache KV cache maths, then break and fix long context Mac
06 Day 5 · Step 1 06-prefix-caching shared prefixes: the ~20× TTFT win, and the anti-pattern Mac → FW
07 Day 6 · Step 1 07-structured-output JSON validity: prompt vs json_object vs json_schema Mac → FW
08 Day 6 · Step 2 08-batch-api evals at half price with the Batch API FW
09 Day 7 · Step 1 09-eval-harness quality + p95 + $/1k tasks in one table, LLM judge Mac → FW
10 Day 7 · Step 2 10-lora-fireworks a real LoRA fine-tune: train → deploy → eval → delete FW $
11 Day 8 · Step 1 11-lora-local-mlx the same LoRA on your Mac (QLoRA) Mac
12 Day 9 · Step 1 12-moe-vs-dense infer MoE active parameters from decode speed Mac
13 Day 9 · Step 2 13-speculative-decoding draft models, acceptance rate, when it hurts Mac → FW $
14 Day 9 · Step 3 14-spark-track vLLM and SGLang on the DGX Spark, same harness Spark
15 Day 10 · Step 1 15-capstone build the customer sizing memo from your results anywhere
16 Day 11 · Step 1, Day 12 · Step 1 16-training-toy full vs LoRA, SFT → reward model → RLHF vs DPO, in numpy anywhere
17 Day 11 · Step 2 17-full-vs-peft-mlx full vs LoRA vs DoRA on a real model; data checker Mac
18 Day 12 · Step 2 18-preference-dpo DPO on your Mac (and optionally Fireworks) Mac → FW

Every lesson folder has a README.md with the same parts: what and why, the code to read first, commands to run, expected output, self-check questions and a short explanation of the result.