Labs overview
Hands-on inference engineering, from a MacBook to a DGX Spark to Fireworks AI. Nineteen short lessons: sixteen that take you from “what is TTFT?” to a customer sizing memo built entirely from numbers you measured yourself, then three on how models are trained: full fine-tuning, SFT, LoRA and its family, RLHF and DPO.
Built to practise the work of an AI Engineer or Field Engineer. Every lesson ends with a short explanation you could give to a customer or teammate.
┌──────────────────┐ ┌────────────────────────┐ ┌─────────────────────────┐ │ 1. WATCH │ ──▶ │ 2. READ │ ──▶ │ 3. RUN │ │ 18 narrated │ │ Lab book web page │ │ this repo │ │ explainer videos│ │ concepts, calculators,│ │ lessons/NN-*/ │ │ (3–4 min each) │ │ cards F1–7 L1–9 T1–3 │ │ README → code → results│ └──────────────────┘ └────────────────────────┘ └─────────────────────────┘ the WHY the WHAT the HOW (your numbers)The lab book is the map. Open lab-book.html in a browser (one
offline file) or use the hosted copy. It explains each concept, has the calculators, and
each lab card names its folder here. This repo is the territory: commented code you
run in order, which writes its measurements to results/. Keep them side by side, with the
page on the left and the terminal on the right.
Got this inside the Field Engineer Kit? The videos are in
../01-videos/(numbered01-…to18-…, the same numbers each lesson’s “Watch:” line uses) and the lab book is../02-lab-book/field-engineer-lab-book.html. Start with../README.md.
Quick start (10 minutes, no model download)
Section titled “Quick start (10 minutes, no model download)”git clone <this repo> field-engineer-labs && cd field-engineer-labspython3 -m venv .venv && source .venv/bin/activatepip install -r requirements.txt && pip install -e . # Linux: drop mlx-lm from requirementscp .env.example .env
make mock & # a fake LLM server with realistic latency behaviour, on :9000make smoke # runs every lesson offline, about 2 minutesOnce the smoke test passes, run bash lessons/00-setup/setup_mac.sh for the real local
stack and start at lesson 00.
The lessons
Section titled “The lessons”Each folder has a README.md (what, why, run, expected output, self-check, concise
explanation) and numbered or commented scripts. Work through them in order, because later
lessons reuse earlier tools.
| # | lesson | lab book | video | runs on | time | cost |
|---|---|---|---|---|---|---|
| 00 | Setup and guardrails | L1 | serving-stack | Mac | 15 m | free |
| 01 | Latency harness: TTFT and ITL | F1 | inference-101 | Mac → FW | 30 m | ~$0.02 |
| 02 | Three local servers | L2 | serving-stack | Mac | 30 m | free |
| 03 | Concurrency sweep | L3 | benchmarking, batching | Mac | 45 m | free |
| 04 | Quantization | L4 | quantization, gpu-bandwidth | Mac | 45 m | free |
| 05 | KV cache and context | L5 | kv-cache | Mac | 45 m | free |
| 06 | Prefix caching | L6 + F2 | prefix-caching | Mac → FW | 30 m | ~$0.03 |
| 07 | Structured output | F3 | platform | Mac → FW | 30 m | ~$0.02 |
| 08 | Batch API | F4 | platform | FW | 20 m | <$0.01 |
| 09 | Eval harness | F6 | training | Mac → FW | 1 h | ~$0.10 |
| 10 | LoRA on Fireworks | F5 | training, qlora | FW | 2 h | ~1 GPU-hour |
| 11 | LoRA on the Mac (QLoRA) | L8 | qlora | Mac | 1–2 h | free |
| 12 | MoE vs dense | L7 | offloading | Mac | 30 m | free |
| 13 | Speculative decoding | F7 | scale-out, batching | Mac → FW | 45 m | ~½ GPU-hour |
| 14 | The Spark track | L9 | serving-stack, scale-out | Spark | 2–3 h | free |
| 15 | Capstone: sizing memo | Capstone | platform | Mac | 2 h | free |
| Training track (after the capstone, or alongside lessons 09–11) | ||||||
| 16 | Training methods in miniature | T1 | full-finetuning, peft, rlhf, dpo | anywhere | 30 m | free |
| 17 | Full vs LoRA vs DoRA on your Mac | T2 | full-finetuning, sft, peft | Mac | 1 h | free |
| 18 | Preference tuning with DPO | T3 | rlhf, dpo | Mac → FW | 1 h | free / optional |
The 10-day plan in the lab book maps onto these lessons: day 1 → 00–01, day 2 → 02, day 3 → 03, day 4 → 04–05, day 5 → 06, day 6 → 07–08, day 7 → 09–10, day 8 → 11, day 9 → 12–14, day 10 → 15. The training track (16–18, videos 14–18) adds two more days: day 11 → 16–17, day 12 → 18.
One harness, every backend
Section titled “One harness, every backend”Every script takes --target, so the same code measures every server and the numbers
stay comparable:
--target |
what | start it with |
|---|---|---|
mock |
offline fake LLM (latency, batching, prefix cache, JSON behaviour) | make mock |
ollama |
Ollama, :11434 | ollama serve |
llamacpp |
llama.cpp llama-server, :8080 |
lessons/02-*/serve_llamacpp.sh |
mlx |
mlx_lm.server, :8081 |
lessons/02-*/serve_mlx.sh |
fireworks |
Fireworks serverless or your deployment (paid) | FIREWORKS_API_KEY in .env |
spark |
vLLM or SGLang on a DGX Spark | lessons/14-*/1_vllm.sh |
Override with --model and --base-url, or set defaults in .env.
Repository layout
Section titled “Repository layout”felab/ shared toolkit every lesson imports targets.py the --target table; loads .env measure.py stream_once(): TTFT and ITL per token; percentiles results.py record() → results/<lesson>.csv; table() tickets.py dataset loader and graders (parse / schema-valid / correct) hardware.py chip, RAM and bandwidth detection; the decode-ceiling formula mock_server.py offline OpenAI-compatible server that behaves like the real thingdata/ make_tickets.py builds the shared triage dataset (800 train / 100 valid / 100 test; de-duplicated, disjoint splits, 30 test tickets in unseen phrasings) triage_schema.json the JSON contract used in lessons 07–11lessons/NN-*/ README.md + commented scripts, in course orderresults/ your measurements (git-ignored), read by the capstonescripts/smoke_test.sh every offline lesson in about 2 minutes (also runs in CI)Spending on Fireworks: the rules
Section titled “Spending on Fireworks: the rules”The whole course costs well under $15 if you follow these rules:
- Serverless is per token and costs cents. Dedicated deployments bill per GPU-hour from the moment they are created, whether or not you send a request.
- Only lessons 10, 13 and (optionally) 18 create deployments. Lesson 10 ends with
5_teardown.sh, and lesson 13 deletes its deployment automatically on exit, even after Ctrl-C. - Set a budget or alert in the Fireworks console before you start. Run
make fw-checkat the end of every session: it lists anything still billing. - Rehearse every paid lesson against
--target mockfirst.
Troubleshooting
Section titled “Troubleshooting”| symptom | fix |
|---|---|
Connection refused |
That server isn’t running. make check shows what is up. |
FIREWORKS_API_KEY is not set |
Put it in .env (never commit it). |
| model id not found on Fireworks | Ids change. Copy a current one from the model library and pass --model. |
| llama.cpp refuses a quantized V cache | Add flash attention: EXTRA="-fa on" bash context_ladder.sh |
| MLX out of memory while training | Use --batch-size 1, lower --num-layers, or a smaller base model. |
| a container tag isn’t found on the Spark | Tags move. See the links in lessons/14-*/1_vllm.sh and 3_sglang.sh. |
Commands and flags follow the vendors’ docs as of September 2026. When a flag has moved,
<tool> --help is the source of truth.
Licence
Section titled “Licence”MIT. Model weights you download keep their own licences.