Skip to content

Troubleshooting

The kit’s own list first, then the fixes from each step of the course.

symptom fix
make smoke says connection refused the mock isn’t running: make mock & first
a lesson says Connection refused that server isn’t running; make check shows what is up
a Fireworks model id is not found ids change; copy a current one from the model library and pass --model
page shows odd symbols open the .html in a browser, not a text editor
unsure whether you’re still paying make fw-check lists anything still billing

Day 1 · Step 1 · Get the code, run it offline

Section titled “Day 1 · Step 1 · Get the code, run it offline”
make smoke stops partway with an error.
Scroll up to the last heading it printed (for example ‘05 kv’): the error under it names the problem, and the rows below cover the usual ones. make smoke starts the practice server itself when none is running. If make check then shows mock down, start it with make mock & and read the error it prints.
‘No module named felab’, ‘No module named openai’ or ‘python: command not found’.
The virtual environment is off, or the install did not finish. From the 03-labs folder run source .venv/bin/activate, then pip install -r requirements.txt && pip install -e . again. Every new Terminal window needs the source line.
pip says the package ‘requires a different Python’.
Your python3 is older than 3.10, the minimum the code needs. Install a current one with brew install python, open a new Terminal window, delete the old environment with rm -rf .venv, and repeat part 2 on this page (Create a virtual environment and install).
‘Address already in use’ when you start the practice server.
One is already running, from earlier or in another window, and it keeps serving. Check with make check: if mock says UP, carry on.
make is not found, or a window asks to install command line developer tools.
Accept the install: Apple’s command line tools include make. Or run xcode-select --install, then try again.

Go to the step

Day 1 · Step 2 · Set up your Mac and the spend cap

Section titled “Day 1 · Step 2 · Set up your Mac and the spend cap”
setup_mac.sh stops with ‘Install Homebrew first’.
Install Homebrew from brew.sh, run the ‘Next steps’ lines it prints at the end, then run the script again. It skips anything already installed.
check_env.py says ‘bandwidth unknown (not an Apple chip)’ on an Apple Mac.
Your chip is newer than the kit’s list (felab/hardware.py knows M1 to M4 and the base M5). Look up your chip’s memory bandwidth on Apple’s tech specs page and work out the limit yourself: bandwidth ÷ 4.9.
ollama pull says it could not connect.
The Ollama server is not running yet. Run ollama serve &, wait two seconds, then pull again.
fireworks_guardrails.sh says ‘firectl is not installed’.
Install it with brew tap fw-ai/firectl && brew install firectl, then run firectl signin, and run the script again.
firectl whoami fails, or the script says ‘Run: firectl signin’.
Run firectl signin (it opens a browser to log in), then run the script again.
A firectl command rejects its options.
Commands change. Run it with --help (for example firectl quota update --help) and follow what it shows. The Fireworks console’s billing page also has usage limits.

Go to the step

The script prints ‘failed — check the id in the model library’ for gpt-oss-20b.
Expected, nothing is wrong with your setup. On 27 September 2026 gpt-oss-20b’s model page said “Serverless: Not supported”: Fireworks no longer runs it per token, so the call fails at once. The script carries on, and your other Fireworks rows still count. Optional: replace the gpt-oss-20b line in the MODELS list at the top of your copy of compare_models.sh with another model the model library (app.fireworks.ai/models) marks as serverless, and run it again.
Every model says ‘failed — check the id in the model library’, and the error above mentions 401 or an invalid API key.
.env still holds the placeholder key fw_xxxxxxxxxxxxxxxx. From the lesson folder run open -e ../../.env, replace the placeholder after FIREWORKS_API_KEY= with your key from app.fireworks.ai (API keys page), save, and run bash compare_models.sh again.
‘FIREWORKS_API_KEY is not set’.
The FIREWORKS_API_KEY= line is missing from .env, or has nothing after the =. From the lesson folder run open -e ../../.env, add the line with your key after the = (create the key at app.fireworks.ai, API keys page), and save. Never commit or share that file.
The script prints ‘failed — check the id in the model library’ for another model too.
Model names change. Copy a current id from the Fireworks model library (app.fireworks.ai/models) into the list at the top of compare_models.sh, or pass it to bench_ttft.py with --model. If all three fail and the error mentions 401, it is the key, not the ids: see the first row.
Fireworks errors with code 429 (too many requests).
Until a payment method is added, the account allows 10 requests a minute. compare_models.sh sends about 17: 6 per model that answers (5 timed plus a warm-up), 1 for gpt-oss-20b’s failed call and 4 for the long prompt (6 + 1 + 6 + 4). Running it again hits the same limit. Add a small prepaid credit (spend-cap checklist item 5), or time one model at a time, a minute apart, for example python bench_ttft.py --target fireworks --model accounts/fireworks/models/gpt-oss-120b --runs 5 (6 requests, counting the warm-up).
During the long-prompt Ollama run, Ollama’s log (in the window where it runs) says ‘truncating input prompt’.
Ollama cut the prompt to its context setting (the most tokens it accepts), so the row timed a shorter prompt. Stop it with pkill ollama, start it with room for the long prompt, OLLAMA_CONTEXT_LENGTH=16384 ollama serve &, and run the long-prompt line again.
‘Connection refused’.
That server is not running. From the 03-labs folder, make check shows which are up. Start the practice server with make mock &, and Ollama with ollama serve &.
make fw-check says ‘No rule to make target’.
You are still in the lesson folder. Run cd ../.. to get back to 03-labs, then run it again.

Go to the step

Day 2 · Step 1 · Serve one model three ways

Section titled “Day 2 · Step 1 · Serve one model three ways”
The comparison says ollama is not running (or llamacpp, or mlx) and skips it.
That server is not up. Look in its terminal for an error and start it again. From the 03-labs folder, with the Python environment on, make check lists which servers answer.
All three are skipped although the servers are running, and the address after ‘at’ is blank.
The script could not load the kit’s Python toolkit. In that terminal, go to the 03-labs folder, run source .venv/bin/activate, then bash lessons/02-three-local-servers/compare_backends.sh (the script moves into its own folder by itself).
llama-server or mlx_lm.server: command not found.
For mlx_lm.server, the Python environment is usually off in that terminal: from the 03-labs folder run source .venv/bin/activate and try again. If a command is still missing, rerun Day 1’s setup script, which is safe to run again: bash lessons/00-setup/setup_mac.sh, then source .venv/bin/activate.
ollama serve says the address is already in use.
Ollama is already running, probably still from Day 1. Leave it running and skip terminal 1.
The mlx row fails with an error about the model.
The model name the script sends must match the one the MLX server loaded. The MLX server prints the exact line to use: run that export MLX_MODEL=... line in terminal 4, then compare again.
The lab book card’s BASE_URL=... python bench_ttft.py --model local lines measure the practice server, or fail to connect, instead of your Mac’s servers.
Those lines belong to the lab book’s own short script. Day 1’s script picks a server with --target instead (--target llamacpp, mlx or ollama), or run bash compare_backends.sh.

Go to the step

The first sweep stops with Connection refused (or Connection error).
The practice server is not running. From the 03-labs folder, with the Python environment on, run make mock &, then run the sweep again. make check shows which servers are up.
The llama.cpp sweep stops with Connection refused (or Connection error).
llama.cpp is not running yet, or is still loading the model. Start it in its own terminal with the serve_llamacpp.sh line above and wait until it says it is listening. make check shows llamacpp as UP when it is ready.
serve_llamacpp.sh says No such file or directory.
The new terminal is in the wrong folder. cd into the kit’s 03-labs/lessons/03-concurrency-sweep folder and run the line again.
Your Mac’s chart jumps after 4 people, not 8.
llama.cpp started with Day 2’s 4 seats. Its first line must say slots=8: stop it with Ctrl+C and start it again with Q4_K_M 16384 8 at the end.
plot_sweep.py prints a text chart, then open says 03-sweep.png does not exist.
The chart library (matplotlib) is missing, so the script falls back to text. In the 03-labs folder run source .venv/bin/activate, then pip install -r requirements.txt, then plot again.
plot_sweep.py says No results yet.
Run a sweep first. Every sweep adds its rows to 03-labs/results/03-sweep.csv, wherever you run it from, and the chart reads them from there.

Go to the step

The download stops because the disk is full.
The three sizes need about 17 GB, and the quality loop later stores the 8-bit and 3-bit versions again (about 12.5 GB). Free some space and run bash quant_bench.sh again: sizes already downloaded are skipped. Once its table has printed, you can delete the models folder inside lessons/04-quantization. If the disk filled during the quality loop instead, delete that folder now (the speed test no longer needs it) and run the loop again.
It stops with Unknown bandwidth — pass --bw.
The script’s list of Apple chips does not include yours. The error lists the chips it knows with their speeds; your Mac’s tech specs page on Apple’s website lists its memory bandwidth. Then run python ceiling_check.py --bw 273, with your figure in place of 273. The timings are already saved, so nothing downloads again.
Efficiency is well below 60%.
Something else is slowing the Mac: another app using memory heavily, a hot Mac slowing itself down, or part of the model running on the Mac’s main processor instead of its graphics cores. Close big apps (video calls, browser tabs playing video), plug in, and run bash quant_bench.sh again; downloads are skipped if you have not deleted models/. Some chips, such as the M3 Max and M4 Max, come in two versions, and the script assumes the faster one (felab/hardware.py). If yours is slower, run python ceiling_check.py --bw followed by your figure in GB per second.
The quality loop prints a connection error (for example Connection refused) for one size.
The server was still downloading that size when the questions started: the loop waits only 20 seconds. Start that size alone, for example bash ../02-three-local-servers/serve_llamacpp.sh Q8_0, wait for the line saying it is listening, stop it with Ctrl+C, then run the loop again.
llama-bench: command not found
llama.cpp is missing. Day 1’s setup installs it; to add it again, run brew install llama.cpp, then bash quant_bench.sh.

Go to the step

Day 4 · Step 2 · KV cache and context length

Section titled “Day 4 · Step 2 · KV cache and context length”
Every 8-bit (q8_0) row fails, even the short 4,096-token one.
Turn on the server’s flash attention setting. (Flash attention is a faster way to run the look-back step; llama.cpp needs it to store the notes in 8 bits.) Run EXTRA="-fa on" bash context_ladder.sh. On older llama.cpp versions use EXTRA="-fa" bash context_ladder.sh.
The 131,072-token 16-bit row says failed to start.
That is the lesson. The notes (17.18 GB) plus the 4.9 GB model need about 22 GB, but a Mac gives the model only about 70% of its memory (Day 1’s rule): about 17 GB on a 24 GB Mac. Write “failed” in Your numbers; the 8-bit row below it is the fix.
Both 131,072-token rows fail (a 16 GB Mac).
That can happen. The 9.13 GB of 8-bit notes plus the 4.9 GB model is about 14 GB, more than the about 11 GB a 16 GB Mac gives the model (70%, Day 1). Write “failed” for both; the 32,768-token pair shows the same halving.
Your rows are exactly half the lesson’s, for example 2,048 and 1,088 at 32,768 tokens.
llama.cpp printed the notes’ two halves (keys and values) on the same line, and the script copied only the last number. In lessons/05-kv-cache, run grep -i "kv buffer size" logs/ctx32768_f16.log, using the log named after the row (its ctx, then f16 or q8_0). The MiB number on that line is the full figure.

Go to the step

Connection refused.
That server is not running. make check (from the 03-labs folder) shows what is up. Start the practice server with make mock &, or start llama.cpp as in the run notes and wait until it says it is listening.
Opening first on the practice server shows a speed-up near 1, and call 1 is already fast.
The practice server still remembers the opening from an earlier run: this command run twice, or make smoke while it was up. Restart it: pkill -f felab.mock_server, then make mock & from the 03-labs folder, and run again.
No such file or directory when you start llama.cpp.
The command’s path starts from this lesson’s folder. From the 03-labs folder run cd lessons/06-prefix-caching, then start it again.
llama.cpp returns an error about the context size, or its opening-first run shows no speed-up.
Each of its 4 seats holds only 2,048 tokens by default, less than the 3,000-token opening. Stop it (Ctrl+C) and, from this lesson’s folder, start it with bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 4 (4,096 tokens per seat).
FIREWORKS_API_KEY is not set.
Put your key in the .env file in 03-labs, as on Day 1, or export it in this terminal.
cached_tokens=0 or cached_tokens=None.
Run the Fireworks command again right away: caches expire when idle (the Prefix Caching video). If it still says None, this model does not report it; pass another with --model.
A Fireworks model id is not found.
Model ids change. Copy a current one from the Fireworks model library and pass it with --model.
mlx_lm.generate: command not found.
The Python environment from Day 1 is not active. From the 03-labs folder run source .venv/bin/activate, then try again.

Go to the step

Every mode prints server rejected request (APIConnectionError: Connection error.) and the table says (no rows).
That server is not running. make check (from the 03-labs folder) shows what is up. Start the practice server with make mock &, or llama.cpp with bash lessons/02-three-local-servers/serve_llamacpp.sh in a second terminal, and wait until it says it is listening.
data/tickets/test.jsonl missing.
The ticket data is not built yet. From the 03-labs folder run python data/make_tickets.py, then run the script again.
No module named 'felab' (or 'openai').
The Python environment from Day 1 is off in this terminal. From the 03-labs folder run source .venv/bin/activate, then try again.
One mode prints server rejected request with an error such as BadRequestError, while the other modes print rows.
That server does not accept that way of asking, so the script skips the mode and carries on. Write it down: which modes a server supports is a real finding for a customer.
FIREWORKS_API_KEY is not set, or every mode prints AuthenticationError on Fireworks.
Your key is missing or wrong. Put your key in the .env file in 03-labs, as on Day 1, or export it in this terminal.
Every mode prints server rejected request (NotFoundError: ...) on Fireworks.
The model id is not found. Model ids change: copy a current one from the Fireworks model library and pass it with --model.
Fireworks shows a low schema_valid_% or empty replies, even with the format enforced.
gpt-oss-120b writes out its thinking before its answer (a reasoning model, check 3). That thinking may use up the script’s 120-token reply limit, so the answer comes back cut off or empty. Write it down as a finding and compare with the prompt row, which follows check 3’s pattern: schema in the prompt, nothing enforced, checked afterwards.

Go to the step

Connection refused on the practice run.
The practice server is not running. From the 03-labs folder run make mock &, then run the command again.
firectl: command not found.
Install it as on Day 1: brew tap fw-ai/firectl && brew install firectl, then firectl signin.
A firectl command rejects its options.
Commands change. Run it with --help, for example firectl batch-inference-job create --help, and follow what it says; the script’s own note says the same.
The job is still running when you have to stop.
Nothing is lost: Ctrl+C stops the checking, not the job. Later, from lessons/08-batch-api, run firectl batch-inference-job get <job name> with the name from the create job: line. When it says COMPLETED, run the last two commands.
The job stays pending and shows no error.
Fireworks’ batch guide (checked 27 September 2026) warns that a job on a model batch does not support can stay pending with no error. Batch should still take gpt-oss-20b, but if yours has not started by the time you finish today’s other work, start a second job on gpt-oss-120b, which is offered per token: from lessons/08-batch-api, run bash submit_batch.sh accounts/fireworks/models/gpt-oss-120b. Write down its new job name and follow that one. Your score is then gpt-oss-120b’s, so label it that way.
The job ends FAILED or EXPIRED.
Read the job details the script printed at the end for the reason. Model ids change: if the model is no longer offered, copy a current one from the Fireworks model library and run bash submit_batch.sh <that id>.
answered is below 100.
Some requests failed and went to the job’s error file instead of the results. The score covers the ones answered, each matched to its ticket by custom_id: that is exactly why the label exists.
Pass a results file, or --simulate.
The path after score_batch.py does not point to a file. Use the exact name of the downloaded .jsonl file; ls lists what is in the folder.

Go to the step

Connection refused on the practice run.
The practice server is not running. From 03-labs, run make mock &, then run the command again.
Connection refused on the ollama mlx run.
One of the two servers is not running: make check (from 03-labs) shows which. Start Ollama with ollama serve, and MLX in another terminal with bash lessons/02-three-local-servers/serve_mlx.sh.
FIREWORKS_API_KEY is not set
Your key belongs in the .env file in 03-labs, as on Day 1: a line FIREWORKS_API_KEY= followed by the key. Then run the command again.
The kit’s gpt-oss-20b line fails, or says the model is not found.
Expected: gpt-oss-20b is no longer offered per token (“Serverless: Not supported” on its model page, checked 27 September 2026). Run the gpt-oss-120b line from the code block instead: python evaluate.py "fireworks:accounts/fireworks/models/gpt-oss-120b@0.15/0.60" --n 50 --judge fireworks.
A Fireworks model id is not found, even on the gpt-oss-120b line.
Model ids change. Copy a current one from the model library on fireworks.ai that is offered serverless, put it in place of accounts/fireworks/models/gpt-oss-120b, and copy its current prices after the @ too.
The Fireworks row shows a low schema_valid_%, or empty replies.
gpt-oss-120b writes out its thinking before its answer (a reasoning model, as on Day 6). evaluate.py allows each reply 120 tokens (max_tokens), and the thinking may use them up, so the answer comes back cut off or empty. Write it down as a finding: it is part of that model’s row, not a broken setup.
The --judge fireworks run says a model is not found.
The judge is your default Fireworks model, gpt-oss-120b. Put a current id in 03-labs/.env as FIREWORKS_MODEL=accounts/fireworks/models/<id>, then run the command again.
judge_1to5 says nan.
No judge was asked: only runs with --judge fill that column. The ollama mlx run has no judge, so nan is expected there.

Go to the step

1_train.sh fails because the base model cannot be fine-tuned, or is not found.
Pick a model from Fireworks’ list of tunable models and run BASE_MODEL=accounts/fireworks/models/<id> bash 1_train.sh. Pass the same BASE_MODEL=... to bash 4_eval.sh later: the scripts do not remember it, and would grade the default base instead.
2_wait.sh stops with FAILED or CANCELLED.
It prints the job’s details: read the error there. Nothing bills by the hour yet. Once it is fixed, run bash 1_train.sh again: it starts a new job with new names.
3_deploy.sh keeps printing dots and never reaches ready.
The deployment already exists, so it may be billing. Press Ctrl-C and check the deployment’s state in the Fireworks console. If it shows ready, carry on with bash 4_eval.sh. If it is not coming up, run bash 5_teardown.sh, then make fw-check from 03-labs.
4_eval.sh stops with an error, such as a model not found.
Your GPU is still billing: if you cannot fix this in a few minutes, run bash 5_teardown.sh first and come back. If the error names your fine-tune, copy the model string from the deployment’s API tab in the Fireworks console and run python ../09-eval-harness/evaluate.py "fireworks:<that string>" --n 60. If it names the base model, its id may have changed: copy a current one from the model library and run BASE_MODEL=accounts/fireworks/models/<id> bash 4_eval.sh.
You closed the terminal while the deployment was running.
Go back to lessons/10-lora-fireworks and run bash 5_teardown.sh. The ids are saved in the .state file, and the clock runs until you do.
5_teardown.sh stops with an error, or you are unsure whether anything is still billing.
From 03-labs, run make fw-check. Delete anything it lists with firectl deployment delete <DEPLOYMENT_ID>, then run make fw-check again until the list is empty.

Go to the step

Day 8 · Step 1 · The same LoRA on your Mac

Section titled “Day 8 · Step 1 · The same LoRA on your Mac”
Training stops with an out-of-memory error, or the Mac slows to a crawl.
Close big apps first. Then in 1_train_mlx.sh change --batch-size 2 to --batch-size 1 (the script’s own advice) and run it again. If it still fails, also change --num-layers 8 to 4: fewer layers with add-ons need less memory, but the add-on can learn less.
mlx_lm.lora: command not found
The Python environment from Day 1 is not on. From the 03-labs folder run source .venv/bin/activate, then cd lessons/11-lora-local-mlx and run the script again.
2_fuse_and_serve.sh fails with ‘Address already in use’.
Another server already uses port 8081, most likely Day 2’s MLX server. Stop it with Ctrl+C in its window, then run the script again.
3_eval_local.sh says Connection refused.
The server is not up yet: fusing comes first and takes a while. Wait for the line ending serving on :8081 (Ctrl-C to stop) in the server window, give it a few seconds to load, then run the eval again. Keep that window open.
The eval gets through the fine-tune, then stops with an error at evaluating mlx mlx-community/Qwen2.5-7B-Instruct-4bit.
Your copy of the MLX tools will not switch models on one server. No table prints, but the fine-tune’s scores are already saved: the last row of 03-labs/results/09-eval.csv, named mlx fused. Copy its category_%, severity_% and p95_ms from there. Stop the server with Ctrl+C in its window to free memory. In that window, start the original model on a second port: mlx_lm.server --model mlx-community/Qwen2.5-7B-Instruct-4bit --port 8082. Then in the eval window (still in the lesson folder) run MLX_URL=http://localhost:8082/v1 python ../09-eval-harness/evaluate.py "mlx:mlx-community/Qwen2.5-7B-Instruct-4bit" --n 60. It prints the original model’s row.
Validation loss starts rising again before step 300.
The model has begun memorising the training tickets. Stop with Ctrl+C; the last save is kept. For a rerun, lower --iters 300 in 1_train_mlx.sh to about the step where the loss was lowest.

Go to the step

Connection refused.
Ollama is not running. Start it with ollama serve &, then run the command again.
gpt-oss:20b not pulled — run: ollama pull gpt-oss:20b.
Run that pull. The name must match exactly, including :20b.
pass --bw <GB/s>.
The script does not know your chip’s memory speed. Look it up and add it by hand, for example python moe_probe.py llama3.1:8b gpt-oss:20b --bw 546 for a chip that moves 546 GB a second.
gpt-oss:20b is very slow or will not load.
Your Mac is probably short of memory: the model needs its whole 13.8 GB, and Day 1’s rule keeps models under about 70% of memory. Close other apps, or use the --demo numbers.

Go to the step

spec_bench.py cannot connect right after a server start.
The server was still loading. The first speculative start also downloads the draft model, which can take longer than the 40-second wait. The ; kill %1 on that line then stops the server too: start it again (bash serve_spec_llamacpp.sh &), wait until it says it is listening, then run the spec_bench.py line again.
The server will not start because port 8080 is in use.
Another llama.cpp server from Day 2, 4 or 5 is still running. Stop it with Ctrl+C in its terminal, then start again.
kill %1 says there is no such job, or stops the wrong program.
%1 is the first background job of this terminal. Run jobs to list them and use the right number, for example kill %2.
The server refuses the draft model.
The draft must share the big model’s tokenizer. If you changed DRAFT_REPO or LLAMACPP_REPO, set them back, or pick a draft from the same model family.
fireworks_spec.sh stops with an error from firectl.
Check that you are signed in (firectl whoami, Day 1) and that your quota allows a deployment (firectl quota list). Then run make fw-check: it should list nothing.
fireworks_spec.sh never says ready.
Press Ctrl+C: the script deletes the deployment on the way out. Run make fw-check to confirm nothing is left, and try again later.

Go to the step

set SPARK_IP in .env.
Add SPARK_IP=<the Spark's address> to the .env file in 03-labs, plus SPARK_USER=<your login> if it is not nvidia.
SSH asks for a password, or says Permission denied (publickey).
remote.sh needs SSH keys. Set them up once with ssh-copy-id nvidia@<the Spark's address> (with your own login if it differs), then try again.
An image tag is not found.
Container tags change often. Look up the current GB10/arm64 tag (links are in the script comments). Then replace the tag after :- in the IMAGE= line near the top of 1_vllm.sh or 3_sglang.sh, in your copy on the Mac (remote.sh sends that file to the Spark). remote.sh passes only HF_TOKEN, so VLLM_IMAGE or SGLANG_IMAGE set on your Mac never arrives.
A model download fails with an access error.
It may be a gated model, one you must request on Hugging Face first. Request access, then put your token in .env as HF_TOKEN=…; remote.sh passes it to the Spark.
2_bench.sh says python: command not found or No module named 'felab'.
It runs on your Mac. From the 03-labs folder, run source .venv/bin/activate first.
The prefix test cannot connect to http://:30000/v1.
$SPARK_IP was empty in your terminal. Run export SPARK_IP=<the Spark's address>, or paste the command 3_sglang.sh printed.
nvidia-smi shows memory as N/A.
Normal on a Spark: its CPU and GPU share one memory, so the usual fields are blank. Use free -h on the Spark instead.

Go to the step

Day 10 · Step 1 · Sizing memo and 10-minute talk

Section titled “Day 10 · Step 1 · Sizing memo and 10-minute talk”
ModuleNotFoundError: No module named 'felab'
The virtual environment is off in this window. From the 03-labs folder run source .venv/bin/activate, then cd lessons/15-capstone and run the command again.
A section says TODO although you did that day.
The memo reads fixed file names in results/. From lessons/15-capstone, run ls ../../results/ and look for that day’s file (the list is under Before you start on the Day 10 overview). If it is missing, run that day’s main command again; the practice-server version is free.
Replicas, the dedicated cost and the crossover say TODO, but 03-sweep.csv is there.
No crowd size in your sweep kept the memo’s speed promise, by default 95 of 100 first words within 1,000 ms. Once your sweep has any real-server rows, the memo ignores the practice-server rows. From lessons/15-capstone, open ../../results/03-sweep.csv and read the ttft_p95 column (in ms). Pass the customer’s real target with --slo-ms, or run Day 3’s sweep again with smaller crowds, for example --levels 1,2,4.
The memo warns that capacity came from a laptop or mock server.
Expected: your Day 3 sweep ran on your Mac or the practice server. Data-center GPUs running a production engine such as vLLM hold far more people per copy, so treat the copies and the dedicated cost as placeholders and say so on your risks slide. Day 9’s optional DGX Spark step, or a Fireworks deployment, gives real figures.
The fine-tune row names mock mock-8b-lora instead of your own model.
The memo compares every row in results/09-eval.csv, including the practice-server rows Day 1’s smoke test saved. Open that file in TextEdit, delete the lines containing mock mock-8b, keep the first line (the column names), save, and run python build_memo.py again.
The crossover is hundreds of millions of tickets a month.
Not a bug in your data: the memo assumes today’s copies carry any volume. With a handful of people per copy, the dedicated option needs many GPUs, so serverless wins by a wide margin: with the practice server’s 8, the example’s crossover is about 744 million tickets, 372 times its volume. Replace the capacity with a real GPU measurement before you quote it.
open says no application knows how to open the memo.
Your Mac has no app set for .md files. Run open -e ../../results/SIZING_MEMO.md to open it in TextEdit instead.
The talk runs well over 10 minutes.
Five slides in 10 minutes is about 2 minutes each. Keep one number per point, and move tables to a backup slide you show only if asked.

Go to the step

I skipped Day 7, so I have no fine-tune of my own.
Use something you did hit: Day 9’s speculative decoding test also needed a dedicated deployment (30 to 45 minutes of one), or a model name that no longer worked. If you ran into nothing, you can use the docs disagreement in section 1 above, but say you found it by reading, not by running.
Everything worked, so I have nothing to criticise.
Look for what cost you time or money rather than what broke: a step that needed a GPU billed by the hour, a price you had to look up in two places, a limit you only found in a changelog (the provider’s list of recent changes). A small, specific friction beats a big, vague one.
My criticism may be out of date.
The lab book read Fireworks’ pages on 23 September 2026 and warns that prices move weekly. Read the pricing page and the LoRA deployment docs again before sharing your recommendation. If it has been fixed, say so and ask what drove the change: that shows you keep up.
It sounds like a complaint.
Start with what works, with a number, and end with your proposal and a question. Cut “always”, “never” and “broken”; keep your one measured number.

Go to the step

ModuleNotFoundError: No module named 'numpy'
The Python environment is off in this terminal. From the 03-labs folder run source .venv/bin/activate and try again. The toys need nothing else, so on another machine pip install numpy is enough.
Your numbers differ a little from the lesson’s.
The toy draws its random numbers from a fixed starting point (--seed 0), so the shape should match: rank 1 stuck, rank 2 and above matching full at 200 examples. Small differences in the last digits do not matter. Run python toy_finetune.py --seed 1 to see the same shape with other random numbers.

Go to the step

The full run stops with an out-of-memory error, or the Mac slows to a crawl.
Full training peaks at about 8 to 10 GB. Close big apps and try again, or skip it. To skip it, first remove what the failed run left behind in the lesson folder, or compare_variants.py will try to test it: rm -rf logs/full.log adapters/full. Then run VARIANTS="lora dora" bash run_variants.sh. Record that full training did not fit on your Mac: it is a useful measured limit for your comparison.
Training takes too long.
Run ITERS=150 bash run_variants.sh: 150 steps instead of 300, about half the time. Note it next to your numbers.
run_variants.sh stops with fix the data first.
sft_data_check.py found a problem in the training file and printed its type with the first line number. If you edited data/tickets/, rebuild it from the 03-labs folder with python data/make_tickets.py.
compare_variants.py says No logs yet: run bash run_variants.sh first.
It reads the logs/ folder inside lessons/17-full-vs-peft-mlx. Run bash run_variants.sh there first, then python compare_variants.py from the same folder.
compare_variants.py says the model evaluation was skipped.
The accuracy and quiz columns need mlx-lm on an Apple-silicon Mac. Turn the environment on (source .venv/bin/activate from 03-labs) and run it again. The memory, time and file-size columns come from the logs either way.
mlx_lm.lora: command not found
mlx-lm is missing from the environment. From 03-labs, run source .venv/bin/activate, then pip install -r requirements.txt (Day 1’s setup), then bash run_variants.sh again.

Go to the step

ModuleNotFoundError: No module named 'numpy'
The Python environment is off. From the 03-labs folder run source .venv/bin/activate, then run the script again. If it still fails, run pip install numpy.
--beta 0, as on the RLHF video’s command card, prints the same as the default.
The script treats 0 as its default, 0.5, because one of its lines divides by β. For a loose leash use a small number instead, such as python toy_alignment.py --beta 0.02.
Your section 5 table has 8 rows, not 3.
That is expected: the script tries 8 leash settings, from 4 down to 0.01. The lesson shows 3 of them (0.50, 0.10 and 0.02).
At --beta 0.1, 5b still favours the right-category answer, but the sweep says β 0.10 ends on the wrong one.
Expected: the sampled run stops short of the optimum. Compare the two KL numbers on the 5b line: 0.94 reached, 1.81 for the optimum. The script’s closing note that 5b lands on the optimum holds at the default β 0.5 (0.34 and 0.34), not here.

Go to the step

Day 12 · Step 2 · Preference tuning with DPO

Section titled “Day 12 · Step 2 · Preference tuning with DPO”
pref_eval.py --target mock says Connection refused.
The practice server is not running. From the 03-labs folder run make mock &, then try again.
pref_eval.py --target mlx says Connection refused.
The two model servers are not running. dpo_mlx.sh printed their commands at the end. In each new terminal, first go to 03-labs and run source .venv/bin/activate, then cd lessons/18-preference-dpo. Start mlx_lm.server --model <START> --port 8081 in one terminal and the same with --adapter-path adapters/dpo --port 8082 in another. When both are up, run the check again.
You do not know what to type for <START>.
Use the model dpo_mlx.sh printed on its DPO starting from: line, such as the full path to sft_model. It is also saved in the file .state in the lesson folder.
dpo_mlx.sh starts from the instruct model, not sft_model.
It did not find Day 11’s LoRA in lessons/17-full-vs-peft-mlx/adapters/lora. DPO still runs, but on a model that gets only about 40% of categories right (check yourself 1). From 03-labs, run VARIANTS=lora bash lessons/17-full-vs-peft-mlx/run_variants.sh (5 to 15 minutes), then bash lessons/18-preference-dpo/dpo_mlx.sh again, and restart both servers with the new <START> it prints (sft_model).
mlx_lm_lora.train stops with an unknown option or argument error.
The tool’s options change between versions, and the script says its --help wins. Run mlx_lm_lora.train --help, match the options at the top of dpo_mlx.sh to it, and run it again.
category_% dropped after DPO.
Taste has cost correctness. The lesson’s fixes: a larger β (BETA=0.2 bash dpo_mlx.sh), fewer steps (ITERS=100 bash dpo_mlx.sh), or more wrong-category pairs, so correctness is part of what “preferred” means.
The pairwise rows are not in results/18-pref.csv.
They have different columns, so the kit’s recorder saves each one in a new file named results/18-pref-<time>.csv. Copy the numbers from the terminal or from those files.
dpo_fireworks.sh says the base model is not found or not DPO-enabled.
Model ids change. Pick one the Fireworks model library marks as DPO-enabled and run BASE_MODEL=accounts/fireworks/models/<id> bash dpo_fireworks.sh. Afterwards run make fw-check: it should list nothing.

Go to the step