Troubleshooting
The kit’s own list first, then the fixes from each step of the course.
From the kit README
Section titled “From the kit README”| symptom | fix |
|---|---|
make smoke says connection refused |
the mock isn’t running: make mock & first |
a lesson says Connection refused |
that server isn’t running; make check shows what is up |
| a Fireworks model id is not found | ids change; copy a current one from the model library and pass --model |
| page shows odd symbols | open the .html in a browser, not a text editor |
| unsure whether you’re still paying | make fw-check lists anything still billing |
From the course steps
Section titled “From the course steps”Day 1 · Step 1 · Get the code, run it offline
Section titled “Day 1 · Step 1 · Get the code, run it offline”make smokestops partway with an error.- Scroll up to the last heading it printed (for example ‘05 kv’): the error under it names the problem, and the rows below cover the usual ones.
make smokestarts the practice server itself when none is running. Ifmake checkthen shows mock down, start it withmake mock &and read the error it prints. - ‘No module named felab’, ‘No module named openai’ or ‘python: command not found’.
- The virtual environment is off, or the install did not finish. From the
03-labsfolder runsource .venv/bin/activate, thenpip install -r requirements.txt && pip install -e .again. Every new Terminal window needs thesourceline. - pip says the package ‘requires a different Python’.
- Your
python3is older than 3.10, the minimum the code needs. Install a current one withbrew install python, open a new Terminal window, delete the old environment withrm -rf .venv, and repeat part 2 on this page (Create a virtual environment and install). - ‘Address already in use’ when you start the practice server.
- One is already running, from earlier or in another window, and it keeps serving. Check with
make check: if mock says UP, carry on. makeis not found, or a window asks to install command line developer tools.- Accept the install: Apple’s command line tools include
make. Or runxcode-select --install, then try again.
Day 1 · Step 2 · Set up your Mac and the spend cap
Section titled “Day 1 · Step 2 · Set up your Mac and the spend cap”setup_mac.shstops with ‘Install Homebrew first’.- Install Homebrew from brew.sh, run the ‘Next steps’ lines it prints at the end, then run the script again. It skips anything already installed.
check_env.pysays ‘bandwidth unknown (not an Apple chip)’ on an Apple Mac.- Your chip is newer than the kit’s list (felab/hardware.py knows M1 to M4 and the base M5). Look up your chip’s memory bandwidth on Apple’s tech specs page and work out the limit yourself: bandwidth ÷ 4.9.
ollama pullsays it could not connect.- The Ollama server is not running yet. Run
ollama serve &, wait two seconds, then pull again. fireworks_guardrails.shsays ‘firectl is not installed’.- Install it with
brew tap fw-ai/firectl && brew install firectl, then runfirectl signin, and run the script again. firectl whoamifails, or the script says ‘Run: firectl signin’.- Run
firectl signin(it opens a browser to log in), then run the script again. - A firectl command rejects its options.
- Commands change. Run it with
--help(for examplefirectl quota update --help) and follow what it shows. The Fireworks console’s billing page also has usage limits.
Day 1 · Step 3 · Measure TTFT and ITL
Section titled “Day 1 · Step 3 · Measure TTFT and ITL”- The script prints ‘failed — check the id in the model library’ for gpt-oss-20b.
- Expected, nothing is wrong with your setup. On 27 September 2026 gpt-oss-20b’s model page said “Serverless: Not supported”: Fireworks no longer runs it per token, so the call fails at once. The script carries on, and your other Fireworks rows still count. Optional: replace the gpt-oss-20b line in the
MODELSlist at the top of your copy ofcompare_models.shwith another model the model library (app.fireworks.ai/models) marks as serverless, and run it again. - Every model says ‘failed — check the id in the model library’, and the error above mentions 401 or an invalid API key.
.envstill holds the placeholder keyfw_xxxxxxxxxxxxxxxx. From the lesson folder runopen -e ../../.env, replace the placeholder afterFIREWORKS_API_KEY=with your key from app.fireworks.ai (API keys page), save, and runbash compare_models.shagain.- ‘FIREWORKS_API_KEY is not set’.
- The
FIREWORKS_API_KEY=line is missing from.env, or has nothing after the=. From the lesson folder runopen -e ../../.env, add the line with your key after the=(create the key at app.fireworks.ai, API keys page), and save. Never commit or share that file. - The script prints ‘failed — check the id in the model library’ for another model too.
- Model names change. Copy a current id from the Fireworks model library (app.fireworks.ai/models) into the list at the top of
compare_models.sh, or pass it tobench_ttft.pywith--model. If all three fail and the error mentions 401, it is the key, not the ids: see the first row. - Fireworks errors with code 429 (too many requests).
- Until a payment method is added, the account allows 10 requests a minute.
compare_models.shsends about 17: 6 per model that answers (5 timed plus a warm-up), 1 for gpt-oss-20b’s failed call and 4 for the long prompt (6 + 1 + 6 + 4). Running it again hits the same limit. Add a small prepaid credit (spend-cap checklist item 5), or time one model at a time, a minute apart, for examplepython bench_ttft.py --target fireworks --model accounts/fireworks/models/gpt-oss-120b --runs 5(6 requests, counting the warm-up). - During the long-prompt Ollama run, Ollama’s log (in the window where it runs) says ‘truncating input prompt’.
- Ollama cut the prompt to its context setting (the most tokens it accepts), so the row timed a shorter prompt. Stop it with
pkill ollama, start it with room for the long prompt,OLLAMA_CONTEXT_LENGTH=16384 ollama serve &, and run the long-prompt line again. - ‘Connection refused’.
- That server is not running. From the
03-labsfolder,make checkshows which are up. Start the practice server withmake mock &, and Ollama withollama serve &. make fw-checksays ‘No rule to make target’.- You are still in the lesson folder. Run
cd ../..to get back to03-labs, then run it again.
Day 2 · Step 1 · Serve one model three ways
Section titled “Day 2 · Step 1 · Serve one model three ways”- The comparison says
ollama is not running(or llamacpp, or mlx) and skips it. - That server is not up. Look in its terminal for an error and start it again. From the
03-labsfolder, with the Python environment on,make checklists which servers answer. - All three are skipped although the servers are running, and the address after ‘at’ is blank.
- The script could not load the kit’s Python toolkit. In that terminal, go to the
03-labsfolder, runsource .venv/bin/activate, thenbash lessons/02-three-local-servers/compare_backends.sh(the script moves into its own folder by itself). llama-serverormlx_lm.server: command not found.- For
mlx_lm.server, the Python environment is usually off in that terminal: from the03-labsfolder runsource .venv/bin/activateand try again. If a command is still missing, rerun Day 1’s setup script, which is safe to run again:bash lessons/00-setup/setup_mac.sh, thensource .venv/bin/activate. ollama servesays the address is already in use.- Ollama is already running, probably still from Day 1. Leave it running and skip terminal 1.
- The mlx row fails with an error about the model.
- The model name the script sends must match the one the MLX server loaded. The MLX server prints the exact line to use: run that
export MLX_MODEL=...line in terminal 4, then compare again. - The lab book card’s
BASE_URL=... python bench_ttft.py --model locallines measure the practice server, or fail to connect, instead of your Mac’s servers. - Those lines belong to the lab book’s own short script. Day 1’s script picks a server with
--targetinstead (--target llamacpp,mlxorollama), or runbash compare_backends.sh.
Day 3 · Step 1 · Concurrency sweep
Section titled “Day 3 · Step 1 · Concurrency sweep”- The first sweep stops with
Connection refused(orConnection error). - The practice server is not running. From the
03-labsfolder, with the Python environment on, runmake mock &, then run the sweep again.make checkshows which servers are up. - The llama.cpp sweep stops with
Connection refused(orConnection error). - llama.cpp is not running yet, or is still loading the model. Start it in its own terminal with the
serve_llamacpp.shline above and wait until it says it is listening.make checkshowsllamacppas UP when it is ready. serve_llamacpp.shsaysNo such file or directory.- The new terminal is in the wrong folder.
cdinto the kit’s03-labs/lessons/03-concurrency-sweepfolder and run the line again. - Your Mac’s chart jumps after 4 people, not 8.
- llama.cpp started with Day 2’s 4 seats. Its first line must say
slots=8: stop it with Ctrl+C and start it again withQ4_K_M 16384 8at the end. plot_sweep.pyprints a text chart, thenopensays03-sweep.pngdoes not exist.- The chart library (matplotlib) is missing, so the script falls back to text. In the
03-labsfolder runsource .venv/bin/activate, thenpip install -r requirements.txt, then plot again. plot_sweep.pysaysNo results yet.- Run a sweep first. Every sweep adds its rows to
03-labs/results/03-sweep.csv, wherever you run it from, and the chart reads them from there.
Day 4 · Step 1 · Quantization
Section titled “Day 4 · Step 1 · Quantization”- The download stops because the disk is full.
- The three sizes need about 17 GB, and the quality loop later stores the 8-bit and 3-bit versions again (about 12.5 GB). Free some space and run
bash quant_bench.shagain: sizes already downloaded are skipped. Once its table has printed, you can delete themodelsfolder insidelessons/04-quantization. If the disk filled during the quality loop instead, delete that folder now (the speed test no longer needs it) and run the loop again. - It stops with
Unknown bandwidth — pass --bw. - The script’s list of Apple chips does not include yours. The error lists the chips it knows with their speeds; your Mac’s tech specs page on Apple’s website lists its memory bandwidth. Then run
python ceiling_check.py --bw 273, with your figure in place of 273. The timings are already saved, so nothing downloads again. - Efficiency is well below 60%.
- Something else is slowing the Mac: another app using memory heavily, a hot Mac slowing itself down, or part of the model running on the Mac’s main processor instead of its graphics cores. Close big apps (video calls, browser tabs playing video), plug in, and run
bash quant_bench.shagain; downloads are skipped if you have not deletedmodels/. Some chips, such as the M3 Max and M4 Max, come in two versions, and the script assumes the faster one (felab/hardware.py). If yours is slower, runpython ceiling_check.py --bwfollowed by your figure in GB per second. - The quality loop prints a connection error (for example
Connection refused) for one size. - The server was still downloading that size when the questions started: the loop waits only 20 seconds. Start that size alone, for example
bash ../02-three-local-servers/serve_llamacpp.sh Q8_0, wait for the line saying it is listening, stop it with Ctrl+C, then run the loop again. llama-bench: command not found- llama.cpp is missing. Day 1’s setup installs it; to add it again, run
brew install llama.cpp, thenbash quant_bench.sh.
Day 4 · Step 2 · KV cache and context length
Section titled “Day 4 · Step 2 · KV cache and context length”- Every 8-bit (q8_0) row fails, even the short 4,096-token one.
- Turn on the server’s flash attention setting. (Flash attention is a faster way to run the look-back step; llama.cpp needs it to store the notes in 8 bits.) Run
EXTRA="-fa on" bash context_ladder.sh. On older llama.cpp versions useEXTRA="-fa" bash context_ladder.sh. - The 131,072-token 16-bit row says failed to start.
- That is the lesson. The notes (17.18 GB) plus the 4.9 GB model need about 22 GB, but a Mac gives the model only about 70% of its memory (Day 1’s rule): about 17 GB on a 24 GB Mac. Write “failed” in Your numbers; the 8-bit row below it is the fix.
- Both 131,072-token rows fail (a 16 GB Mac).
- That can happen. The 9.13 GB of 8-bit notes plus the 4.9 GB model is about 14 GB, more than the about 11 GB a 16 GB Mac gives the model (70%, Day 1). Write “failed” for both; the 32,768-token pair shows the same halving.
- Your rows are exactly half the lesson’s, for example 2,048 and 1,088 at 32,768 tokens.
- llama.cpp printed the notes’ two halves (keys and values) on the same line, and the script copied only the last number. In
lessons/05-kv-cache, rungrep -i "kv buffer size" logs/ctx32768_f16.log, using the log named after the row (its ctx, then f16 or q8_0). The MiB number on that line is the full figure.
Day 5 · Step 1 · Prefix caching
Section titled “Day 5 · Step 1 · Prefix caching”Connection refused.- That server is not running.
make check(from the03-labsfolder) shows what is up. Start the practice server withmake mock &, or start llama.cpp as in the run notes and wait until it says it is listening. - Opening first on the practice server shows a speed-up near 1, and call 1 is already fast.
- The practice server still remembers the opening from an earlier run: this command run twice, or
make smokewhile it was up. Restart it:pkill -f felab.mock_server, thenmake mock &from the03-labsfolder, and run again. No such file or directorywhen you start llama.cpp.- The command’s path starts from this lesson’s folder. From the
03-labsfolder runcd lessons/06-prefix-caching, then start it again. - llama.cpp returns an error about the context size, or its opening-first run shows no speed-up.
- Each of its 4 seats holds only 2,048 tokens by default, less than the 3,000-token opening. Stop it (Ctrl+C) and, from this lesson’s folder, start it with
bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 4(4,096 tokens per seat). FIREWORKS_API_KEY is not set.- Put your key in the
.envfile in03-labs, as on Day 1, or export it in this terminal. cached_tokens=0orcached_tokens=None.- Run the Fireworks command again right away: caches expire when idle (the Prefix Caching video). If it still says None, this model does not report it; pass another with
--model. - A Fireworks model id is not found.
- Model ids change. Copy a current one from the Fireworks model library and pass it with
--model. mlx_lm.generate: command not found.- The Python environment from Day 1 is not active. From the
03-labsfolder runsource .venv/bin/activate, then try again.
Day 6 · Step 1 · Structured output
Section titled “Day 6 · Step 1 · Structured output”- Every mode prints
server rejected request (APIConnectionError: Connection error.)and the table says(no rows). - That server is not running.
make check(from the03-labsfolder) shows what is up. Start the practice server withmake mock &, or llama.cpp withbash lessons/02-three-local-servers/serve_llamacpp.shin a second terminal, and wait until it says it is listening. data/tickets/test.jsonl missing.- The ticket data is not built yet. From the
03-labsfolder runpython data/make_tickets.py, then run the script again. No module named 'felab'(or'openai').- The Python environment from Day 1 is off in this terminal. From the
03-labsfolder runsource .venv/bin/activate, then try again. - One mode prints
server rejected requestwith an error such asBadRequestError, while the other modes print rows. - That server does not accept that way of asking, so the script skips the mode and carries on. Write it down: which modes a server supports is a real finding for a customer.
FIREWORKS_API_KEY is not set, or every mode printsAuthenticationErroron Fireworks.- Your key is missing or wrong. Put your key in the
.envfile in03-labs, as on Day 1, or export it in this terminal. - Every mode prints
server rejected request (NotFoundError: ...)on Fireworks. - The model id is not found. Model ids change: copy a current one from the Fireworks model library and pass it with
--model. - Fireworks shows a low schema_valid_% or empty replies, even with the format enforced.
- gpt-oss-120b writes out its thinking before its answer (a reasoning model, check 3). That thinking may use up the script’s 120-token reply limit, so the answer comes back cut off or empty. Write it down as a finding and compare with the prompt row, which follows check 3’s pattern: schema in the prompt, nothing enforced, checked afterwards.
Day 6 · Step 2 · Batch API
Section titled “Day 6 · Step 2 · Batch API”Connection refusedon the practice run.- The practice server is not running. From the
03-labsfolder runmake mock &, then run the command again. firectl: command not found.- Install it as on Day 1:
brew tap fw-ai/firectl && brew install firectl, thenfirectl signin. - A firectl command rejects its options.
- Commands change. Run it with
--help, for examplefirectl batch-inference-job create --help, and follow what it says; the script’s own note says the same. - The job is still running when you have to stop.
- Nothing is lost: Ctrl+C stops the checking, not the job. Later, from
lessons/08-batch-api, runfirectl batch-inference-job get <job name>with the name from thecreate job:line. When it says COMPLETED, run the last two commands. - The job stays pending and shows no error.
- Fireworks’ batch guide (checked 27 September 2026) warns that a job on a model batch does not support can stay pending with no error. Batch should still take gpt-oss-20b, but if yours has not started by the time you finish today’s other work, start a second job on gpt-oss-120b, which is offered per token: from
lessons/08-batch-api, runbash submit_batch.sh accounts/fireworks/models/gpt-oss-120b. Write down its new job name and follow that one. Your score is then gpt-oss-120b’s, so label it that way. - The job ends FAILED or EXPIRED.
- Read the job details the script printed at the end for the reason. Model ids change: if the model is no longer offered, copy a current one from the Fireworks model library and run
bash submit_batch.sh <that id>. answeredis below 100.- Some requests failed and went to the job’s error file instead of the results. The score covers the ones answered, each matched to its ticket by
custom_id: that is exactly why the label exists. Pass a results file, or --simulate.- The path after
score_batch.pydoes not point to a file. Use the exact name of the downloaded.jsonlfile;lslists what is in the folder.
Day 7 · Step 1 · Eval harness
Section titled “Day 7 · Step 1 · Eval harness”Connection refusedon the practice run.- The practice server is not running. From
03-labs, runmake mock &, then run the command again. Connection refusedon theollama mlxrun.- One of the two servers is not running:
make check(from03-labs) shows which. Start Ollama withollama serve, and MLX in another terminal withbash lessons/02-three-local-servers/serve_mlx.sh. FIREWORKS_API_KEY is not set- Your key belongs in the
.envfile in03-labs, as on Day 1: a lineFIREWORKS_API_KEY=followed by the key. Then run the command again. - The kit’s gpt-oss-20b line fails, or says the model is not found.
- Expected: gpt-oss-20b is no longer offered per token (“Serverless: Not supported” on its model page, checked 27 September 2026). Run the gpt-oss-120b line from the code block instead:
python evaluate.py "fireworks:accounts/fireworks/models/gpt-oss-120b@0.15/0.60" --n 50 --judge fireworks. - A Fireworks model id is not found, even on the gpt-oss-120b line.
- Model ids change. Copy a current one from the model library on fireworks.ai that is offered serverless, put it in place of
accounts/fireworks/models/gpt-oss-120b, and copy its current prices after the@too. - The Fireworks row shows a low
schema_valid_%, or empty replies. - gpt-oss-120b writes out its thinking before its answer (a reasoning model, as on Day 6).
evaluate.pyallows each reply 120 tokens (max_tokens), and the thinking may use them up, so the answer comes back cut off or empty. Write it down as a finding: it is part of that model’s row, not a broken setup. - The
--judge fireworksrun says a model is not found. - The judge is your default Fireworks model, gpt-oss-120b. Put a current id in
03-labs/.envasFIREWORKS_MODEL=accounts/fireworks/models/<id>, then run the command again. judge_1to5says nan.- No judge was asked: only runs with
--judgefill that column. Theollama mlxrun has no judge, so nan is expected there.
Day 7 · Step 2 · LoRA on Fireworks
Section titled “Day 7 · Step 2 · LoRA on Fireworks”1_train.shfails because the base model cannot be fine-tuned, or is not found.- Pick a model from Fireworks’ list of tunable models and run
BASE_MODEL=accounts/fireworks/models/<id> bash 1_train.sh. Pass the sameBASE_MODEL=...tobash 4_eval.shlater: the scripts do not remember it, and would grade the default base instead. 2_wait.shstops with FAILED or CANCELLED.- It prints the job’s details: read the error there. Nothing bills by the hour yet. Once it is fixed, run
bash 1_train.shagain: it starts a new job with new names. 3_deploy.shkeeps printing dots and never reachesready.- The deployment already exists, so it may be billing. Press Ctrl-C and check the deployment’s state in the Fireworks console. If it shows ready, carry on with
bash 4_eval.sh. If it is not coming up, runbash 5_teardown.sh, thenmake fw-checkfrom03-labs. 4_eval.shstops with an error, such as a model not found.- Your GPU is still billing: if you cannot fix this in a few minutes, run
bash 5_teardown.shfirst and come back. If the error names your fine-tune, copy the model string from the deployment’s API tab in the Fireworks console and runpython ../09-eval-harness/evaluate.py "fireworks:<that string>" --n 60. If it names the base model, its id may have changed: copy a current one from the model library and runBASE_MODEL=accounts/fireworks/models/<id> bash 4_eval.sh. - You closed the terminal while the deployment was running.
- Go back to
lessons/10-lora-fireworksand runbash 5_teardown.sh. The ids are saved in the.statefile, and the clock runs until you do. 5_teardown.shstops with an error, or you are unsure whether anything is still billing.- From
03-labs, runmake fw-check. Delete anything it lists withfirectl deployment delete <DEPLOYMENT_ID>, then runmake fw-checkagain until the list is empty.
Day 8 · Step 1 · The same LoRA on your Mac
Section titled “Day 8 · Step 1 · The same LoRA on your Mac”- Training stops with an out-of-memory error, or the Mac slows to a crawl.
- Close big apps first. Then in
1_train_mlx.shchange--batch-size 2to--batch-size 1(the script’s own advice) and run it again. If it still fails, also change--num-layers 8to 4: fewer layers with add-ons need less memory, but the add-on can learn less. mlx_lm.lora: command not found- The Python environment from Day 1 is not on. From the
03-labsfolder runsource .venv/bin/activate, thencd lessons/11-lora-local-mlxand run the script again. 2_fuse_and_serve.shfails with ‘Address already in use’.- Another server already uses port 8081, most likely Day 2’s MLX server. Stop it with Ctrl+C in its window, then run the script again.
3_eval_local.shsaysConnection refused.- The server is not up yet: fusing comes first and takes a while. Wait for the line ending
serving on :8081 (Ctrl-C to stop)in the server window, give it a few seconds to load, then run the eval again. Keep that window open. - The eval gets through the fine-tune, then stops with an error at
evaluating mlx mlx-community/Qwen2.5-7B-Instruct-4bit. - Your copy of the MLX tools will not switch models on one server. No table prints, but the fine-tune’s scores are already saved: the last row of
03-labs/results/09-eval.csv, namedmlx fused. Copy itscategory_%,severity_%andp95_msfrom there. Stop the server with Ctrl+C in its window to free memory. In that window, start the original model on a second port:mlx_lm.server --model mlx-community/Qwen2.5-7B-Instruct-4bit --port 8082. Then in the eval window (still in the lesson folder) runMLX_URL=http://localhost:8082/v1 python ../09-eval-harness/evaluate.py "mlx:mlx-community/Qwen2.5-7B-Instruct-4bit" --n 60. It prints the original model’s row. - Validation loss starts rising again before step 300.
- The model has begun memorising the training tickets. Stop with Ctrl+C; the last save is kept. For a rerun, lower
--iters 300in1_train_mlx.shto about the step where the loss was lowest.
Day 9 · Step 1 · MoE vs dense
Section titled “Day 9 · Step 1 · MoE vs dense”Connection refused.- Ollama is not running. Start it with
ollama serve &, then run the command again. gpt-oss:20b not pulled — run: ollama pull gpt-oss:20b.- Run that pull. The name must match exactly, including
:20b. pass --bw <GB/s>.- The script does not know your chip’s memory speed. Look it up and add it by hand, for example
python moe_probe.py llama3.1:8b gpt-oss:20b --bw 546for a chip that moves 546 GB a second. - gpt-oss:20b is very slow or will not load.
- Your Mac is probably short of memory: the model needs its whole 13.8 GB, and Day 1’s rule keeps models under about 70% of memory. Close other apps, or use the
--demonumbers.
Day 9 · Step 2 · Speculative decoding
Section titled “Day 9 · Step 2 · Speculative decoding”spec_bench.pycannot connect right after a server start.- The server was still loading. The first speculative start also downloads the draft model, which can take longer than the 40-second wait. The
; kill %1on that line then stops the server too: start it again (bash serve_spec_llamacpp.sh &), wait until it says it is listening, then run thespec_bench.pyline again. - The server will not start because port 8080 is in use.
- Another llama.cpp server from Day 2, 4 or 5 is still running. Stop it with Ctrl+C in its terminal, then start again.
kill %1says there is no such job, or stops the wrong program.%1is the first background job of this terminal. Runjobsto list them and use the right number, for examplekill %2.- The server refuses the draft model.
- The draft must share the big model’s tokenizer. If you changed
DRAFT_REPOorLLAMACPP_REPO, set them back, or pick a draft from the same model family. fireworks_spec.shstops with an error fromfirectl.- Check that you are signed in (
firectl whoami, Day 1) and that your quota allows a deployment (firectl quota list). Then runmake fw-check: it should list nothing. fireworks_spec.shnever says ready.- Press Ctrl+C: the script deletes the deployment on the way out. Run
make fw-checkto confirm nothing is left, and try again later.
Day 9 · Step 3 · The Spark track
Section titled “Day 9 · Step 3 · The Spark track”set SPARK_IP in .env.- Add
SPARK_IP=<the Spark's address>to the.envfile in03-labs, plusSPARK_USER=<your login>if it is notnvidia. - SSH asks for a password, or says
Permission denied (publickey). remote.shneeds SSH keys. Set them up once withssh-copy-id nvidia@<the Spark's address>(with your own login if it differs), then try again.- An image tag is not found.
- Container tags change often. Look up the current GB10/arm64 tag (links are in the script comments). Then replace the tag after
:-in theIMAGE=line near the top of1_vllm.shor3_sglang.sh, in your copy on the Mac (remote.shsends that file to the Spark).remote.shpasses onlyHF_TOKEN, soVLLM_IMAGEorSGLANG_IMAGEset on your Mac never arrives. - A model download fails with an access error.
- It may be a gated model, one you must request on Hugging Face first. Request access, then put your token in
.envasHF_TOKEN=…;remote.shpasses it to the Spark. 2_bench.shsayspython: command not foundorNo module named 'felab'.- It runs on your Mac. From the
03-labsfolder, runsource .venv/bin/activatefirst. - The prefix test cannot connect to
http://:30000/v1. $SPARK_IPwas empty in your terminal. Runexport SPARK_IP=<the Spark's address>, or paste the command3_sglang.shprinted.nvidia-smishows memory as N/A.- Normal on a Spark: its CPU and GPU share one memory, so the usual fields are blank. Use
free -hon the Spark instead.
Day 10 · Step 1 · Sizing memo and 10-minute talk
Section titled “Day 10 · Step 1 · Sizing memo and 10-minute talk”ModuleNotFoundError: No module named 'felab'- The virtual environment is off in this window. From the
03-labsfolder runsource .venv/bin/activate, thencd lessons/15-capstoneand run the command again. - A section says TODO although you did that day.
- The memo reads fixed file names in
results/. Fromlessons/15-capstone, runls ../../results/and look for that day’s file (the list is under Before you start on the Day 10 overview). If it is missing, run that day’s main command again; the practice-server version is free. - Replicas, the dedicated cost and the crossover say TODO, but
03-sweep.csvis there. - No crowd size in your sweep kept the memo’s speed promise, by default 95 of 100 first words within 1,000 ms. Once your sweep has any real-server rows, the memo ignores the practice-server rows. From
lessons/15-capstone, open../../results/03-sweep.csvand read thettft_p95column (in ms). Pass the customer’s real target with--slo-ms, or run Day 3’s sweep again with smaller crowds, for example--levels 1,2,4. - The memo warns that capacity came from a laptop or mock server.
- Expected: your Day 3 sweep ran on your Mac or the practice server. Data-center GPUs running a production engine such as vLLM hold far more people per copy, so treat the copies and the dedicated cost as placeholders and say so on your risks slide. Day 9’s optional DGX Spark step, or a Fireworks deployment, gives real figures.
- The fine-tune row names
mock mock-8b-lorainstead of your own model. - The memo compares every row in
results/09-eval.csv, including the practice-server rows Day 1’s smoke test saved. Open that file in TextEdit, delete the lines containingmock mock-8b, keep the first line (the column names), save, and runpython build_memo.pyagain. - The crossover is hundreds of millions of tickets a month.
- Not a bug in your data: the memo assumes today’s copies carry any volume. With a handful of people per copy, the dedicated option needs many GPUs, so serverless wins by a wide margin: with the practice server’s 8, the example’s crossover is about 744 million tickets, 372 times its volume. Replace the capacity with a real GPU measurement before you quote it.
opensays no application knows how to open the memo.- Your Mac has no app set for
.mdfiles. Runopen -e ../../results/SIZING_MEMO.mdto open it in TextEdit instead. - The talk runs well over 10 minutes.
- Five slides in 10 minutes is about 2 minutes each. Keep one number per point, and move tables to a backup slide you show only if asked.
Day 10 · Step 2 · One product criticism
Section titled “Day 10 · Step 2 · One product criticism”- I skipped Day 7, so I have no fine-tune of my own.
- Use something you did hit: Day 9’s speculative decoding test also needed a dedicated deployment (30 to 45 minutes of one), or a model name that no longer worked. If you ran into nothing, you can use the docs disagreement in section 1 above, but say you found it by reading, not by running.
- Everything worked, so I have nothing to criticise.
- Look for what cost you time or money rather than what broke: a step that needed a GPU billed by the hour, a price you had to look up in two places, a limit you only found in a changelog (the provider’s list of recent changes). A small, specific friction beats a big, vague one.
- My criticism may be out of date.
- The lab book read Fireworks’ pages on 23 September 2026 and warns that prices move weekly. Read the pricing page and the LoRA deployment docs again before sharing your recommendation. If it has been fixed, say so and ask what drove the change: that shows you keep up.
- It sounds like a complaint.
- Start with what works, with a number, and end with your proposal and a question. Cut “always”, “never” and “broken”; keep your one measured number.
Day 11 · Step 1 · Training toys, part 1
Section titled “Day 11 · Step 1 · Training toys, part 1”ModuleNotFoundError: No module named 'numpy'- The Python environment is off in this terminal. From the
03-labsfolder runsource .venv/bin/activateand try again. The toys need nothing else, so on another machinepip install numpyis enough. - Your numbers differ a little from the lesson’s.
- The toy draws its random numbers from a fixed starting point (
--seed 0), so the shape should match: rank 1 stuck, rank 2 and above matching full at 200 examples. Small differences in the last digits do not matter. Runpython toy_finetune.py --seed 1to see the same shape with other random numbers.
Day 11 · Step 2 · Full vs LoRA vs DoRA
Section titled “Day 11 · Step 2 · Full vs LoRA vs DoRA”- The full run stops with an out-of-memory error, or the Mac slows to a crawl.
- Full training peaks at about 8 to 10 GB. Close big apps and try again, or skip it. To skip it, first remove what the failed run left behind in the lesson folder, or
compare_variants.pywill try to test it:rm -rf logs/full.log adapters/full. Then runVARIANTS="lora dora" bash run_variants.sh. Record that full training did not fit on your Mac: it is a useful measured limit for your comparison. - Training takes too long.
- Run
ITERS=150 bash run_variants.sh: 150 steps instead of 300, about half the time. Note it next to your numbers. run_variants.shstops withfix the data first.sft_data_check.pyfound a problem in the training file and printed its type with the first line number. If you editeddata/tickets/, rebuild it from the03-labsfolder withpython data/make_tickets.py.compare_variants.pysaysNo logs yet: run bash run_variants.sh first.- It reads the
logs/folder insidelessons/17-full-vs-peft-mlx. Runbash run_variants.shthere first, thenpython compare_variants.pyfrom the same folder. compare_variants.pysays the model evaluation was skipped.- The accuracy and quiz columns need mlx-lm on an Apple-silicon Mac. Turn the environment on (
source .venv/bin/activatefrom03-labs) and run it again. The memory, time and file-size columns come from the logs either way. mlx_lm.lora: command not found- mlx-lm is missing from the environment. From
03-labs, runsource .venv/bin/activate, thenpip install -r requirements.txt(Day 1’s setup), thenbash run_variants.shagain.
Day 12 · Step 1 · Training toys, part 2
Section titled “Day 12 · Step 1 · Training toys, part 2”ModuleNotFoundError: No module named 'numpy'- The Python environment is off. From the
03-labsfolder runsource .venv/bin/activate, then run the script again. If it still fails, runpip install numpy. --beta 0, as on the RLHF video’s command card, prints the same as the default.- The script treats 0 as its default, 0.5, because one of its lines divides by β. For a loose leash use a small number instead, such as
python toy_alignment.py --beta 0.02. - Your section 5 table has 8 rows, not 3.
- That is expected: the script tries 8 leash settings, from 4 down to 0.01. The lesson shows 3 of them (0.50, 0.10 and 0.02).
- At
--beta 0.1, 5b still favours the right-category answer, but the sweep says β 0.10 ends on the wrong one. - Expected: the sampled run stops short of the optimum. Compare the two KL numbers on the 5b line: 0.94 reached, 1.81 for the optimum. The script’s closing note that 5b lands on the optimum holds at the default β 0.5 (0.34 and 0.34), not here.
Day 12 · Step 2 · Preference tuning with DPO
Section titled “Day 12 · Step 2 · Preference tuning with DPO”pref_eval.py --target mocksaysConnection refused.- The practice server is not running. From the
03-labsfolder runmake mock &, then try again. pref_eval.py --target mlxsaysConnection refused.- The two model servers are not running.
dpo_mlx.shprinted their commands at the end. In each new terminal, first go to03-labsand runsource .venv/bin/activate, thencd lessons/18-preference-dpo. Startmlx_lm.server --model <START> --port 8081in one terminal and the same with--adapter-path adapters/dpo --port 8082in another. When both are up, run the check again. - You do not know what to type for
<START>. - Use the model
dpo_mlx.shprinted on itsDPO starting from:line, such as the full path tosft_model. It is also saved in the file.statein the lesson folder. dpo_mlx.shstarts from the instruct model, notsft_model.- It did not find Day 11’s LoRA in
lessons/17-full-vs-peft-mlx/adapters/lora. DPO still runs, but on a model that gets only about 40% of categories right (check yourself 1). From03-labs, runVARIANTS=lora bash lessons/17-full-vs-peft-mlx/run_variants.sh(5 to 15 minutes), thenbash lessons/18-preference-dpo/dpo_mlx.shagain, and restart both servers with the new<START>it prints (sft_model). mlx_lm_lora.trainstops with an unknown option or argument error.- The tool’s options change between versions, and the script says its
--helpwins. Runmlx_lm_lora.train --help, match the options at the top ofdpo_mlx.shto it, and run it again. category_%dropped after DPO.- Taste has cost correctness. The lesson’s fixes: a larger β (
BETA=0.2 bash dpo_mlx.sh), fewer steps (ITERS=100 bash dpo_mlx.sh), or more wrong-category pairs, so correctness is part of what “preferred” means. - The pairwise rows are not in
results/18-pref.csv. - They have different columns, so the kit’s recorder saves each one in a new file named
results/18-pref-<time>.csv. Copy the numbers from the terminal or from those files. dpo_fireworks.shsays the base model is not found or not DPO-enabled.- Model ids change. Pick one the Fireworks model library marks as DPO-enabled and run
BASE_MODEL=accounts/fireworks/models/<id> bash dpo_fireworks.sh. Afterwards runmake fw-check: it should list nothing.