Day 8 · Training day, local
Part 1 · Day 8 of 12
0 of 12 days done
About 1 h 35 Free Mac
Today: Train Day 7’s small add-on for a model (a LoRA adapter) again, from scratch, this time on your own Mac, for free. Then put the two side by side: accuracy, time, cost and privacy. That comparison is your answer when a customer asks whether to build it themselves or buy a service.
Example
Imagine the support team from Day 7 cannot send its customer tickets to a cloud service. It still wants the model to learn to sort them better. Changing every number in a large model needs more memory than its Mac has; training a small add-on while leaving the model alone fits. Today you will try that local route and compare its accuracy, cost and privacy with the cloud route.
By the end you’ll have
A filled-in build-versus-buy page (COMPARISON.md): the same fine-tune on Fireworks and on your Mac, side by side, with your own numbers.
Expected: In the lesson’s sample, the Mac’s fine-tune puts 96% of test tickets in the right category, up from 78% before training, for $0; Day 7’s Fireworks sample went from 80% to 97% for about $5 of rented GPU time.
The 6 numbers you write down
Right category on your Mac, before → after training (%)
The share of the 60 test tickets whose category (billing, outage, how-to or abuse) the model got right, before and after training. One ticket is worth 1.67 points.
Example: 78.0 → 96.0 (the lesson’s sample; Day 7’s Fireworks sample: 80.0 → 97.0)
Right severity on your Mac, before → after training (%)
The share of test tickets whose urgency score (1 to 5) the model got exactly right.
Example: 66.0 → 94.0 (the lesson’s sample; Day 7’s Fireworks sample: 70.0 → 95.0)
Reply time of your fine-tune on the Mac, p95 (ms)
How long a reply took: of the 60 test tickets, sent one at a time, 57 got theirs within this time and 3 took longer. ms means thousandths of a second, so 380 ms is 0.38 seconds. It goes in the p95 latency row of
COMPARISON.md, next to Fireworks.Example: No Mac sample in the lesson; Day 7’s Fireworks sample was 380 ms.
Adapter size on disk
The whole trained add-on: the file you would send, keep versions of and swap. Compare it with the 4.3 GB model.
Example:
7.1Mon the sample’sadapter saved:line, about 7.4 MB (the lesson’s sample)Time from starting training to your Mac answering requests (minutes)
The real time that passes (wall-clock row of
COMPARISON.md): how long a customer waits for a working tuned model.Example: Training alone takes about 15 to 40 minutes, the lesson’s range.
Your recommendation: local, managed, or both
Your build-versus-buy answer, in a sentence a customer can use.
Example: Prototype locally, free and private; ship on Fireworks once it must serve many users.
How you will use this: When a team asks whether to train locally or use a managed service, you can compare the same ticket task both ways. Day 7 gives you the cloud result and bill; today adds the Mac result and its privacy tradeoff. Use both to recommend a route instead of relying on a product claim.
Before you start
Day 7’s Fireworks results at hand.
Today you set your Mac’s results next to Day 7’s cloud run on Fireworks. You compare how often each model put tickets in the right category, before and after training. You also compare how often it rated their urgency (severity) exactly right, how fast it replied and what it cost. Day 7’s scores are in
03-labs/results/09-eval.csv; today’s test adds its own rows below them, withmlxin the model name. If you skipped Day 7, use its sample: right category 80% before training, 97% after, about $5 of rented GPU time.About 9 GB of free disk and 6 to 10 GB of free memory.
The starting model, stored at 4 bits per number, downloads once: about 4.3 GB (the lesson’s estimate).
2_fuse_and_serve.shthen writes a second full copy with the add-on baked in, about as big again: 4.3 + 4.3 = 8.6, so about 9 GB. Training uses about 6 to 10 GB of memory, so close video calls and heavy browser tabs first.The Python environment from Day 1.
Every command today runs with it on, because it holds
mlx_lm, Apple’s tool that trains and serves models on your Mac. In each new terminal window, from the03-labsfolder:source .venv/bin/activate(.venv)appears at the start of the prompt.
Warm-up from Day 7
Section titled “Warm-up from Day 7”About 2 minutes. Say your answer out loud, then tap to check it. From Day 7 · Training day, managed.
A judge model (a second AI model that grades answers against a scoring guide) gives the fine-tuned model 4.5 out of 5 and the original model 3.9. Is that proof the fine-tune is better?Show answerHide
In plain words
No. First check that the judge agrees with you: score about 20 of the same answers yourself and compare. AI judges tend to reward longer, more polished answers even when they are no more correct.
Picture it
A teacher who gives longer essays in neat handwriting higher marks, whatever they say. Before you trust their grades, you mark 20 of the same essays yourself and see whether you agree.
With real numberslesson 09’s example table and its scoring guide
- The judge scores each suggested next step from 1 to 5: 5 = specific, right owner, safe; 3 = plausible but vague; 1 = wrong or unsafe.
- Average scores: 3.9 for the original model, 4.5 for the fine-tune, a gap of 0.6 points (4.5 - 3.9).
- The check: score about 20 of the same answers by hand. If you and the judge often disagree, the 0.6 means little.
- Code checks need no such trust: 81% against 96% of categories right is counted against the known answers.
Words to know
- Judge (LLM judge)
- A second model that scores answers against a rubric. Example: gpt-oss-120b in the Fireworks run.
- Rubric
- The scoring guide a judge follows. Example: 5 = specific, right owner, safe; 1 = wrong or unsafe.
- Hand-labelled sample
- Answers a person has marked right or wrong, used to check a judge. Example: about 20 scores, redone by you.
- Bias (judge bias)
- A judge’s habit of favouring something that is not correctness. Example: longer, more polished answers.
Go deeper: the engineer version
The kit's question
The judge gives the fine-tune 4.5 and the base 3.9. Is that proof?
The kit's answer
No. Check that the judge agrees with you on a hand-labelled sample first. Judges favour length and style.
More detail: The lesson also says to change the RUBRIC prompt and watch the scores move, which shows why judges need calibrating. evaluate.py prints only the average (judge_1to5), so to hand-check, print each judge_score() result next to its ticket. In the practice run the stand-in judge is rigged: it gives 4 or 5 when the next step names the ticket’s true category and 1 to 3 otherwise (felab/mock_server.py), so its scores track category accuracy. A real judge has no such anchor.
After fine-tuning, accuracy went up 17 points: 17 more tickets in every 100 sorted correctly. But p95 latency doubled: the time within which 95 of 100 replies come back is now twice as long. What do you recommend?Show answerHide
In plain words
Recommend it only if the slower replies still meet the customer’s speed promise (their SLO, service level objective). Show them both numbers so they choose. Then offer ways to keep the accuracy without the wait: fine-tune a smaller starting model, or add a small helper model that drafts the next few words for the big one to check (speculative decoding).
Picture it
A delivery service that gets more parcels to the right door but takes twice as long. For flowers that must arrive today, that is a no; for a monthly supplies order, it is fine.
With real numberslesson 09’s example table
- The question’s trade: 17 more tickets right in every 100, but the time within which 95 of 100 replies come back is twice as long.
- For scale, the eval lesson’s example table has p95 figures of 310 ms and 640 ms: 640 ÷ 310 = 2.1, about a third of a second against about two thirds.
- In that same table the small fine-tuned model is the fast one: 96% right, 310 ms, $0.02 per 1,000 tickets, against the big model’s 81%, 640 ms, $0.42. That is why a smaller starting model is one of the answers.
- The lesson’s own example is harsher, 3 times the wait: that can be the wrong call for a chat product and the right one for batch work that runs overnight (like Day 6’s Batch API).
Words to know
- SLO (service level objective)
- The speed promise. Example: 95 of 100 replies back within a set time.
- p95 latency
- The time within which 95 of 100 replies come back. Example: 640 ms against 310 ms.
- Speculative decoding
- A small draft model guesses several tokens and the big model checks them all in one pass; the output does not change. Example: Day 9.
- Base model
- The model you start from, before any fine-tuning. Example: a smaller one can be faster and cheaper.
Go deeper: the engineer version
The kit's question
Accuracy went up 17 points but p95 doubled. What do you recommend?
The kit's answer
It depends on the SLO. Put both in the memo, and consider a smaller base model or speculative decoding.
More detail: The p95 here is whole-request latency: evaluate.py times each call from sending it to the full reply, not time to first token. Speculative decoding (Day 9) cuts the time per token when a small draft model’s guesses are accepted. A smaller base model cuts time and cost together, and the lesson’s table shows a small fine-tune beating a big general model on all three columns.
The customer wants 12 fine-tuned versions of a model, one for each business unit. Served separately, each one needs its own reserved GPU (the chip that runs the model), billed by the hour. How do you keep serving costs sensible?Show answerHide
In plain words
Share one GPU. Keep one copy of the base model running and load the right small add-on (adapter) for each request. This is multi-LoRA: one bill instead of 12.
Picture it
One kitchen with one set of equipment and 12 spice packets, one per restaurant brand, instead of 12 separate kitchens. Each order says which packet to use.
With real numberslesson 10’s deployment rate and the lab book’s prices
- One dedicated GPU: about $8 an hour, charged whether or not anyone uses it.
- Running all month (about 730 hours): 730 x $8 = $5,840.
- 12 separate deployments: 12 x $5,840 = $70,080 a month.
- Multi-LoRA, 1 deployment holding 12 adapters: $5,840, one twelfth, as long as one GPU keeps up with the traffic.
- Each adapter is small: about 30 MB in the Training Techniques video. One deployment holds up to 100 by default, the lab book says.
Words to know
- Multi-LoRA
- Serving many LoRA add-ons on one running copy of the base model, picking the add-on for each request. Example: 12 business units, 1 GPU.
- Adapter
- The small add-on LoRA trains, applied on top of the frozen model. Example: about 30 MB.
- Dedicated deployment
- A GPU reserved for you on Fireworks, billed by the second at an hourly rate, used or not. Example: about $8 an hour.
- Base model
- The model you start from, before any fine-tuning. Example: Qwen2.5 7B Instruct in this lesson.
Go deeper: the engineer version
The kit's question
The customer wants 12 fine-tunes, one per business unit. How do you keep serving costs sane?
The kit's answer
Multi-LoRA: one base deployment with 12 adapters, loaded per request.
More detail: Adapters are loaded per request: each request names its adapter, and the server applies it on top of the shared base weights, batching requests for different adapters together (the field guide: hundreds of fine-tunes on one base at the base model’s price). The lab book adds the Fireworks details: a default quota of 100 adapters per deployment, and multi-LoRA needs a BF16 (16-bit) deployment shape with --enable-addons; FP8 and FP4 (8-bit and 4-bit) shapes cannot host adapters.
Today’s steps
Section titled “Today’s steps”1 step, then the wrap-up.
-
The same LoRA on your Mac
You repeat Day 7’s fine-tune (extra training on your own examples) on the same 800 support tickets, but on your own Mac, for free. Doing it both ways gives you the build-versus-buy comparison (do it yourself or pay a service), with your own numbers, that you can explain to a customer.
- Wrap-up and drill Put your numbers in one table, answer 3 questions out loud, explain 1 result and work 1 example.
Your progress is saved on this device.