Skip to content

Day 7 wrap-up and drill

  1. Overview
  2. Step 1
  3. Step 2
  4. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

Your notes show what the fine-tune gained and what it cost. Gain: categories right before and after (the lesson’s sample: 80% to about 97%). Cost: minutes deployed and their dollars (38 minutes, $5.07), plus about 15 cents of training. Step 1’s untuned Fireworks row is written down too, and make fw-check lists nothing.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 1 · Eval harness

The bar a fine-tune must clear, on quality, speed and cost at once, to be worth paying for.

Example: 81 / 640 / 0.42 (the lesson’s sample big general model, written for gpt-oss-20b; yours will differ)

How a second model rated the suggested next steps. Trust it only after checking some of its scores by hand.

Example: 3.9 (the lesson’s big general model)

What a free model on your own machine does on the same task, before any training.

Example: ollama __ / __; mlx __ / __ (the lesson has no sample for these)

Step 2 · LoRA on Fireworks

The gain your fine-tune bought, measured on tickets it never saw in training.

Example: 80 / 97 (the lesson’s sample, a 17-point gain; yours land on values like 96.7 or 98.3)

95 of 100 replies came back within this many milliseconds. A fine-tune should not be much slower than its base. Here the base runs on Fireworks’ shared servers and the fine-tune on its own GPU, so part of any difference is the setup, not the training.

Example: 420 / 380 (the lesson’s sample)

What the whole experiment cost. Serving is nearly all of it; training is about 15 cents.

Example: 38 min, $5.07 serving + about $0.15 of training (the lesson’s sample)

Every question from today’s steps. Say your answer out loud, then show the answer and rate yourself. 5 cards

QuestionStep 1 · Eval harness

A judge model (a second AI model that grades answers against a scoring guide) gives the fine-tuned model 4.5 out of 5 and the original model 3.9. Is that proof the fine-tune is better?

Show answerHide answer

In plain words

No. First check that the judge agrees with you: score about 20 of the same answers yourself and compare. AI judges tend to reward longer, more polished answers even when they are no more correct.

Picture it

A teacher who gives longer essays in neat handwriting higher marks, whatever they say. Before you trust their grades, you mark 20 of the same essays yourself and see whether you agree.

With real numberslesson 09’s example table and its scoring guide

  • The judge scores each suggested next step from 1 to 5: 5 = specific, right owner, safe; 3 = plausible but vague; 1 = wrong or unsafe.
  • Average scores: 3.9 for the original model, 4.5 for the fine-tune, a gap of 0.6 points (4.5 - 3.9).
  • The check: score about 20 of the same answers by hand. If you and the judge often disagree, the 0.6 means little.
  • Code checks need no such trust: 81% against 96% of categories right is counted against the known answers.

Words to know

Judge (LLM judge)
A second model that scores answers against a rubric. Example: gpt-oss-120b in the Fireworks run.
Rubric
The scoring guide a judge follows. Example: 5 = specific, right owner, safe; 1 = wrong or unsafe.
Hand-labelled sample
Answers a person has marked right or wrong, used to check a judge. Example: about 20 scores, redone by you.
Bias (judge bias)
A judge’s habit of favouring something that is not correctness. Example: longer, more polished answers.
Go deeper: the engineer version

The kit's question

The judge gives the fine-tune 4.5 and the base 3.9. Is that proof?

The kit's answer

No. Check that the judge agrees with you on a hand-labelled sample first. Judges favour length and style.

More detail: The lesson also says to change the RUBRIC prompt and watch the scores move, which shows why judges need calibrating. evaluate.py prints only the average (judge_1to5), so to hand-check, print each judge_score() result next to its ticket. In the practice run the stand-in judge is rigged: it gives 4 or 5 when the next step names the ticket’s true category and 1 to 3 otherwise (felab/mock_server.py), so its scores track category accuracy. A real judge has no such anchor.

How did you do?

QuestionStep 1 · Eval harness

After fine-tuning, accuracy went up 17 points: 17 more tickets in every 100 sorted correctly. But p95 latency doubled: the time within which 95 of 100 replies come back is now twice as long. What do you recommend?

Show answerHide answer

In plain words

Recommend it only if the slower replies still meet the customer’s speed promise (their SLO, service level objective). Show them both numbers so they choose. Then offer ways to keep the accuracy without the wait: fine-tune a smaller starting model, or add a small helper model that drafts the next few words for the big one to check (speculative decoding).

Picture it

A delivery service that gets more parcels to the right door but takes twice as long. For flowers that must arrive today, that is a no; for a monthly supplies order, it is fine.

With real numberslesson 09’s example table

  • The question’s trade: 17 more tickets right in every 100, but the time within which 95 of 100 replies come back is twice as long.
  • For scale, the eval lesson’s example table has p95 figures of 310 ms and 640 ms: 640 ÷ 310 = 2.1, about a third of a second against about two thirds.
  • In that same table the small fine-tuned model is the fast one: 96% right, 310 ms, $0.02 per 1,000 tickets, against the big model’s 81%, 640 ms, $0.42. That is why a smaller starting model is one of the answers.
  • The lesson’s own example is harsher, 3 times the wait: that can be the wrong call for a chat product and the right one for batch work that runs overnight (like Day 6’s Batch API).

Words to know

SLO (service level objective)
The speed promise. Example: 95 of 100 replies back within a set time.
p95 latency
The time within which 95 of 100 replies come back. Example: 640 ms against 310 ms.
Speculative decoding
A small draft model guesses several tokens and the big model checks them all in one pass; the output does not change. Example: Day 9.
Base model
The model you start from, before any fine-tuning. Example: a smaller one can be faster and cheaper.
Go deeper: the engineer version

The kit's question

Accuracy went up 17 points but p95 doubled. What do you recommend?

The kit's answer

It depends on the SLO. Put both in the memo, and consider a smaller base model or speculative decoding.

More detail: The p95 here is whole-request latency: evaluate.py times each call from sending it to the full reply, not time to first token. Speculative decoding (Day 9) cuts the time per token when a small draft model’s guesses are accepted. A smaller base model cuts time and cost together, and the lesson’s table shows a small fine-tune beating a big general model on all three columns.

How did you do?

QuestionStep 2 · LoRA on Fireworks

The customer wants 12 fine-tuned versions of a model, one for each business unit. Served separately, each one needs its own reserved GPU (the chip that runs the model), billed by the hour. How do you keep serving costs sensible?

Show answerHide answer

In plain words

Share one GPU. Keep one copy of the base model running and load the right small add-on (adapter) for each request. This is multi-LoRA: one bill instead of 12.

Picture it

One kitchen with one set of equipment and 12 spice packets, one per restaurant brand, instead of 12 separate kitchens. Each order says which packet to use.

With real numberslesson 10’s deployment rate and the lab book’s prices

  • One dedicated GPU: about $8 an hour, charged whether or not anyone uses it.
  • Running all month (about 730 hours): 730 x $8 = $5,840.
  • 12 separate deployments: 12 x $5,840 = $70,080 a month.
  • Multi-LoRA, 1 deployment holding 12 adapters: $5,840, one twelfth, as long as one GPU keeps up with the traffic.
  • Each adapter is small: about 30 MB in the Training Techniques video. One deployment holds up to 100 by default, the lab book says.

Words to know

Multi-LoRA
Serving many LoRA add-ons on one running copy of the base model, picking the add-on for each request. Example: 12 business units, 1 GPU.
Adapter
The small add-on LoRA trains, applied on top of the frozen model. Example: about 30 MB.
Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate, used or not. Example: about $8 an hour.
Base model
The model you start from, before any fine-tuning. Example: Qwen2.5 7B Instruct in this lesson.
Go deeper: the engineer version

The kit's question

The customer wants 12 fine-tunes, one per business unit. How do you keep serving costs sane?

The kit's answer

Multi-LoRA: one base deployment with 12 adapters, loaded per request.

More detail: Adapters are loaded per request: each request names its adapter, and the server applies it on top of the shared base weights, batching requests for different adapters together (the field guide: hundreds of fine-tunes on one base at the base model’s price). The lab book adds the Fireworks details: a default quota of 100 adapters per deployment, and multi-LoRA needs a BF16 (16-bit) deployment shape with --enable-addons; FP8 and FP4 (8-bit and 4-bit) shapes cannot host adapters.

How did you do?

QuestionStep 2 · LoRA on Fireworks

A fine-tune gains 17 points of accuracy: 17 more tickets in every 100 sorted correctly. When would you still advise against it?

Show answerHide answer

In plain words

When a cheaper fix gets most of the way there: a better prompt, or a few worked examples inside it. Or when there is too little traffic to pay for a GPU that bills every hour.

Picture it

Hiring a full-time specialist to handle one call a day. First try better instructions for the staff you have; if that closes most of the gap, or the calls are rare, the salary is not worth it.

With real numberslesson 10’s sample, lesson 09’s example table and the lab book’s prices

  • The gain: 80% to about 97% of categories right, 17 points (lesson 10’s sample).
  • The fine-tune needs its own GPU: about $5,840 a month if it stays on (730 hours x $8).
  • Paying per token instead: the big general model in the eval lesson’s table costs $0.42 per 1,000 tickets.
  • Break-even: $0.42 buys 1,000 tickets, so $5,840 buys 5,840 ÷ 0.42 = about 13,900 lots of 1,000: 13.9 million tickets a month, about 5 every second, day and night. Below that, paying per token is cheaper.
  • Say no if traffic is below that volume, or if a better prompt closes most of the 17 points. Or share the GPU with other fine-tunes (multi-LoRA), so each pays less.

Words to know

Few-shot examples
A few worked examples placed in the prompt to show the model what you want, with no training. Example: a cheaper fix to try first.
Serverless (pay-per-token)
Fireworks’ shared, always-on models, billed per token you send and receive. Example: $0.42 per 1,000 tickets.
Break-even
The volume at which two ways of paying cost the same. Example: about 13.9 million tickets a month.
GPU-hour
One GPU for one hour, the unit dedicated deployments are billed in. Example: about $8.
Go deeper: the engineer version

The kit's question

The fine-tune gains 17 points of accuracy. When would you still say no?

The kit's answer

When a prompt change or few-shot examples close most of the gap, or when traffic is too low to justify a dedicated deployment.

More detail: Try the free levers first: the Training Techniques video puts better prompts and retrieval before any training. Then compare the monthly dedicated cost with the pay-per-token bill at the real volume. The lab book’s cost calculator draws that line against one H100 at $8 an hour, 24x7 ($5,840 a month). Multi-LoRA changes the sum by splitting that GPU across adapters. The sum above pairs a big general model’s price with a small fine-tune’s GPU; for a real answer, use the customer’s own candidates.

How did you do?

QuestionStep 2 · LoRA on Fireworks

During training the loss (how wrong the model is on its training examples) keeps falling, but its accuracy on the held-out test tickets stays flat. What does that mean?

Show answerHide answer

In plain words

The model is memorising its training examples instead of learning the task, so it does no better on new tickets. Train less (fewer passes), give it more varied examples, or make the add-on smaller (a lower rank).

Picture it

A student who memorises last year’s exam answers: their practice scores keep rising, but on a new exam they do no better. More varied practice, fewer repeats and a smaller notebook to cram into all help.

With real numbersthe QLoRA On Your Own Box video and lesson 10’s training settings

  • The video’s example, 200 messy, unchecked examples: 97% right on the examples it trained on, 74% on held-out ones. A 23-point gap: memorised.
  • 2,000 cleaner examples: 93% and 89%. A 4-point gap: it learned the task.
  • Today’s settings are already small: 2 passes (epochs) over 800 examples, adapter rank 8. The script suggests 16 to 32 only for harder tasks.
  • The fixes: fewer passes, more varied examples, or a lower rank.

Words to know

Loss (training loss)
A score of how wrong the model is on its training examples; training pushes it down. Example: it can fall while test accuracy stays flat.
Overfitting (memorising)
Training until the model remembers its examples instead of learning the task. Example: 97% on training examples, 74% on new ones.
Epoch
One full pass over the training examples. Example: 2 in 1_train.sh.
Rank (LoRA rank)
The size setting of a LoRA add-on; a higher rank trains more numbers. Example: 8 here, 16 to 32 for harder tasks.
Go deeper: the engineer version

The kit's question

The loss is falling but eval accuracy is flat. What does that mean?

The kit's answer

The model is memorising. Use fewer epochs, more varied data, or a lower rank.

More detail: Falling training loss with flat held-out accuracy is overfitting. 1_train.sh says it in one line: more epochs means memorising, so watch the eval, not the loss. On Day 8 the MLX run prints validation loss (loss on examples kept out of training) every 50 steps: stop when it flattens. A lower rank means fewer trainable numbers, so less room to memorise.

How did you do?

Under 20 seconds each. Record yourself once and listen back.

Step 1 · Eval harness

Question A new model tops the public benchmarks. How would you decide whether it is right for us?

One clear answer

Benchmarks are for launches. For a customer I use three graders on their data and one table: quality, p95 latency and dollars per thousand tasks.

What this means
  • “Benchmarks are for launches”: Public benchmarks (standard tests every model maker reports scores on) help compare models when they come out. They do not measure this customer’s job.
  • “For a customer I use three graders”: I grade each answer three ways, cheapest first: code checks (free and exact), an AI judge with a scoring guide, and a person checking a sample.
  • “on their data”: On the customer’s own examples, kept aside from any training (held out), like the 100 test tickets here.
  • “and one table”: All the results side by side, so nobody chooses a model on one number.
  • “quality, p95 latency and dollars per thousand tasks”: How often it is right, how long 95 of 100 answers take, and what 1,000 tasks cost. The lesson’s example: 81% against 96%, 640 ms against 310 ms, $0.42 against $0.02.

Step 2 · LoRA on Fireworks

Question Fine-tuning sounds expensive. What will it really cost us?

One clear answer

Training was about a dollar. The thing to plan is serving: a LoRA needs dedicated capacity, so we either share it with multi-LoRA or justify it with volume.

What this means
  • “Training was about a dollar”: Teaching the model cost little. This lesson’s 800 examples are about 0.3 million training tokens, about $0.15; the Training Techniques video’s bigger job, 2.4 million tokens, is $1.20.
  • “The thing to plan is serving”: The real cost is running the fine-tuned model for users afterwards.
  • “a LoRA needs dedicated capacity”: On Fireworks a LoRA fine-tune cannot be served pay-per-token. It needs a GPU reserved for it, about $8 an hour whether or not anyone uses it: the lesson’s 38-minute test cost $5.07.
  • “so we either share it with multi-LoRA”: Put many fine-tunes on that one GPU so they split its cost: 12 business units, 1 GPU.
  • “or justify it with volume”: Send enough traffic that the GPU is cheaper than paying per token: about 13.9 million tasks a month, at the eval lesson’s $0.42 per 1,000.

2 worked examples from today’s videos. Work each on paper, then reveal.

From the video: Training Techniques 3:14

Whiteboard 1 of 2

A customer’s support-ticket model (up to 16 billion numbers) sorts 81% of tickets into the right category. You fine-tune it with LoRA on 2,000 examples of about 600 tokens each, going through them twice. How many tokens do you pay for, what does it cost, and what do you do next?

Given, in plain words

Training on a service such as Fireworks bills every token the model reads, on every pass: $0.50 per million tokens for models up to 16 billion numbers. After training, the model sorts 89% correctly. The add-on it produces (the adapter) is about 30 MB.

Reveal the answerHide the answer

Answer · in plain words

About 2.4 million tokens and about $1.20, for 8 more points of accuracy: the video calls it the cheapest 8 points you will ever buy. Next, look at the 11% it still gets wrong, because the kind of mistake decides the next method.

Picture it

A tutor for a student who knows the subject but answers in the wrong format. One short course of 2,000 worked examples, gone over twice, costs less than a coffee and lifts the grade from 81 to 89. The remaining mistakes tell you whether they need more examples, a different kind of teaching, or a textbook.

Careful

The $0.50 per million is the lab book’s reading of Fireworks’ prices on 23 September 2026. Prices move weekly: check the pricing page before you quote one.

Worked answer, step by step

  1. Tokens per pass: 2,000 examples x 600 tokens = 1,200,000 tokens (1.2 million).
  2. Two passes (epochs): 1.2 million x 2 = 2.4 million training tokens.
  3. Cost: 2.4 million x $0.50 per million = $1.20.
  4. Gain: 89% - 81% = 8 more tickets in every 100 sorted correctly, for $1.20.
  5. Still wrong: 100% - 89% = 11%. Read those mistakes. Wrong format or wording: more examples of the right answer (SFT). Right answer but poor judgement: pairs of answers with the better one marked (DPO). An answer a program can check: train against that checker (RFT). Missing facts: look them up and add them to the prompt (retrieval), not training.
  6. Serving: the 30 MB adapter runs on top of a base model that can carry many adapters, so each extra fine-tune adds little to the serving bill.
  7. Day 7’s lesson run is smaller: 800 short examples, 2 passes, about 0.3 million tokens (the script’s estimate), so 0.3 x $0.50 = about $0.15. In the lesson’s sample, 38 minutes of serving cost $5.07, over 30 times more.
Go deeper: the engineer version

The kit's question

2,000 examples × 600 tokens × 2 epochs = 2.4M training tokens

The kit's answer

Make it concrete with a ticket triage model. The base model gets eighty one percent of categories right. Two thousand examples of supervised fine tuning, about two point four million training tokens, costs a little over a dollar on a managed platform and takes the model to eighty nine percent. That is the cheapest eight points you will ever buy.

More detail: The scene’s rows: base model accuracy 81%, after LoRA SFT 89%, training cost about $1.20 at $0.50 per 1M, adapter about 30 MB, serving one base with many adapters; 2.4M x $0.50 = $1.20 exactly. The lab book’s fine-tune calculator starts on the same case (2,000 examples, 600 tokens, 2 epochs, LoRA SFT, up to 16B) and adds 1 hour of dedicated serving at $8, so its total is $9.20: serving is most of it. As LoRA DPO or full-parameter SFT ($1.00 per million at up to 16B) the same job would cost $2.40.

Words to know

Epoch
One full pass over the training examples. Example: 2 here, so every token is billed twice.
Training token
One token read during training; managed training bills every token on every pass. Example: 2.4 million here.
SFT (supervised fine-tuning)
Training on examples of the right answer, to teach format, tone and domain patterns. Example: 2,000 labelled tickets.
Adapter
The small add-on LoRA trains, applied on top of the frozen model. Example: about 30 MB.

From the video: QLoRA On Your Own Box 2:49

Whiteboard 2 of 2

A customer asks whether they can fine-tune an 8-billion-number model on an Apple-silicon laptop, with 2,000 examples. Will it fit in memory, how long will it take and what will it cost?

Given, in plain words

QLoRA keeps the model frozen (unchanged) and stored at 4 bits, half a byte per number. It trains only a small add-on. Full fine-tuning changes every number, so it holds about 16 bytes per number while training. Those are the number itself, a correction for it (its gradient), the training method’s running notes and a spare full-precision copy (the Full Fine-Tuning video, Day 11). Day 1’s rule of thumb for serving, used here too: plan to use at most about 70% of the Mac’s memory.

Reveal the answerHide the answer

Answer · in plain words

Yes. Training peaks at about 11 GB: tight on a 16 GB Mac, where that is right at Day 1’s 70% guide (11.2 GB), and roomy on a 36 GB one. It takes 30 to 50 minutes and costs nothing. Full fine-tuning the same model would need about 128 GB.

Picture it

Full fine-tuning rewrites a whole textbook, with a draft copy and editor’s notes for every page: a warehouse job. QLoRA keeps the textbook closed and writes on sticky notes on top: a desk job.

Careful

Memory is not the hard part; data is. In the same video, 200 messy, unchecked examples reached only 74% on held-out examples, while 2,000 clean ones reached 89%.

Worked answer, step by step

  1. The frozen model at 4 bits: 8 billion x 0.5 bytes = 4 GB. 4-bit files also store small shared scales, so about 5 GB in practice (Day 4’s 4-bit Llama 3.1 8B file was 4.9 GB).
  2. The adapter: about 30 MB, 0.6% of the 5 GB model (30 ÷ 5,000 = 0.006).
  3. Peak memory while training: about 11 GB (the video’s figure). The model and add-on are about 5 GB; the other 11 - 5 = about 6 GB is training’s working space: in-between results plus the add-on’s corrections and notes.
  4. Does it fit? Day 1’s 70% guide on a 16 GB Mac: 0.7 x 16 = 11.2 GB. About 11 GB is right at the limit, so close other apps or train with --batch-size 1. A 36 GB Mac has room (0.7 x 36 = 25.2 GB).
  5. Full fine-tuning instead: 8 billion x 16 bytes = 128 GB before working space. 128 ÷ 11 = about 12 times the QLoRA peak, which is why it needs a cluster (several GPUs working as one).
  6. Time and cost: 300 training steps in 30 to 50 minutes, $0 on a machine you already own. Then fuse (bake) the adapter into the model and serve it with mlx_lm.server (Day 8).
Go deeper: the engineer version

The kit's question

mlx_lm.lora · 2,000 examples · 300 iterations

The kit's answer

Concretely, on an Apple silicon laptop: an eight billion parameter model in four bit is about five gigabytes of weights, the adapter adds tens of megabytes, and training on two thousand examples for three hundred iterations takes well under an hour. Peak memory stays around eleven gigabytes, which is why this works on a machine you already own.

More detail: The scene’s rows: 4-bit base weights about 5 GB, LoRA adapter about 30 MB, peak memory about 11 GB, training time 30 to 50 min, cost $0. Day 8’s lesson trains Qwen2.5 7B at 4 bits with --iters 300 --batch-size 2, about 0.75 of a pass over its 800 examples (300 x 2 ÷ 800), and its memory budget puts QLoRA at about 6 GB against about 120 GB for a full fine-tune of a 7B model. The 16 bytes per parameter are the weights, a same-size gradient, Adam’s two running averages in full precision and a full-precision master copy (the Full Fine-Tuning video).

Words to know

QLoRA
LoRA on a model stored at 4 bits: the frozen model is small and only the add-on trains. Example: about 11 GB at the peak for an 8B model.
Frozen
Not changed during training; only the add-on learns. Example: the 5 GB of 4-bit weights.
Gradient
The correction training works out for each number, saying which way to nudge it. Example: full fine-tuning holds one per number.
Optimizer state
The running notes the training method keeps for every number it updates. Example: part of the 16 bytes per number.

How you will use this

Day 8 trains the same kind of small add-on (a LoRA adapter) on your Mac and compares it with today’s cloud run: your answer to “do it ourselves, or pay a service?”. Day 10’s memo quotes the lowest and highest share of categories right in all your eval results (today’s and Day 8’s) as the before and after. It reads every row, practice runs included, and names only the top one. So on Day 10, check that the before figure is a real model’s, not the practice server’s 75.