Skip to content

Day 8 wrap-up and drill

  1. Overview
  2. Step 1
  3. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

Your build-versus-buy page (COMPARISON.md) is filled in. It shows the same fine-tune on Fireworks (Day 7) and on your Mac (today), side by side. It covers accuracy before and after, reply time, time to a working model, cost and privacy. The last line is your one-sentence recommendation.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 1 · The same LoRA on your Mac

The share of the 60 test tickets whose category (billing, outage, how-to or abuse) the model got right, before and after training. One ticket is worth 1.67 points.

Example: 78.0 → 96.0 (the lesson’s sample; Day 7’s Fireworks sample: 80.0 → 97.0)

The share of test tickets whose urgency score (1 to 5) the model got exactly right.

Example: 66.0 → 94.0 (the lesson’s sample; Day 7’s Fireworks sample: 70.0 → 95.0)

How long a reply took: of the 60 test tickets, sent one at a time, 57 got theirs within this time and 3 took longer. ms means thousandths of a second, so 380 ms is 0.38 seconds. It goes in the p95 latency row of COMPARISON.md, next to Fireworks.

Example: No Mac sample in the lesson; Day 7’s Fireworks sample was 380 ms.

The whole trained add-on: the file you would send, keep versions of and swap. Compare it with the 4.3 GB model.

Example: 7.1M on the sample’s adapter saved: line, about 7.4 MB (the lesson’s sample)

The real time that passes (wall-clock row of COMPARISON.md): how long a customer waits for a working tuned model.

Example: Training alone takes about 15 to 40 minutes, the lesson’s range.

Your build-versus-buy answer, in a sentence a customer can use.

Example: Prototype locally, free and private; ship on Fireworks once it must serve many users.

Every question from today’s steps. Say your answer out loud, then show the answer and rate yourself. 3 cards

QuestionStep 1 · The same LoRA on your Mac

Day 8’s lesson keeps two sets of tickets out of training: 100 validation tickets, checked during training, and 100 test tickets. Why is the final score taken on the test tickets (test.jsonl) and not the validation ones (valid.jsonl)?

Show answerHide answer

In plain words

You used the validation tickets to decide when to stop training, so they helped pick the model and can flatter it. The test tickets played no part in any decision, so their score is honest.

Picture it

A coach decides when the team has practised enough by watching its scores on one practice test. A good score on that same test afterwards proves little. The fair check is a fresh test nobody has looked at.

With real numbersthe kit’s ticket data (data/make_tickets.py) and Day 8’s training script

  • 1,000 tickets, split three ways: 800 to train on, 100 to check progress during training (validation), 100 kept back for the final grade (test).
  • Training checks the validation tickets every 50 steps, and you stop where their loss flattens: in the sample, 0.198 at step 100 and 0.121 at step 300.
  • That stopping point was chosen by looking at the validation tickets, so they are no longer neutral.
  • The test tickets never touch training or the stop decision. No test ticket repeats a training ticket, and 30 of the 100 use wordings training never saw.
  • The final grade uses the first 60 test tickets, so each one is worth 1.67 points (100 ÷ 60).

Words to know

Validation set
Examples kept out of training and checked during it, to decide when to stop. Example: valid.jsonl, 100 tickets.
Test set (held-out set)
Examples kept out of training and out of every decision, used only for the final score. Example: test.jsonl, 100 tickets; the eval uses 60.
Validation loss
A score of how wrong the model still is on the validation set; lower is better. Example: 0.412, 0.198 and 0.121 at steps 50, 100 and 300.
Leak (data leakage)
Data used to judge a model has also shaped it, so the score looks better than it is. Example: the validation tickets, once they set the stopping point.
Go deeper: the engineer version

The kit's question

Why do we evaluate on test.jsonl and not valid.jsonl?

The kit's answer

Validation loss guided when to stop, so it has leaked into the decision. Test is untouched.

More detail: Early stopping on validation loss is a form of model selection, so the validation score is optimistic; the test split stays untouched for the final number. make_tickets.py de-duplicates across all splits and keeps one phrasing per category for the test set only, so the test score measures generalisation, not memory. One caution: the lab book’s standalone snippet builds its validation file with head -50 train.jsonl after copying all of train.jsonl into training, so those 50 rows are also training rows and its validation loss would look better than it is. Day 8’s script uses the separate data/tickets/valid.jsonl.

How did you do?

QuestionStep 1 · The same LoRA on your Mac

When would you tell a customer to fine-tune on their own machines (a laptop, or AI chips called GPUs that they own) rather than on a managed platform such as Fireworks?

Show answerHide answer

In plain words

When their data is not allowed to leave their own machines. Or when they want fast, free experiments first, before paying a managed platform to train the final version and serve it to real users.

Picture it

A chef tests new recipes in their own kitchen: it costs nothing, and the secret sauce stays secret. For a 500-guest wedding, they hire a caterer with the staff and ovens to serve everyone at once.

With real numbersDay 8’s lesson, Day 7’s Fireworks run (lesson sample figures and its summary line), and the lab book’s GPU price

  • On the Mac: $0 to train, and the 800 tickets never leave the machine. Training takes about 15 to 40 minutes.
  • But one Mac serves one person at a time with no uptime promise: COMPARISON.md answers “no” for 100 users at once.
  • On Fireworks (Day 7’s sample): training cost cents to about a dollar, but answering requests (serving) needs a GPU reserved by the hour. 38 minutes cost $5.07: $5.07 ÷ 38 x 60 = about $8 an hour, the lab book’s price.
  • Accuracy is close either way in the samples: category 78% to 96% on the Mac, 80% to 97% on Fireworks.
  • The usual pattern, from COMPARISON.md: try ideas on your own machine first (prototype locally), then run the real service on a managed platform (ship managed).

Words to know

Managed platform
A service that runs everything between the GPUs and the customer’s app: the engine, scaling, routing and the models. Example: Fireworks on Day 7.
Endpoint
A web address that answers requests for one model. Example: http://localhost:8081/v1 for Day 8’s fine-tune.
Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: about $8 an hour.
SLA (service level agreement)
A written promise to a customer about uptime or speed. Example: COMPARISON.md says the Mac has none.
Go deeper: the engineer version

The kit's question

When would you tell a customer to fine-tune locally or on their own GPUs?

The kit's answer

When data can’t leave their environment, or for fast iteration before a managed production run.

More detail: Local or self-hosted training fits data that must stay inside the customer’s environment, and cheap, fast iteration. Managed wins on time to a scalable endpoint, autoscaling and an SLA. On Fireworks a LoRA cannot run on serverless (shared, pay-per-token) models; it needs a dedicated deployment, so the serving hours, not the training, are the cost to plan (the lab book). Day 7’s lesson calls training “a few cents” in its sample and “about a dollar” in its summary line; the QLoRA On Your Own Box video says about a dollar. If data residency is the only blocker, the lab book notes that Fireworks can also run inside the customer’s own cloud account (BYOC), or pin a dedicated deployment to a region at 1.5 times the rate. The QLoRA On Your Own Box video puts it this way: “the local versus managed trade-off is a conversation, not a religion.”

How did you do?

QuestionStep 1 · The same LoRA on your Mac

The add-on you trained (the adapter) is about 7 MB. Why does that small size matter when the model runs as a service?

Show answerHide answer

In plain words

One copy of the base model can keep hundreds of these small add-ons in memory and apply the right one to each request. So every customer can have their own tuned model for about the price of one shared server.

Picture it

One base recipe, a different spice packet per customer (the field guide’s comparison). The kitchen cooks one big pot and adds a small packet to each order, instead of running a separate kitchen for every customer.

With real numbersDay 7’s lesson (12 business units), Day 8’s sample run and memory script, and the lab book’s $8-an-hour GPU price

  • Adapter: 7.1M in Day 8’s sample, about 7.4 MB. Base model: about 4.3 GB, which is 4,300 MB. 4,300 ÷ 7.4 = about 580, so the base is hundreds of times bigger.
  • 100 adapters: 100 x 7.4 MB = 740 MB, about a sixth of one base model (740 ÷ 4,300 = 0.17).
  • Day 7’s example customer wants 12 tuned models, one per business unit (department). Each on its own dedicated GPU at $8 an hour: 12 x $8 = $96 an hour.
  • The same 12 as adapters on one shared base: about $8 an hour, if one deployment can carry all their traffic.
  • Running all month (730 hours): $8 x 730 = $5,840 for one deployment, against 12 x $5,840 = $70,080 for twelve.

Words to know

Adapter (LoRA adapter)
The small set of extra numbers LoRA trains and saves; the base model stays unchanged. Example: about 7 MB in Day 8’s sample.
Base model
The model a fine-tune starts from and leaves unchanged. Example: Qwen2.5 7B Instruct at 4 bits, about 4.3 GB.
Multi-LoRA
One server keeps one base model and many adapters, applying the right adapter to each request. Example: 12 business units on one deployment.
Tenant
One customer, or business unit, sharing a platform with others. Example: per-tenant fine-tunes, one adapter each.
Go deeper: the engineer version

The kit's question

The adapter is 7 MB. Why does that matter for serving?

The kit's answer

One base model can hold hundreds of adapters in memory, which makes per-tenant fine-tunes affordable.

More detail: Multi-LoRA servers keep one base in GPU memory and apply the right adapter per request, batching requests for different adapters together; the field guide says Fireworks serves hundreds of fine-tunes on one base model at the base model’s price. Fusing the adapter into the weights, as Day 8’s 2_fuse_and_serve.sh does for simplicity, gives that up: each fused model is a full copy. The size is the script’s estimate: 8 layers x 4 grids x 2 x 3,584 x rank 8 = 1,835,008 numbers, 3.7 MB at 2 bytes each. It treats all four grids as 3,584 x 3,584; Qwen2.5 7B’s key and value grids are 3,584 x 512 (lesson 17’s peft_params.py), which gives 1,441,792, and mlx-lm’s own defaults decide which grids get an add-on. Lesson 11’s README says about 4 MB, its sample prints 7.1M, and it does not say why they differ. du -sh adapters measures the whole folder, which also keeps copies saved during training, so quote your own ls -lh adapters/adapters.safetensors figure.

How did you do?

Under 20 seconds each. Record yourself once and listen back.

Step 1 · The same LoRA on your Mac

Question Should we fine-tune on our own hardware, or pay for a managed platform?

One clear answer

I ran the same LoRA locally and managed. Local is free and private; managed got me to a scalable endpoint in minutes. The adapter is a few megabytes, so multi-LoRA makes per-customer tuning economical.

What this means
  • “I ran the same LoRA locally and managed”: I trained the same small add-on (a LoRA adapter) on the same 800 support tickets twice. Once on my own Mac, once on Fireworks, a managed platform that runs the hardware for you.
  • “Local is free and private”: Training on my Mac cost $0, and the tickets never left it.
  • “managed got me to a scalable endpoint in minutes”: Fireworks put the add-on behind a live web address other programs can call (an endpoint). It was ready within minutes, and it can add copies as more people use it. My Mac serves one person at a time.
  • “The adapter is a few megabytes”: The trained add-on is about 7 MB in the lesson’s sample, against about 4.3 GB for the model it sits on.
  • “so multi-LoRA makes per-customer tuning economical”: Multi-LoRA: one server keeps one copy of the model and many add-ons, one per customer, and uses the right add-on for each request. Twelve tuned models then cost about one rented GPU, $8 an hour, instead of twelve ($96 an hour).

A worked example from today’s videos. Work it on paper, then reveal.

From the video: QLoRA On Your Own Box 2:49

Whiteboard 1 of 1

Two fine-tunes of the same model. One learned from 200 messy examples, the other from 2,000 clean ones. On their own training examples they score 97% and 93%. Which one do you ship, and why?

Given, in plain words

Training accuracy is how often the model is right on the examples it learned from. Held-out accuracy is how often it is right on examples it never saw, set aside before training: 74% for the messy run, 89% for the clean run.

Reveal the answerHide the answer

Answer · in plain words

Ship the one trained on 2,000 clean examples. It scores a little lower on its own homework but much higher on new questions, and new questions are what customers send.

Picture it

One student memorises last year’s answer sheet and aces the practice quiz, then stumbles on the real exam. Another learns the material: slightly lower on the quiz, far higher on the exam.

Careful

Set the held-out examples aside before training starts. That is why today’s final grade uses test.jsonl, which training never touches.

Worked answer, step by step

  1. Messy run: 97% on its training examples, 74% on new ones. Gap: 97 - 74 = 23 points.
  2. Clean run: 93% on its training examples, 89% on new ones. Gap: 93 - 89 = 4 points.
  3. A big gap means the model memorised its examples instead of learning the task: that is overfitting.
  4. On new examples the clean run wins by 89 - 74 = 15 points, though it looks 4 points worse on its own examples (97 - 93).
  5. Per 100 new examples: 100 - 89 = 11 wrong instead of 100 - 74 = 26, less than half the mistakes.
  6. So judge a fine-tune only on held-out examples, and stop training when validation loss stops improving.
Go deeper: the engineer version

The kit's question

Data quality decides it · 200 noisy examples · 2,000 clean examples

The kit's answer

What matters is the data, not the hardware. Two hundred examples overfit fast, and the model got worse on anything slightly different. Two thousand cleaner examples, with a held out set to tell you when to stop, gave the improvement. The most common fine tuning failure is not compute, it is training on data nobody checked. Hold out the eval set before you start; stop when validation loss stops improving.

More detail: Rising training accuracy with falling held-out accuracy is the overfitting signature, and a LoRA on a small, noisy set memorises fast. Today’s lesson has all three defences: a held-out split made before training (make_tickets.py), validation loss reported every 50 steps (--steps-per-eval 50) so you can stop by hand, and a final score on the untouched test split. The video does not say what made the 200 examples noisy; its point is to check the data before you train.

Words to know

Overfitting
Learning the training examples by heart instead of the task, so new examples go worse. Example: 97% on training examples, 74% on new ones.
Held-out accuracy
How often the model is right on examples set aside before training and never trained on. Example: 74% for the messy run, 89% for the clean run.
Training accuracy
How often the model is right on the examples it learned from. Example: 97% for the messy run.
Generalise
Do well on new examples, not only the ones trained on. Example: 89% held out for the clean run.

How you will use this

When a team asks whether to train locally or use a managed service, you can compare the same ticket task both ways. Day 7 gives you the cloud result and bill; today adds the Mac result and its privacy tradeoff. Use both to recommend a route instead of relying on a product claim.