Skip to content

Day 7 · Training day, managed

Part 1 · Day 7 of 12

0 of 12 days done

About 3 h 10 About $5 Mac and Fireworks

Long day. Split it: step 1 one evening, step 2 the next.

Today: Build the table that decides between models: how often each is right, how fast it answers, what it costs. Every model takes the same test, sorting customer support tickets (help requests) by type and urgency. Then train your own version on Fireworks (a fine-tune), measure the gain and shut down its GPU (the chip it runs on) before it runs up a bill.

Example

Imagine a support team whose AI sends one in five tickets to the wrong team. In the lesson’s sample, training it on past tickets reduces that error to about three in a hundred. The training costs only cents, but keeping a rented chip running to use the result costs several dollars even for a short test. Today you will measure both the improvement and the full bill before recommending it to a customer.

By the end you’ll have

Your fine-tune’s measured gain and total cost, next to step 1’s scores for the untuned models.

Expected: In the lesson’s sample: 80% to about 97% of ticket categories right (a 17-point gain), for $5.07 of serving plus about 15 cents of training.

The 6 numbers you write down

  1. gpt-oss-120b on Fireworks, untuned: categories right (%), p95 (ms), $ per 1,000 tickets

    The bar a fine-tune must clear, on quality, speed and cost at once, to be worth paying for.

    Example: 81 / 640 / 0.42 (the lesson’s sample big general model, written for gpt-oss-20b; yours will differ)

  2. Judge score for that model (1 to 5)

    How a second model rated the suggested next steps. Trust it only after checking some of its scores by hand.

    Example: 3.9 (the lesson’s big general model)

  3. Your Mac, Ollama and MLX rows: categories right (%) and p95 (ms)

    What a free model on your own machine does on the same task, before any training.

    Example: ollama __ / __; mlx __ / __ (the lesson has no sample for these)

  4. Categories right, base then fine-tune (%)

    The gain your fine-tune bought, measured on tickets it never saw in training.

    Example: 80 / 97 (the lesson’s sample, a 17-point gain; yours land on values like 96.7 or 98.3)

  5. p95 latency, base then fine-tune (ms)

    95 of 100 replies came back within this many milliseconds. A fine-tune should not be much slower than its base. Here the base runs on Fireworks’ shared servers and the fine-tune on its own GPU, so part of any difference is the setup, not the training.

    Example: 420 / 380 (the lesson’s sample)

  6. Total cost: minutes deployed and serving dollars, plus training

    What the whole experiment cost. Serving is nearly all of it; training is about 15 cents.

    Example: 38 min, $5.07 serving + about $0.15 of training (the lesson’s sample)

How you will use this: Day 8 trains the same kind of small add-on (a LoRA adapter) on your Mac and compares it with today’s cloud run: your answer to “do it ourselves, or pay a service?”. Day 10’s memo quotes the lowest and highest share of categories right in all your eval results (today’s and Day 8’s) as the before and after. It reads every row, practice runs included, and names only the top one. So on Day 10, check that the before figure is a real model’s, not the practice server’s 75.

Before you start

  1. About 3 hours 10 minutes, best split over two sittings.

    Step 1 takes about an hour. Step 2 takes about 2 hours, mostly waiting: 10 to 30 minutes of training, then up to an hour with a GPU running. Do step 1 one day and step 2 the next. Run step 2’s last three commands (start the GPU, test, shut down) in one sitting: the GPU bills by the hour in between.

  2. Enough Fireworks credit, and nothing already running.

    Step 2 costs about $5 to $8: a dedicated deployment (a GPU reserved for you) bills about $8 an hour. Check that your prepaid credit covers it. Check too that your monthly spend cap from Day 1 has that much room left: once spending reaches the cap, Fireworks stops answering requests. Then, from the 03-labs folder:

    make fw-check

    The deployments list should be empty. Anything listed bills by the hour: delete it with firectl deployment delete <DEPLOYMENT_ID>.

  3. The practice server.

    Step 1 starts with a free rehearsal on the kit’s practice server (made-up numbers, nothing to download). Start it from 03-labs if it is not already running:

    make mock &

    It prints a line starting mock LLM on http://localhost:9000/v1.

  4. Ollama and MLX from Day 2.

    Step 1 also grades your Mac’s model from Day 2, served two ways, so both servers must be running. Start Ollama with ollama serve (skip it if the Ollama app is open). Then start MLX in a second terminal, from 03-labs:

    bash lessons/02-three-local-servers/serve_mlx.sh

    Leave both running. make check shows which servers are up.

About 2 minutes. Say your answer out loud, then tap to check it. From Day 6 · Structured output and batch.

Every reply now comes back well-formed (100% valid), but only 78 of every 100 put the ticket in the right category (78% accuracy). What do you do next?Show answerHide

In plain words

Work on the answers, not the format. Improve the instructions or add a few worked examples to the prompt first; if that is not enough, fine-tune the model. The format problem is already solved.

Picture it

The online form can now always be filed, but some people pick the wrong option. Clearer instructions on the form help first; if people keep getting it wrong, you train them.

With real numberslesson 07’s practice table and the Training Techniques video (Day 7)

  • In the practice run, 39 of 50 replies are right (78%) and 11 are wrong, all perfectly formatted.
  • Enforcement has nothing left to fix: json_object and json_schema both score 78%.
  • Cheapest fix first: a better prompt, which the Training Techniques video (Day 7) calls free.
  • Then train. The Training Techniques video (Day 7) works an example with a different ticket-sorting model: 2,000 solved tickets of about 600 tokens each, read twice in training, so 2,000 x 600 x 2 = 2.4 million tokens.
  • At $0.50 per million training tokens: 2.4 x $0.50 = about $1.20. That model goes from 81% to 89% right.

Words to know

Accuracy
The share of replies with the right answer, checked against each ticket’s known label. Example: 78% in the practice run.
Few-shot examples
A few solved examples placed in the prompt to show the model what a good answer looks like. Example: tickets shown with their correct JSON before the real one.
Fine-tune
Further training of an existing model on your own examples. Example: Days 7 and 8 train one on these tickets.
Label
The known right answer stored with each test example. Example: each test ticket’s category and severity.
Go deeper: the engineer version

The kit's question

Validity is 100% but accuracy is 78%. What next?

The kit's answer

Improve the prompt or add few-shot examples, then fine-tune. Constrained decoding has done its part.

More detail: Constrained decoding decides which tokens are allowed, not which allowed token is right: billing is a valid category even when the label is outage. Accuracy is a model problem. Fix the prompt first (clearer instructions, a few labelled examples), then fine-tune on labelled data (the LoRA fine-tunes of Days 7 and 8), and measure each change on the same held-out tickets.

The server already forces every reply to follow the schema (json_schema mode). Why also write the schema into the prompt?Show answerHide

In plain words

So the model plans its answer in the right shape from the start. Forcing alone only blocks wrong tokens as they come, which can push the model into awkward or worse answers.

Picture it

Tell a writer the word limit before they start and they plan to fit it. Cut them off mid-sentence at the limit instead, and the ending comes out garbled.

With real numbersthe kit’s triage schema and lesson 07’s script

  • The schema allows severity 1 to 5 only. A model that never saw the schema may start to write “high”. The server allows only a digit, so the model must pick a number it never planned.
  • next_action may hold at most 120 characters. A model unaware of the limit can run long, and a strict server can cut the text off at the limit.
  • Lesson 07’s script puts the schema in the prompt in all 3 modes, and Day 6’s batch file (step 2) does the same.
  • The cost is small: the schema is 341 characters, a few lines of text, sent with each request.

Words to know

Schema (JSON schema)
The rulebook a JSON reply must follow: which fields, what type each holds, which values are allowed. Example: 4 categories, severity 1 to 5.
Enforce
Make the server block anything that breaks the rules. Example: json_schema mode.
Completion
The text the model writes back. Example: the JSON reply to one ticket.
Go deeper: the engineer version

The kit's question

Why keep the schema in the prompt when json_schema already enforces it?

The kit's answer

The model writes better content when it knows the shape. Enforcement alone can force awkward completions.

More detail: Constrained decoding only removes tokens; it does not tell the model what shape it is heading for. A model that has read the schema already puts its choices on valid continuations, so the mask rarely has to override it. Without the schema in the prompt, the mask may force a digit where the model meant a word, or close a string at its length limit. The lab book passes this on as a Fireworks docs rule: put the schema in the prompt too. The lesson’s script does it in every mode, with the comment “the model should know the shape it is being held to, not just be forced into it”.

Some models write out their reasoning before giving a final answer (reasoning models). How do you get well-formed replies from one?Show answerHide

In plain words

Write the schema into the prompt, but do not make the server enforce it. Let the model think freely, then check its final answer with code afterwards.

Picture it

An exam that allows scrap paper: students work things out in their own words, then fill in the answer grid. Forcing their working into the grid’s boxes would wreck it, so you mark the grid afterwards.

With real numberslesson 07’s code and the lab book’s structured-output card

  • First: the schema goes in the prompt, as the lesson’s script does in all 3 of its modes.
  • Then: leave out response_format, the enforcement setting, so nothing blocks tokens while the model reasons. The lab book passes this on as Fireworks’ own advice.
  • Next: check the final answer with the lesson’s code: does it read as JSON, and does it follow the schema (one of 4 categories, severity 1 to 5)?
  • Finally: report the share that passes on 50 tickets, like the other modes, so the numbers compare.

Words to know

Reasoning model
A model that writes out its thinking before its final answer. Some servers return that thinking separately.
response_format
The request setting that turns on enforcement: json_object or json_schema. Leaving it out means asking in the prompt only.
Validate
Check a reply with code after it arrives. Example: the lesson’s “does it parse” and “does it fit the schema” checks.
Go deeper: the engineer version

The kit's question

What do you do with a reasoning model?

The kit's answer

Put the schema in the prompt, let it reason freely, and validate afterwards.

More detail: Constrained decoding restricts every token the server generates. The lesson’s script notes that some reasoning models return their thinking separately and that constrained decoding can interfere with it, so the documented pattern is schema in the prompt, and validate after. Validating afterwards keeps the result measurable: the same parse and schema checks give a rate you can quote. The kit’s timing code also reads a reasoning model’s separate reasoning_content stream (felab/measure.py), because the user waits for it.

Start step 1: Eval harness (1 h)

Your progress is saved on this device.