Skip to content

Day 6 wrap-up and drill

  1. Overview
  2. Step 1
  3. Step 2
  4. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

You have the well-formed share you would quote to a customer (the json_schema row, about 100% on your Mac) and the right-category share beside it, for your Mac and for Fireworks. Your 100 test tickets are scored through the Batch API, or submitted with the job name written down.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 1 · Structured output

How often the real model on your Mac follows the schema when only asked, and when the server enforces it. Only the second is guaranteed.

Example: 86 / 100 on the practice server

How often a well-formed reply is also right. Format and correctness are two separate numbers.

Example: 78 on the practice server

The figure you would quote to a customer: the share of replies their code can rely on, measured on 50 test tickets the model never trained on, with how often they are right beside it. Day 10’s memo reads your latest run from results/07-structured.csv.

Example: 100 / 78 on the practice server

Step 2 · Batch API

Proves the loop works end to end before you pay: file in, answers matched back by label, scored.

Example: 100 of 100, 78%

Lets you check or finish the job later with firectl batch-inference-job get <job name>.

Example: triage-batch- plus month, day, hour and minute

A real model’s score on the 100 test tickets, which it never trained on, bought at about half the usual price.

Example: 100 / 100 / 78 (the practice run’s figures; yours come from a real model)

Every question from today’s steps. Say your answer out loud, then show the answer and rate yourself. 5 cards

QuestionStep 1 · Structured output

Every reply now comes back well-formed (100% valid), but only 78 of every 100 put the ticket in the right category (78% accuracy). What do you do next?

Show answerHide answer

In plain words

Work on the answers, not the format. Improve the instructions or add a few worked examples to the prompt first; if that is not enough, fine-tune the model. The format problem is already solved.

Picture it

The online form can now always be filed, but some people pick the wrong option. Clearer instructions on the form help first; if people keep getting it wrong, you train them.

With real numberslesson 07’s practice table and the Training Techniques video (Day 7)

  • In the practice run, 39 of 50 replies are right (78%) and 11 are wrong, all perfectly formatted.
  • Enforcement has nothing left to fix: json_object and json_schema both score 78%.
  • Cheapest fix first: a better prompt, which the Training Techniques video (Day 7) calls free.
  • Then train. The Training Techniques video (Day 7) works an example with a different ticket-sorting model: 2,000 solved tickets of about 600 tokens each, read twice in training, so 2,000 x 600 x 2 = 2.4 million tokens.
  • At $0.50 per million training tokens: 2.4 x $0.50 = about $1.20. That model goes from 81% to 89% right.

Words to know

Accuracy
The share of replies with the right answer, checked against each ticket’s known label. Example: 78% in the practice run.
Few-shot examples
A few solved examples placed in the prompt to show the model what a good answer looks like. Example: tickets shown with their correct JSON before the real one.
Fine-tune
Further training of an existing model on your own examples. Example: Days 7 and 8 train one on these tickets.
Label
The known right answer stored with each test example. Example: each test ticket’s category and severity.
Go deeper: the engineer version

The kit's question

Validity is 100% but accuracy is 78%. What next?

The kit's answer

Improve the prompt or add few-shot examples, then fine-tune. Constrained decoding has done its part.

More detail: Constrained decoding decides which tokens are allowed, not which allowed token is right: billing is a valid category even when the label is outage. Accuracy is a model problem. Fix the prompt first (clearer instructions, a few labelled examples), then fine-tune on labelled data (the LoRA fine-tunes of Days 7 and 8), and measure each change on the same held-out tickets.

How did you do?

QuestionStep 1 · Structured output

The server already forces every reply to follow the schema (json_schema mode). Why also write the schema into the prompt?

Show answerHide answer

In plain words

So the model plans its answer in the right shape from the start. Forcing alone only blocks wrong tokens as they come, which can push the model into awkward or worse answers.

Picture it

Tell a writer the word limit before they start and they plan to fit it. Cut them off mid-sentence at the limit instead, and the ending comes out garbled.

With real numbersthe kit’s triage schema and lesson 07’s script

  • The schema allows severity 1 to 5 only. A model that never saw the schema may start to write “high”. The server allows only a digit, so the model must pick a number it never planned.
  • next_action may hold at most 120 characters. A model unaware of the limit can run long, and a strict server can cut the text off at the limit.
  • Lesson 07’s script puts the schema in the prompt in all 3 modes, and Day 6’s batch file (step 2) does the same.
  • The cost is small: the schema is 341 characters, a few lines of text, sent with each request.

Words to know

Schema (JSON schema)
The rulebook a JSON reply must follow: which fields, what type each holds, which values are allowed. Example: 4 categories, severity 1 to 5.
Enforce
Make the server block anything that breaks the rules. Example: json_schema mode.
Completion
The text the model writes back. Example: the JSON reply to one ticket.
Go deeper: the engineer version

The kit's question

Why keep the schema in the prompt when json_schema already enforces it?

The kit's answer

The model writes better content when it knows the shape. Enforcement alone can force awkward completions.

More detail: Constrained decoding only removes tokens; it does not tell the model what shape it is heading for. A model that has read the schema already puts its choices on valid continuations, so the mask rarely has to override it. Without the schema in the prompt, the mask may force a digit where the model meant a word, or close a string at its length limit. The lab book passes this on as a Fireworks docs rule: put the schema in the prompt too. The lesson’s script does it in every mode, with the comment “the model should know the shape it is being held to, not just be forced into it”.

How did you do?

QuestionStep 1 · Structured output

Some models write out their reasoning before giving a final answer (reasoning models). How do you get well-formed replies from one?

Show answerHide answer

In plain words

Write the schema into the prompt, but do not make the server enforce it. Let the model think freely, then check its final answer with code afterwards.

Picture it

An exam that allows scrap paper: students work things out in their own words, then fill in the answer grid. Forcing their working into the grid’s boxes would wreck it, so you mark the grid afterwards.

With real numberslesson 07’s code and the lab book’s structured-output card

  • First: the schema goes in the prompt, as the lesson’s script does in all 3 of its modes.
  • Then: leave out response_format, the enforcement setting, so nothing blocks tokens while the model reasons. The lab book passes this on as Fireworks’ own advice.
  • Next: check the final answer with the lesson’s code: does it read as JSON, and does it follow the schema (one of 4 categories, severity 1 to 5)?
  • Finally: report the share that passes on 50 tickets, like the other modes, so the numbers compare.

Words to know

Reasoning model
A model that writes out its thinking before its final answer. Some servers return that thinking separately.
response_format
The request setting that turns on enforcement: json_object or json_schema. Leaving it out means asking in the prompt only.
Validate
Check a reply with code after it arrives. Example: the lesson’s “does it parse” and “does it fit the schema” checks.
Go deeper: the engineer version

The kit's question

What do you do with a reasoning model?

The kit's answer

Put the schema in the prompt, let it reason freely, and validate afterwards.

More detail: Constrained decoding restricts every token the server generates. The lesson’s script notes that some reasoning models return their thinking separately and that constrained decoding can interfere with it, so the documented pattern is schema in the prompt, and validate after. Validating afterwards keeps the result measurable: the same parse and schema checks give a rate you can quote. The kit’s timing code also reads a reasoning model’s separate reasoning_content stream (felab/measure.py), because the user waits for it.

How did you do?

QuestionStep 2 · Batch API

Which of a customer’s jobs (workloads) would you move to the Batch API, where requests go in as one file and the answers come back later at about half price?

Show answerHide answer

In plain words

Any job where nobody is waiting for the answer: testing a model on a fixed set of examples, reprocessing old records, sorting documents into categories, or generating practice data.

Picture it

A restaurant cooks for the diners at its tables first. Tomorrow’s catering order waits for the quiet hours, when the kitchen is idle and the work costs less.

With real numberslesson 08 and the Platform Map video’s example prices

  • Work where someone waits, such as a chat window, pays for speed: the lesson’s example is an answer in 300 ms (0.3 seconds). It stays on serverless or dedicated.
  • A test of the model on 100 tickets can wait minutes to hours: batch, at about half price.
  • The Platform Map video’s app, 20 million tokens a day: $4.80 on serverless, about $2.40 if it could all go to batch.
  • Discounts for a repeated prompt opening (prompt caching, Day 5) apply on top of the batch price.

Words to know

Back-fill
Running a job over all your existing records at once, for example after adding a new field.
Embedding
A list of numbers that captures a text’s meaning, used for search and matching. Recomputing them for every document is a typical back-fill.
Synthetic data
Examples made by a program or a model instead of collected from real users. Example: the kit’s tickets.
Eval
A repeatable test of a model on a fixed set of prompts. Example: the 100 test tickets.
Go deeper: the engineer version

The kit's question

Which customer workloads would you move to batch?

The kit's answer

Anything where no user is waiting: evals, embeddings back-fills, document classification, synthetic data.

More detail: The test is whether anyone waits, not the size of the job. Batch jobs run when the provider has idle GPUs (submit_batch.sh), so they can take minutes to hours; the lab book lists 12 to 72 hour windows. Prompt-caching discounts stack on top of the batch price. Anything user-facing (chat, agents, code completion) stays on serverless or a dedicated deployment.

How did you do?

QuestionStep 2 · Batch API

In the batch file, every request carries its own label, custom_id (such as t-000). Why does it need one?

Show answerHide answer

In plain words

So each answer can be matched to its question. Results can come back in any order, and failed requests go to a separate error file, so answer number 5 is not always for question number 5.

Picture it

A cloakroom gives each coat a numbered tag. If staff instead handed coats back in line order and one coat had gone missing, everyone behind it would get the wrong coat. The tag makes the order irrelevant.

With real numberslesson 08’s scripts

  • The lesson’s file holds 100 requests, labelled t-000 to t-099.
  • Suppose request t-042 fails. The results then hold 99 answers.
  • Even if the other 99 came back in order, matching by line order would pair every answer after the gap with the wrong ticket: 57 wrong pairings (t-043 to t-099).
  • Matched by custom_id, the 99 are scored correctly and the table says answered 99 of 100.

Words to know

custom_id
The label on each request in a batch file. Example: t-042.
Error file
Where a batch job puts the requests that failed, apart from the results. Example: why answered can be below 100.
JSONL
A file with one JSON record per line. Example: one line per request, one line per result.
Go deeper: the engineer version

The kit's question

Why custom_id?

The kit's answer

Results aren’t ordered, and some requests may fail and land in the error file.

More detail: Batch output is a set of result records, not an ordered response stream: order is not guaranteed and failures are written separately. score_batch.py builds a table of labels keyed t-000 to t-099 and looks each result up by custom_id. Its find_content() then searches the record for the reply text, because the exact output shape can differ between providers and versions.

How did you do?

Under 20 seconds each. Record yourself once and listen back.

Step 1 · Structured output

Question Our software keeps breaking on the model’s replies. How do you make them reliable?

One clear answer

Schema-constrained decoding for the format, then a validity and accuracy rate on 50 of their real tickets. The first is solved by the platform, the second by the model.

What this means
  • “Schema-constrained decoding for the format”: I have the server enforce the customer’s schema (the rulebook for the reply’s shape) one token at a time, so every reply is readable data with the right fields. On the practice server: 50 of 50 well-formed, against 43 of 50 when only asked in the prompt.
  • “then a validity and accuracy rate”: Then I measure two separate numbers: the share of replies that are well-formed (validity) and the share that are right (accuracy).
  • “on 50 of their real tickets”: Measured on 50 of the customer’s own examples, not on a public test. The lesson rehearses this on 50 made-up practice tickets the model never trained on.
  • “The first is solved by the platform”: Format is fixed by a server setting, not by the model: 100% well-formed with json_schema on the practice server.
  • “the second by the model”: Being right depends on the model: a better prompt first, then a fine-tune (further training on their examples, Days 7 and 8). On the practice server, 78% right even with perfect format.

Step 2 · Batch API

Question Testing every change on our own data sounds expensive. How do we afford it?

One clear answer

Evaluation is batch work. At half price you can afford to evaluate every prompt or model change, not just the big ones.

What this means
  • “Evaluation is batch work”: Testing a model on a fixed set of examples has nobody waiting for the answers, so it can go in one batch file and come back later.
  • “At half price”: The Batch API charges about half the normal per-token price (serverless). The Platform Map video’s $4.80 a day would be about $2.40.
  • “you can afford to evaluate every prompt or model change”: The same budget buys twice as many test runs, so every change to the instructions or the model gets tested. Today’s 100 tickets cost well under a cent, by the script’s own estimate.
  • “not just the big ones”: Small tweaks cause surprises too. Testing only the big releases lets those slip through.

2 worked examples from today’s videos. Work each on paper, then reveal.

From the video: The Platform Map 2:51

Whiteboard 1 of 2

A customer’s app uses 20 million tokens (word pieces) a day: 16 million sent in, 4 million written back. On Fireworks serverless, the model (gpt-oss-120b) costs $0.15 per million tokens in and $0.60 per million out. Should they keep paying per token, or rent one dedicated H100 (a top data-centre GPU) at $8 an hour?

Given, in plain words

Serverless bills only the tokens used. A dedicated GPU bills every hour it runs, busy or idle. A month is about 730 hours (365 days x 24 hours ÷ 12 months).

Reveal the answerHide the answer

Answer · in plain words

Stay on serverless: about $146 a month, against $5,840 for one dedicated H100, 40 times more. Move to dedicated only when traffic is big and steady enough to keep the GPU busy.

Picture it

Taxis against leasing a car: if you ride twice a week, taxis cost far less than a lease you pay whether you drive or not. The lease wins only when you are on the road most of the day. One difference: a dedicated GPU also guarantees steady speed, which some customers need whatever the price.

Careful

Price is only half the check. At break-even, the GPU must also serve 800 million tokens a day at the customer’s speed target; only a benchmark like Day 3’s can tell you that. And if none of this traffic needed an instant answer, the Batch API would bill about half: about $2.40 a day.

Worked answer, step by step

  1. Input: 16 million tokens x $0.15 per million = $2.40 a day.
  2. Output: 4 million x $0.60 per million = $2.40 a day. Output costs 4 times more per token, so a quarter of the tokens costs the same.
  3. Per day: $2.40 + $2.40 = $4.80.
  4. Per month: 730 hours ÷ 24 = 30.4 days; $4.80 x 30.4 = $146 (the video says about $150).
  5. One dedicated H100 for a month: $8 x 730 hours = $5,840, used or not.
  6. $5,840 ÷ $146 = 40: the GPU costs 40 times the token bill.
  7. Break-even on price: 40 times the traffic, 20 million x 40 = 800 million tokens a day, with the same 4-to-1 split of tokens sent and written.
  8. Answer: serverless until the traffic justifies the GPU, and say so plainly.
Go deeper: the engineer version

The kit's question

Put money on it. Twenty million tokens a day, four to one input to output, on a mid sized open model at fifteen cents per million in and sixty cents out.

The kit's answer

That is under five dollars a day, about a hundred and fifty a month. One dedicated H one hundred running all month is five thousand eight hundred and forty. The break even is a long way above where most products start.

More detail: On screen: input 16M x $0.15 = $2.40; output 4M x $0.60 = $2.40; per day $4.80; per month about $146; one dedicated H100 $8/hr x 730 h = $5,840; recommendation “Serverless until utilisation justifies the GPU — and say so plainly”. The prices match the lab book’s tables (gpt-oss 120B: $0.15 in, $0.015 cached, $0.60 out per 1M; H100 on-demand $8.00 an hour, billed per GPU-second). 800 million tokens a day is about 9,300 tokens a second around the clock (800,000,000 ÷ 86,400), 80% of them input; whether one H100 sustains that at the customer’s p95 target is a benchmark question, not a pricing one. Cached input and batch pricing push the break-even further out. Fireworks on-demand deployments can scale to zero after an hour idle by default (the lab book), which stops the bill but not the 730-hour sum for a GPU kept up all month.

Words to know

Serverless
Fireworks’ shared, always-on models, billed per token you send and receive.
Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not.
GPU-hour
One GPU rented for one hour. Example: $8 for an H100; a month is about 730 of them.
Break-even
The point where two options cost the same. Example: about 800 million tokens a day here.

From the video: The Platform Map 2:51

Whiteboard 2 of 2

A customer sends all its traffic to a big frontier model at $6.00 per million tokens written. You move the routine 90% (the simple, repetitive requests) to a fine-tuned 8B model (8 billion learned numbers) at $0.20 per million. How much cheaper is the moved traffic, and how much cheaper is the whole bill for tokens written?

Given, in plain words

The video’s example table, with numbers made up to show the shape. Big model: $6.00 per million tokens written, 87% right, 95 of 100 replies within 1.9 seconds. Fine-tuned 8B: $0.20 per million, 91% right, 0.6 seconds, built in one afternoon for about $2. The other 10% of traffic stays on the big model.

Reveal the answerHide the answer

Answer · in plain words

The moved traffic gets 30 times cheaper. The whole bill for tokens written gets about 7.7 times cheaper, because the 10% still on the big model becomes most of what you pay.

Picture it

A firm sends every question to its most expensive senior expert. Handing the routine 90% to a trained specialist makes those answers 30 times cheaper, and often better, but the few still sent to the expert become most of the bill.

Careful

The table is illustrative; the video says its shape matches published case studies. Before quoting a saving, measure the customer’s own accuracy and speed with Day 7’s eval harness. And remember that on Fireworks a fine-tune runs on a dedicated GPU billed by the hour, so its real cost depends on keeping that GPU busy.

Worked answer, step by step

  1. Moved traffic: $6.00 ÷ $0.20 = 30 times cheaper per million tokens written.
  2. Before, for every million tokens written: $6.00.
  3. Assume the routine 90% of traffic also writes 90% of the tokens. After: 10% on the big model, 0.1 x $6.00 = $0.60; 90% on the 8B, 0.9 x $0.20 = $0.18.
  4. Total after: $0.60 + $0.18 = $0.78 per million.
  5. Whole bill: $6.00 ÷ $0.78 = 7.7 times cheaper.
  6. The 10% left behind is $0.60 ÷ $0.78 = 77% of the new bill.
  7. In the same example, quality and speed also improve: 87% to 91% right, 1.9 to 0.6 seconds (3.2 times faster).
Go deeper: the engineer version

The kit's question

Specialisation is the other lever, and it compounds with the first.

The kit's answer

Move the routine ninety percent of traffic to a fine tuned eight billion parameter model and the unit cost falls by roughly thirty times, while quality on that narrow task usually goes up, not down. That is the whole argument for owning your weights, in one table.

More detail: On screen (“Specialise the routine traffic”): $ per 1M output $6.00 against $0.20; task accuracy 87% against 91%; p95 latency 1.9 s against 0.6 s; what it cost: one afternoon plus about $2; note “Illustrative numbers, but the shape matches the published case studies.” The video’s “roughly thirty times” is the unit cost of the moved traffic ($6.00 ÷ $0.20). Blended over all output (assuming equal output length per request): 0.1 x $6.00 + 0.9 x $0.20 = $0.78 per 1M, 7.7 times lower, with 77% of it on the frontier share, so the routing split becomes the next lever. The table prices output tokens only. On Fireworks a LoRA fine-tune is served on a dedicated deployment (the lab book), so its real unit cost depends on keeping that GPU busy: the first whiteboard’s question again.

Words to know

Frontier model
One of the largest, most capable models, usually priced highest. Example: $6.00 per million here.
Fine-tune
Further training of an existing model on your own examples. Example: the 8B model here.
Unit cost
The price of one unit of work, such as a million tokens written. Example: $6.00 against $0.20.
Output tokens
The tokens the model writes back; usually priced higher than input. Example: $0.60 against $0.15 per million.

How you will use this

On Day 7, a test rig (the eval harness) grades the same tickets more ways: a second model acting as judge, speed, and cost per 1,000 tickets. On Day 10, your sizing memo (a written hardware and cost recommendation) quotes your latest two well-formed shares, asked in the prompt and enforced, from results/07-structured.csv. It names batch as the way to halve the cost of tests and bulk re-processing.