Structured output
course 9 of 22
lesson 07 · lessons/07-structured-output
30 min About $0.02 Mac, then Fireworks
What you will do and why
Business software reads a model’s replies with code, and one badly formed reply can crash it. You try three ways of asking for the format, measure how often replies come back well-formed, and see that well-formed is not the same as right.
Why it matters: Sorting 50 support tickets on the practice server: asking nicely gave 43 readable replies. Making the server enforce the format gave 50 of 50, yet only 39 had the right category.
You are done when: Your Mac and Fireworks runs each printed a table, and for each you wrote down the json_schema row’s well-formed share and right-category share. Expect about 100% well-formed on your Mac. If Fireworks shows less, or a server rejected a mode, you wrote that down instead.
No video covers how the server forces a reply’s format. “In plain words” below is the explanation.
In plain words
Section titled “In plain words”The server can guarantee a reply’s format, but not that the answer is right. Business software reads a model’s reply with code, not with eyes; if the reply is not exactly in the expected format, the code stops with an error. A schema is the rulebook for that format. The server can enforce it (constrained decoding): the model writes one token (a word piece) at a time, and at every step the server blocks any token that would break the rules.
Picture it
A paper form lets people write anything in any box, so some come back in a shape the office’s computer cannot file, such as “sometime next week” in a date box. An online form with dropdowns offers only allowed choices, so every form can be filed. It still cannot stop someone choosing the wrong option.
With real numbersthe kit’s triage schema (data/triage_schema.json) and lesson 07’s practice-server table (made-up numbers for teaching), 50 tickets for each way of asking
- The task: read one customer support ticket and reply with 3 fields. The schema (the kit calls it the contract) sets them:
categoryis one of 4 words (billing, outage, how-to, abuse);severitya whole number from 1 (least urgent) to 5 (most urgent);next_actiontext of up to 120 characters. Nothing else is allowed. - With enforcement on, the reply must start with
{, so a chattySure! Here is the JSON:is blocked at the very first token. After{"category": "the server allows only tokens that start one of the 4 words. - Asking in the prompt only: 43 of 50 replies could be read as data (86%); 7 came back broken.
- Enforced, in either of the server’s two settings (
json_objectorjson_schema): 50 of 50 readable and following the schema (100%). Onlyjson_schemaguarantees the schema’s fields;json_objectguarantees readable JSON, and the practice server happens to give it the right fields too. - Right category: 39 of 50 enforced (78%), 32 of 50 prompt-only (64%). The 7-ticket gap is the 7 broken replies: a reply the code cannot read counts as wrong.
- Even enforced, 11 of 50 are wrong (22%): well-formed is not the same as right.
Words to know
- JSON
- A plain-text format for data that programs can read: named fields inside curly brackets. Example:
{"category": "billing", "severity": 3}. - Schema (JSON schema)
- The rulebook a JSON reply must follow: which fields, what type each holds, which values are allowed. Example:
data/triage_schema.json. - Constrained decoding
- At every step the server blocks any token that would break the schema, so the reply cannot come out malformed, unless the reply length limit cuts it off. Example:
json_schemamode, 50 of 50 valid. - Parse
- Read text as data. Parsing fails when the text is not valid JSON. Example: the
parses_%column.
From the lesson
lessons/07-structured-output/README.md
What and why
Section titled “What and why”Enterprise code parses the model’s reply. If the reply isn’t valid JSON, the pipeline throws. Constrained decoding fixes this at the source: at every step the server masks out tokens that would break the grammar or schema, so the model cannot produce invalid output. Think of a form with dropdowns instead of free-text boxes.
| mode | guarantee | typical validity |
|---|---|---|
| prompt only (“reply with JSON”) | none | 85–98% |
response_format: json_object |
valid JSON, any keys | ~100% parse, keys can still be wrong |
response_format: json_schema |
valid JSON matching your schema | ~100% |
Validity is not correctness. A perfectly formatted wrong answer is still wrong, and that is the job of the eval harness in lesson 09 (Day 7) and the fine-tunes in lessons 10 and 11 (Days 7 and 8).
Read the code first
Section titled “Read the code first”data/triage_schema.jsonis the contract: an enum category, an integer severity from 1 to 5, and a short string.felab/tickets.py → parse()is deliberately strict. Leniency would hide real integration failures.structured_output.py → request_kwargs()holds the three modes. It is one dict each.
python data/make_tickets.py # once — builds data/tickets/*.jsonlcd lessons/07-structured-outputpython structured_output.py --target mockpython structured_output.py --target llamacpp # llama.cpp turns json_schema into a grammarpython structured_output.py --target fireworksWhat each command does
python data/make_tickets.pyRun it from the
03-labsfolder. It builds the kit’s practice support tickets, the same for everyone: 800 for training, 100 for checking during training, 100 kept back for tests. Day 1’s smoke test already built them; running it again writes identical files. Look for threewrotelines: 800, 100 and 100 rows.python structured_output.py --target mockSends the first 50 test tickets to the practice server three times: asking for JSON in the prompt only, then with
json_object, then withjson_schema. The schema is in the prompt every time. Look for schema_valid_% at 86 in the prompt row and 100 in both enforced rows, and category_acc_% at 64 in the prompt row and 78 in both enforced rows.python structured_output.py --target llamacppThe same 150 requests (50 tickets x 3 ways), now to a real model on your Mac: Day 2’s Llama 3.1 8B, stored at 4 bits per number. Start its llama.cpp server first, as in Before you start. llama.cpp turns the schema into a grammar, a rulebook of allowed text it enforces token by token. Look for the json_schema row at 100% valid, and note its category_acc_%. If one mode prints
server rejected request(for example withBadRequestError), that server does not support it: write that down, it is a finding too. If every mode prints it withAPIConnectionError, the server is not running: see Stuck.python structured_output.py --target fireworksThe same 150 requests on Fireworks (paid, about $0.02), to the model named in your
.envfile (gpt-oss-120b unless you changed it). Look for the json_schema row: its schema_valid_% is the figure you would quote to a customer, with its category_acc_% beside it. gpt-oss-120b writes out its thinking before it answers (a reasoning model, check 3 below). If its rows show low validity or empty replies, that thinking may have used up the script’s 120-token reply limit (max_tokens): write it down as a finding, and compare with the prompt row, which follows check 3’s pattern.
What you should see (mock)
Section titled “What you should see (mock)”How to read it
Each row is one way of asking, 50 tickets each. parses_% is the share that read as JSON; schema_valid_% the share that also follow the schema; category_acc_% the share with the right category; p95_ms the wait 95 of 100 replies stay under, in milliseconds (thousandths of a second). The practice server breaks about 1 in 8 prompt-only replies on purpose (7 of 50 here): it puts Sure! Here is the JSON: in front and cuts off the end. Those replies are longer, and longer replies take longer to write, so that row is slower (536 against 447 ms).
mode n parses_% schema_valid_% category_acc_% p95_msprompt 50 86.0 86.0 64.0 535.8json_object 50 100.0 100.0 78.0 446.7json_schema 50 100.0 100.0 78.0 446.6Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
Every reply now comes back well-formed (100% valid), but only 78 of every 100 put the ticket in the right category (78% accuracy). What do you do next?Show answerHide
In plain words
Work on the answers, not the format. Improve the instructions or add a few worked examples to the prompt first; if that is not enough, fine-tune the model. The format problem is already solved.
Picture it
The online form can now always be filed, but some people pick the wrong option. Clearer instructions on the form help first; if people keep getting it wrong, you train them.
With real numberslesson 07’s practice table and the Training Techniques video (Day 7)
- In the practice run, 39 of 50 replies are right (78%) and 11 are wrong, all perfectly formatted.
- Enforcement has nothing left to fix:
json_objectandjson_schemaboth score 78%. - Cheapest fix first: a better prompt, which the Training Techniques video (Day 7) calls free.
- Then train. The Training Techniques video (Day 7) works an example with a different ticket-sorting model: 2,000 solved tickets of about 600 tokens each, read twice in training, so 2,000 x 600 x 2 = 2.4 million tokens.
- At $0.50 per million training tokens: 2.4 x $0.50 = about $1.20. That model goes from 81% to 89% right.
Words to know
- Accuracy
- The share of replies with the right answer, checked against each ticket’s known label. Example: 78% in the practice run.
- Few-shot examples
- A few solved examples placed in the prompt to show the model what a good answer looks like. Example: tickets shown with their correct JSON before the real one.
- Fine-tune
- Further training of an existing model on your own examples. Example: Days 7 and 8 train one on these tickets.
- Label
- The known right answer stored with each test example. Example: each test ticket’s category and severity.
Go deeper: the engineer version
The kit's question
Validity is 100% but accuracy is 78%. What next?
The kit's answer
Improve the prompt or add few-shot examples, then fine-tune. Constrained decoding has done its part.
More detail: Constrained decoding decides which tokens are allowed, not which allowed token is right: billing is a valid category even when the label is outage. Accuracy is a model problem. Fix the prompt first (clearer instructions, a few labelled examples), then fine-tune on labelled data (the LoRA fine-tunes of Days 7 and 8), and measure each change on the same held-out tickets.
The server already forces every reply to follow the schema (json_schema mode). Why also write the schema into the prompt?Show answerHide
In plain words
So the model plans its answer in the right shape from the start. Forcing alone only blocks wrong tokens as they come, which can push the model into awkward or worse answers.
Picture it
Tell a writer the word limit before they start and they plan to fit it. Cut them off mid-sentence at the limit instead, and the ending comes out garbled.
With real numbersthe kit’s triage schema and lesson 07’s script
- The schema allows
severity1 to 5 only. A model that never saw the schema may start to write “high”. The server allows only a digit, so the model must pick a number it never planned. next_actionmay hold at most 120 characters. A model unaware of the limit can run long, and a strict server can cut the text off at the limit.- Lesson 07’s script puts the schema in the prompt in all 3 modes, and Day 6’s batch file (step 2) does the same.
- The cost is small: the schema is 341 characters, a few lines of text, sent with each request.
Words to know
- Schema (JSON schema)
- The rulebook a JSON reply must follow: which fields, what type each holds, which values are allowed. Example: 4 categories, severity 1 to 5.
- Enforce
- Make the server block anything that breaks the rules. Example:
json_schemamode. - Completion
- The text the model writes back. Example: the JSON reply to one ticket.
Go deeper: the engineer version
The kit's question
Why keep the schema in the prompt when json_schema already enforces it?
The kit's answer
The model writes better content when it knows the shape. Enforcement alone can force awkward completions.
More detail: Constrained decoding only removes tokens; it does not tell the model what shape it is heading for. A model that has read the schema already puts its choices on valid continuations, so the mask rarely has to override it. Without the schema in the prompt, the mask may force a digit where the model meant a word, or close a string at its length limit. The lab book passes this on as a Fireworks docs rule: put the schema in the prompt too. The lesson’s script does it in every mode, with the comment “the model should know the shape it is being held to, not just be forced into it”.
Some models write out their reasoning before giving a final answer (reasoning models). How do you get well-formed replies from one?Show answerHide
In plain words
Write the schema into the prompt, but do not make the server enforce it. Let the model think freely, then check its final answer with code afterwards.
Picture it
An exam that allows scrap paper: students work things out in their own words, then fill in the answer grid. Forcing their working into the grid’s boxes would wreck it, so you mark the grid afterwards.
With real numberslesson 07’s code and the lab book’s structured-output card
- First: the schema goes in the prompt, as the lesson’s script does in all 3 of its modes.
- Then: leave out
response_format, the enforcement setting, so nothing blocks tokens while the model reasons. The lab book passes this on as Fireworks’ own advice. - Next: check the final answer with the lesson’s code: does it read as JSON, and does it follow the schema (one of 4 categories, severity 1 to 5)?
- Finally: report the share that passes on 50 tickets, like the other modes, so the numbers compare.
Words to know
- Reasoning model
- A model that writes out its thinking before its final answer. Some servers return that thinking separately.
- response_format
- The request setting that turns on enforcement:
json_objectorjson_schema. Leaving it out means asking in the prompt only. - Validate
- Check a reply with code after it arrives. Example: the lesson’s “does it parse” and “does it fit the schema” checks.
Go deeper: the engineer version
The kit's question
What do you do with a reasoning model?
The kit's answer
Put the schema in the prompt, let it reason freely, and validate afterwards.
More detail: Constrained decoding restricts every token the server generates. The lesson’s script notes that some reasoning models return their thinking separately and that constrained decoding can interfere with it, so the documented pattern is schema in the prompt, and validate after. Validating afterwards keeps the result measurable: the same parse and schema checks give a rate you can quote. The kit’s timing code also reads a reasoning model’s separate reasoning_content stream (felab/measure.py), because the user waits for it.
Explain what you learned
Section titled “Explain what you learned”Question Our software keeps breaking on the model’s replies. How do you make them reliable?
One clear answer
Schema-constrained decoding for the format, then a validity and accuracy rate on 50 of their real tickets. The first is solved by the platform, the second by the model.
What this means
- “Schema-constrained decoding for the format”: I have the server enforce the customer’s schema (the rulebook for the reply’s shape) one token at a time, so every reply is readable data with the right fields. On the practice server: 50 of 50 well-formed, against 43 of 50 when only asked in the prompt.
- “then a validity and accuracy rate”: Then I measure two separate numbers: the share of replies that are well-formed (validity) and the share that are right (accuracy).
- “on 50 of their real tickets”: Measured on 50 of the customer’s own examples, not on a public test. The lesson rehearses this on 50 made-up practice tickets the model never trained on.
- “The first is solved by the platform”: Format is fixed by a server setting, not by the model: 100% well-formed with
json_schemaon the practice server. - “the second by the model”: Being right depends on the model: a better prompt first, then a fine-tune (further training on their examples, Days 7 and 8). On the practice server, 78% right even with perfect format.
Your numbersSaved on this device and collected in the Day 6 wrap-up.
Hint: From python structured_output.py --target llamacpp: the schema_valid_% column of the prompt row and the json_schema row. Write both, for example “86 / 100”.
Hint: The category_acc_% of the json_schema row in the same run.
Hint: From python structured_output.py --target fireworks: schema_valid_% and category_acc_% of the json_schema row. Write both.
Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.
Done when
Section titled “Done when”Your Mac and Fireworks runs each printed a table, and for each you wrote down the json_schema row’s well-formed share and right-category share. Expect about 100% well-formed on your Mac. If Fireworks shows less, or a server rejected a mode, you wrote that down instead.
Stuck?
Section titled “Stuck?”- Every mode prints
server rejected request (APIConnectionError: Connection error.)and the table says(no rows). - That server is not running.
make check(from the03-labsfolder) shows what is up. Start the practice server withmake mock &, or llama.cpp withbash lessons/02-three-local-servers/serve_llamacpp.shin a second terminal, and wait until it says it is listening. data/tickets/test.jsonl missing.- The ticket data is not built yet. From the
03-labsfolder runpython data/make_tickets.py, then run the script again. No module named 'felab'(or'openai').- The Python environment from Day 1 is off in this terminal. From the
03-labsfolder runsource .venv/bin/activate, then try again. - One mode prints
server rejected requestwith an error such asBadRequestError, while the other modes print rows. - That server does not accept that way of asking, so the script skips the mode and carries on. Write it down: which modes a server supports is a real finding for a customer.
FIREWORKS_API_KEY is not set, or every mode printsAuthenticationErroron Fireworks.- Your key is missing or wrong. Put your key in the
.envfile in03-labs, as on Day 1, or export it in this terminal. - Every mode prints
server rejected request (NotFoundError: ...)on Fireworks. - The model id is not found. Model ids change: copy a current one from the Fireworks model library and pass it with
--model. - Fireworks shows a low schema_valid_% or empty replies, even with the format enforced.
- gpt-oss-120b writes out its thinking before its answer (a reasoning model, check 3). That thinking may use up the script’s 120-token reply limit, so the answer comes back cut off or empty. Write it down as a finding and compare with the prompt row, which follows check 3’s pattern: schema in the prompt, nothing enforced, checked afterwards.
Go deeper: the lab book's Fireworks lab for this step All fixes
Code in this step
Section titled “Code in this step”structured_output.py Python · 89 lines
"""Lesson 07 — structured output that doesn't break.
Enterprise integrations parse the model's reply with code. One stray "Sure! Here'sthe JSON:" and the pipeline throws. Three ways to ask for JSON, from weakest to strongest:
prompt "Reply with JSON only" → the model *usually* complies json_object response_format={"type":"json_object"} → guaranteed JSON, any shape json_schema response_format={"type":"json_schema"} → guaranteed JSON *matching the schema* (constrained decoding: tokens that would break the schema are masked out)
We run N held-out tickets through each mode and report a validity RATE — a numberyou can put in a customer doc, instead of "it seemed fine".
python structured_output.py --target mock --n 50 python structured_output.py --target llamacpp --modes prompt,json_schema python structured_output.py --target fireworks --n 50 # ~ $0.02
Reasoning models: some return their thinking separately and constrained decoding caninterfere with it; the documented pattern is schema in the prompt, and validate after."""import argparseimport jsonimport time
from felab import add_target_args, banner, client, percentile, record, resolve, tablefrom felab.tickets import SCHEMA, SYSTEM, grade, load_test
def request_kwargs(mode: str) -> dict: if mode == "json_object": return {"response_format": {"type": "json_object"}} if mode == "json_schema": return {"response_format": {"type": "json_schema", "json_schema": {"name": "triage", "schema": SCHEMA, "strict": True}}} return {}
def run(cli, model: str, mode: str, rows: list[dict]) -> dict: # The schema goes in the prompt in EVERY mode: the model should know the shape it is # being held to, not just be forced into it. system = SYSTEM + "\nSchema:\n" + json.dumps(SCHEMA) g, lat = [], [] for r in rows: t0 = time.perf_counter() try: resp = cli.chat.completions.create(model=model, temperature=0, max_tokens=120, messages=[{"role": "system", "content": system}, {"role": "user", "content": r["ticket"]}], **request_kwargs(mode)) reply = resp.choices[0].message.content except Exception as e: # some servers reject a mode they don't support — that's a finding too print(f" {mode}: server rejected request ({type(e).__name__}: {str(e)[:80]})") return {"mode": mode, "n": 0} lat.append((time.perf_counter() - t0) * 1000) g.append(grade(reply, r)) n = len(g) return {"mode": mode, "n": n, "parses_%": 100 * sum(x["parses"] for x in g) / n, "schema_valid_%": 100 * sum(x["valid"] for x in g) / n, "category_acc_%": 100 * sum(x["category_ok"] for x in g) / n, "p95_ms": percentile(lat, 95)}
def main() -> None: ap = add_target_args(argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)) ap.add_argument("--n", type=int, default=50, help="tickets per mode") ap.add_argument("--modes", default="prompt,json_object,json_schema") a = ap.parse_args() t = resolve(a) banner(t) cli = client(t) rows = load_test(a.n)
results = [] for mode in a.modes.split(","): print(f" {mode} …", flush=True) r = run(cli, t.model, mode, rows) if r["n"]: results.append(r) record("07-structured", {"target": t.name, "model": t.model.split("/")[-1], **r}) print("\n" + table(results)) print("\nschema_valid_% is the reliability number. category_acc_% is a different question\n" "(is the answer RIGHT?) — constrained decoding fixes the first, not the second. Lesson 09.")
if __name__ == "__main__": main()