QuestionStep 1 · Structured output
Every reply now comes back well-formed (100% valid), but only 78 of every 100 put the ticket in the right category (78% accuracy). What do you do next?
Show answerHide answer
In plain words
Work on the answers, not the format. Improve the instructions or add a few worked examples to the prompt first; if that is not enough, fine-tune the model. The format problem is already solved.
Picture it
The online form can now always be filed, but some people pick the wrong option. Clearer instructions on the form help first; if people keep getting it wrong, you train them.
With real numberslesson 07’s practice table and the Training Techniques video (Day 7)
- In the practice run, 39 of 50 replies are right (78%) and 11 are wrong, all perfectly formatted.
- Enforcement has nothing left to fix:
json_objectandjson_schemaboth score 78%. - Cheapest fix first: a better prompt, which the Training Techniques video (Day 7) calls free.
- Then train. The Training Techniques video (Day 7) works an example with a different ticket-sorting model: 2,000 solved tickets of about 600 tokens each, read twice in training, so 2,000 x 600 x 2 = 2.4 million tokens.
- At $0.50 per million training tokens: 2.4 x $0.50 = about $1.20. That model goes from 81% to 89% right.
Words to know
- Accuracy
- The share of replies with the right answer, checked against each ticket’s known label. Example: 78% in the practice run.
- Few-shot examples
- A few solved examples placed in the prompt to show the model what a good answer looks like. Example: tickets shown with their correct JSON before the real one.
- Fine-tune
- Further training of an existing model on your own examples. Example: Days 7 and 8 train one on these tickets.
- Label
- The known right answer stored with each test example. Example: each test ticket’s category and severity.
Go deeper: the engineer version
The kit's question
Validity is 100% but accuracy is 78%. What next?
The kit's answer
Improve the prompt or add few-shot examples, then fine-tune. Constrained decoding has done its part.
More detail: Constrained decoding decides which tokens are allowed, not which allowed token is right: billing is a valid category even when the label is outage. Accuracy is a model problem. Fix the prompt first (clearer instructions, a few labelled examples), then fine-tune on labelled data (the LoRA fine-tunes of Days 7 and 8), and measure each change on the same held-out tickets.
How did you do?