QuestionStep 1 · Eval harness
A judge model (a second AI model that grades answers against a scoring guide) gives the fine-tuned model 4.5 out of 5 and the original model 3.9. Is that proof the fine-tune is better?
Show answerHide answer
In plain words
No. First check that the judge agrees with you: score about 20 of the same answers yourself and compare. AI judges tend to reward longer, more polished answers even when they are no more correct.
Picture it
A teacher who gives longer essays in neat handwriting higher marks, whatever they say. Before you trust their grades, you mark 20 of the same essays yourself and see whether you agree.
With real numberslesson 09’s example table and its scoring guide
- The judge scores each suggested next step from 1 to 5: 5 = specific, right owner, safe; 3 = plausible but vague; 1 = wrong or unsafe.
- Average scores: 3.9 for the original model, 4.5 for the fine-tune, a gap of 0.6 points (4.5 - 3.9).
- The check: score about 20 of the same answers by hand. If you and the judge often disagree, the 0.6 means little.
- Code checks need no such trust: 81% against 96% of categories right is counted against the known answers.
Words to know
- Judge (LLM judge)
- A second model that scores answers against a rubric. Example: gpt-oss-120b in the Fireworks run.
- Rubric
- The scoring guide a judge follows. Example: 5 = specific, right owner, safe; 1 = wrong or unsafe.
- Hand-labelled sample
- Answers a person has marked right or wrong, used to check a judge. Example: about 20 scores, redone by you.
- Bias (judge bias)
- A judge’s habit of favouring something that is not correctness. Example: longer, more polished answers.
Go deeper: the engineer version
The kit's question
The judge gives the fine-tune 4.5 and the base 3.9. Is that proof?
The kit's answer
No. Check that the judge agrees with you on a hand-labelled sample first. Judges favour length and style.
More detail: The lesson also says to change the RUBRIC prompt and watch the scores move, which shows why judges need calibrating. evaluate.py prints only the average (judge_1to5), so to hand-check, print each judge_score() result next to its ticket. In the practice run the stand-in judge is rigged: it gives 4 or 5 when the next step names the ticket’s true category and 1 to 3 otherwise (felab/mock_server.py), so its scores track category accuracy. A real judge has no such anchor.
How did you do?