QuestionStep 1 · The same LoRA on your Mac
Day 8’s lesson keeps two sets of tickets out of training: 100 validation tickets, checked during training, and 100 test tickets. Why is the final score taken on the test tickets (test.jsonl) and not the validation ones (valid.jsonl)?
Show answerHide answer
In plain words
You used the validation tickets to decide when to stop training, so they helped pick the model and can flatter it. The test tickets played no part in any decision, so their score is honest.
Picture it
A coach decides when the team has practised enough by watching its scores on one practice test. A good score on that same test afterwards proves little. The fair check is a fresh test nobody has looked at.
With real numbersthe kit’s ticket data (data/make_tickets.py) and Day 8’s training script
- 1,000 tickets, split three ways: 800 to train on, 100 to check progress during training (validation), 100 kept back for the final grade (test).
- Training checks the validation tickets every 50 steps, and you stop where their loss flattens: in the sample, 0.198 at step 100 and 0.121 at step 300.
- That stopping point was chosen by looking at the validation tickets, so they are no longer neutral.
- The test tickets never touch training or the stop decision. No test ticket repeats a training ticket, and 30 of the 100 use wordings training never saw.
- The final grade uses the first 60 test tickets, so each one is worth 1.67 points (100 ÷ 60).
Words to know
- Validation set
- Examples kept out of training and checked during it, to decide when to stop. Example:
valid.jsonl, 100 tickets. - Test set (held-out set)
- Examples kept out of training and out of every decision, used only for the final score. Example:
test.jsonl, 100 tickets; the eval uses 60. - Validation loss
- A score of how wrong the model still is on the validation set; lower is better. Example: 0.412, 0.198 and 0.121 at steps 50, 100 and 300.
- Leak (data leakage)
- Data used to judge a model has also shaped it, so the score looks better than it is. Example: the validation tickets, once they set the stopping point.
Go deeper: the engineer version
The kit's question
Why do we evaluate on test.jsonl and not valid.jsonl?
The kit's answer
Validation loss guided when to stop, so it has leaked into the decision. Test is untouched.
More detail: Early stopping on validation loss is a form of model selection, so the validation score is optimistic; the test split stays untouched for the final number. make_tickets.py de-duplicates across all splits and keeps one phrasing per category for the test set only, so the test score measures generalisation, not memory. One caution: the lab book’s standalone snippet builds its validation file with head -50 train.jsonl after copying all of train.jsonl into training, so those 50 rows are also training rows and its validation loss would look better than it is. Day 8’s script uses the separate data/tickets/valid.jsonl.
How did you do?