Supervised Fine-Tuning
file 15 · 4:02 · subtitles burned in
Chapters
In this video What training-by-example data looks like, why only the answer is graded, how much data you need, and how to read the training curves.
What it explains: data format, loss masking, data volume, loss curves
3 key points
Supervised fine-tuning (SFT) teaches by example: a prompt in, the ideal answer out, and only the answer is graded.
Each ticket example holds a system message (the standing instructions), the customer’s ticket and the ideal JSON answer. The trainer skips the first two when scoring (masking the prompt);
run_variants.shturns this on with--mask-prompt.A few hundred clean examples beat many noisy ones.
Typical ranges for ticket triage with LoRA on a 7B model: 50 examples mostly fix the format, 500 add 8 to 12 points of accuracy, 2,000 add 12 to 16, then it flattens. 20,000 noisy ones often do worse than 2,000 clean. Today’s data: 800 training tickets.
Read the two loss curves together: training loss (how wrong it is on the examples it learns from) and validation loss (how wrong on examples held back).
Both falling: it is learning. Training falling while validation rises: it is memorising, so stop earlier or add data. Both flat and high: check the data, or raise the learning rate (the size of each nudge). The video’s usual starting values: about 0.0001 (written 1e-4) for LoRA, 10 times smaller for full.
Used in the course
The title card in the video says “Video 15 of 18”: that is the file order. The course plays the videos in the order of its days.