Skip to content

Supervised Fine-Tuning

file 15 · 4:02 · subtitles burned in

Chapters

In this video What training-by-example data looks like, why only the answer is graded, how much data you need, and how to read the training curves.

What it explains: data format, loss masking, data volume, loss curves

3 key points

  1. Supervised fine-tuning (SFT) teaches by example: a prompt in, the ideal answer out, and only the answer is graded.

    Each ticket example holds a system message (the standing instructions), the customer’s ticket and the ideal JSON answer. The trainer skips the first two when scoring (masking the prompt); run_variants.sh turns this on with --mask-prompt.

  2. A few hundred clean examples beat many noisy ones.

    Typical ranges for ticket triage with LoRA on a 7B model: 50 examples mostly fix the format, 500 add 8 to 12 points of accuracy, 2,000 add 12 to 16, then it flattens. 20,000 noisy ones often do worse than 2,000 clean. Today’s data: 800 training tickets.

  3. Read the two loss curves together: training loss (how wrong it is on the examples it learns from) and validation loss (how wrong on examples held back).

    Both falling: it is learning. Training falling while validation rises: it is memorising, so stop earlier or add data. Both flat and high: check the data, or raise the learning rate (the size of each nudge). The video’s usual starting values: about 0.0001 (written 1e-4) for LoRA, 10 times smaller for full.

The title card in the video says “Video 15 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (15.2 MB)