Skip to content

Training Techniques

file 06 · 3:14 · subtitles burned in

Chapters

In this video The four standard ways to change how an existing model behaves, how to choose one from what is going wrong, and why LoRA (training a small add-on instead of the whole model) keeps it cheap.

What it explains: overview: SFT, LoRA, DPO, RFT and how to choose

3 key points

  1. Pick the method from the kind of mistake: examples teach format, comparisons teach taste, an automatic score teaches correctness.

    Can you write the right answer? Supervised fine-tuning (SFT: training on examples of it). Can you only say which of two answers is better? Preference tuning (DPO, direct preference optimization). Can a program score an answer? Reinforcement fine-tuning (RFT). Missing facts need looking up (retrieval), not training.

  2. LoRA (low-rank adaptation) keeps it cheap: the model stays frozen (unchanged) and only a small add-on is trained, so one running model can carry many add-ons.

    The video’s ticket example: 2,000 examples x 600 tokens (word pieces) x 2 passes = 2.4 million training tokens, about $1.20 at $0.50 per million. Accuracy rose from 81% to 89%, and the add-on is about 30 MB, a tiny fraction of the model.

  3. Try a better prompt and retrieval first: they are free.

    Train only when those stop helping. Then look at what is still wrong: the video’s example stops at 89% and asks what the other 11% of mistakes look like, because that picks the next method.

Used in the course

The title card in the video says “Video 6 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (11.4 MB)