Skip to content

Direct Preference Optimization

file 18 · 3:50 · subtitles burned in

Chapters

In this video How DPO gets RLHF’s result from the same pairs with one simple loss, what β does, and how your own product already produces preference pairs.

What it explains: the DPO loss, a worked step, pairs from product data

3 key points

  1. DPO learns straight from chosen and rejected pairs with one simple loss: no reward model, no trial-and-error loop.

    Each step makes the chosen answer more likely and the rejected one less, compared with a frozen copy. In the video’s step the gap reaches 1.2, and with β (the leash setting) at 0.1 the loss falls from 0.69 to 0.63.

  2. It needs 2 models in memory instead of RLHF’s 4, and trains almost as steadily as ordinary fine-tuning.

    The 2 are the model being trained and the frozen reference copy. Pairs often come free from your product: an agent’s edit of a draft is chosen, the original draft rejected.

  3. Fine-tune on examples first, then DPO, and measure both win rate (how often the new answer is preferred to the old) and task accuracy.

    DPO refines a model that already does the task. Start β around 0.1; lower lets the model move further. Improving taste can quietly cost correctness, so check accuracy every time.

The title card in the video says “Video 18 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (14.7 MB)