Direct Preference Optimization
file 18 · 3:50 · subtitles burned in
Chapters
In this video How DPO gets RLHF’s result from the same pairs with one simple loss, what β does, and how your own product already produces preference pairs.
What it explains: the DPO loss, a worked step, pairs from product data
3 key points
DPO learns straight from chosen and rejected pairs with one simple loss: no reward model, no trial-and-error loop.
Each step makes the chosen answer more likely and the rejected one less, compared with a frozen copy. In the video’s step the gap reaches 1.2, and with β (the leash setting) at 0.1 the loss falls from 0.69 to 0.63.
It needs 2 models in memory instead of RLHF’s 4, and trains almost as steadily as ordinary fine-tuning.
The 2 are the model being trained and the frozen reference copy. Pairs often come free from your product: an agent’s edit of a draft is chosen, the original draft rejected.
Fine-tune on examples first, then DPO, and measure both win rate (how often the new answer is preferred to the old) and task accuracy.
DPO refines a model that already does the task. Start β around 0.1; lower lets the model move further. Improving taste can quietly cost correctness, so check accuracy every time.
Used in the course
The title card in the video says “Video 18 of 18”: that is the file order. The course plays the videos in the order of its days.