Skip to content

RLHF

file 17 · 4:10 · subtitles burned in

Chapters

In this video Why models are trained on people’s comparisons, the three stages of RLHF (reinforcement learning from human feedback), and the leash that limits a model gaming its scorer.

What it explains: reward model, PPO, KL leash, reward hacking, GRPO/RFT

3 key points

  1. RLHF learns from comparisons, for tasks with no single right answer.

    People struggle to write the perfect reply, but easily pick the better of two. A reward model turns those picks into scores: 2.1 for an apology with a fix, -0.4 for a curt reply, so about a 92% chance a person prefers the first.

  2. Three stages: train on examples first (SFT), then a reward model, then a loop (PPO) that chases the score on a leash.

    Unleashed, the model games the scorer with flattery, padding and repeated phrases: reward hacking. The KL penalty charges it for drifting from the starting model, and β (beta) sets the leash length.

  3. Powerful but heavy: 4 models in memory at once, against 1 for ordinary fine-tuning, and many settings to tune.

    The 4 are the model being trained, a frozen copy, the reward model and a value model. Relatives: GRPO (used for many reasoning models) compares a group of answers with each other and drops the value model; Fireworks’ RFT (reinforcement fine-tuning) scores each answer with a program, such as tests passing.

The title card in the video says “Video 17 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (14.8 MB)