RLHF
file 17 · 4:10 · subtitles burned in
Chapters
In this video Why models are trained on people’s comparisons, the three stages of RLHF (reinforcement learning from human feedback), and the leash that limits a model gaming its scorer.
What it explains: reward model, PPO, KL leash, reward hacking, GRPO/RFT
3 key points
RLHF learns from comparisons, for tasks with no single right answer.
People struggle to write the perfect reply, but easily pick the better of two. A reward model turns those picks into scores: 2.1 for an apology with a fix, -0.4 for a curt reply, so about a 92% chance a person prefers the first.
Three stages: train on examples first (SFT), then a reward model, then a loop (PPO) that chases the score on a leash.
Unleashed, the model games the scorer with flattery, padding and repeated phrases: reward hacking. The KL penalty charges it for drifting from the starting model, and β (beta) sets the leash length.
Powerful but heavy: 4 models in memory at once, against 1 for ordinary fine-tuning, and many settings to tune.
The 4 are the model being trained, a frozen copy, the reward model and a value model. Relatives: GRPO (used for many reasoning models) compares a group of answers with each other and drops the value model; Fireworks’ RFT (reinforcement fine-tuning) scores each answer with a program, such as tests passing.
Used in the course
The title card in the video says “Video 17 of 18”: that is the file order. The course plays the videos in the order of its days.