Skip to content

Training toys, part 2

course 21 of 22

lesson 16 · lessons/16-training-toy

25 min Free Anywhere

What you will do and why

You watch a model learn to game its scorer (reward hacking). The scorer learned from people who could not check the ticket’s category, so the model trained against it stops getting the category right.

Why it matters: In the lesson’s table, as the leash loosens (the model may drift further from how it started), the scorer’s average rating rises from +1.69 to +2.11. The average real value of the answers, a hidden score of what the customer wants, falls from +4.30 to +2.15.

You are done when: Section 5 of your run shows the scorer’s average rating (E[RM]) rising as β shrinks, while the average true value (E[true]) peaks and then falls. In the --expert-labels run, every row beats the imitation step.

RLHF · 4:10

Download mp4 (14.8 MB)

Chapters

In this video Why models are trained on people’s comparisons, the three stages of RLHF (reinforcement learning from human feedback), and the leash that limits a model gaming its scorer.

3 key points

  1. RLHF learns from comparisons, for tasks with no single right answer.

    People struggle to write the perfect reply, but easily pick the better of two. A reward model turns those picks into scores: 2.1 for an apology with a fix, -0.4 for a curt reply, so about a 92% chance a person prefers the first.

  2. Three stages: train on examples first (SFT), then a reward model, then a loop (PPO) that chases the score on a leash.

    Unleashed, the model games the scorer with flattery, padding and repeated phrases: reward hacking. The KL penalty charges it for drifting from the starting model, and β (beta) sets the leash length.

  3. Powerful but heavy: 4 models in memory at once, against 1 for ordinary fine-tuning, and many settings to tune.

    The 4 are the model being trained, a frozen copy, the reward model and a value model. Relatives: GRPO (used for many reasoning models) compares a group of answers with each other and drops the value model; Fireworks’ RFT (reinforcement fine-tuning) scores each answer with a program, such as tests passing.

A reward model is a scorer trained on people’s choices between two answers. It can only learn what those people could judge. Push a model hard against it and the model learns to please the scorer instead of the customer: reward hacking. A leash (the KL penalty: a charge for drifting far from how the model first answered) limits the damage; better labellers fix it.

Picture it

A teacher with no time to check facts grades essays on neatness and brevity. Students learn fast: neat, short essays with wrong facts get top marks. A rule that students may not stray far from how they wrote at the start of term slows this down; only a teacher who checks the facts stops it.

With real numberslesson 16’s toy_alignment.py (one support ticket, six possible answers) and its README table

  • Hidden true value of an answer: +3 for the right category, +1.5 for valid JSON (a text format programs can read), +0.8 for a friendly line, minus 1.5 x its length (0.05 to 0.70).
  • Best answer, JSON plus a friendly line with the right category: 3 + 1.5 - (1.5 x 0.35) + 0.8 = 4.775.
  • The fast labellers cannot check the category, so for them it is worth 0. Their favourite is short friendly JSON with the wrong category: 1.5 - (1.5 x 0.10) + 0.8 = 2.15.
  • Leash β at 0.50 (firm), 0.10, 0.02 (loose), from the lesson’s table: the scorer’s average rating rises +1.69, +1.96, +2.11. The average true value of the model’s answers falls +4.30, +3.45, +2.15.
  • At β 0.02 the model nearly always gives the labellers’ favourite: average true value 2.15, less than half of the best answer’s 4.775.

Words to know

Reward model
A model trained on people’s comparisons to give any answer one score. Example: 2.1 for the apology, -0.4 for the curt reply.
Reward hacking
Raising the score by exploiting the scorer’s blind spots instead of getting better. Example: short friendly JSON with the wrong category.
KL leash (KL penalty)
A charge for drifting away from the starting model, measured by KL (Kullback-Leibler divergence), so training cannot run off after loopholes. Example: the KL column grows from 0.34 to 3.90 in the lesson’s table.
β (beta)
The leash setting: smaller means a longer leash. Example: the toy tries 4 down to 0.01.

From the lesson

lessons/16-training-toy/README.md

Real training runs take minutes to hours, which makes it hard to see what an algorithm does. These two scripts shrink the problem until every number fits on one screen, while keeping the real algorithms intact.

script shrinks lets you see
toy_finetune.py a model → one 64×64 layer full fine-tuning vs LoRA ranks: trainable parameters, optimizer memory, held-out error as the data grows
toy_alignment.py all possible text → 6 answers to one ticket base → SFT → reward model → RLHF (β sweep and a sampled policy-gradient run) → DPO, including reward hacking
  • toy_alignment.py:
    • LABEL_U: fast labellers can’t check the ticket category, which is the realistic blind spot.
    • VIS: what the reward model is allowed to see.
    • π* ∝ π_sft · exp(RM/β): the exact optimum that PPO approximates by sampling.
    • The DPO loop is the loss from step 2’s video (Direct Preference Optimization), written out gradient by gradient.
Terminal window
cd lessons/16-training-toy
python toy_alignment.py # ~2 s
python toy_alignment.py --expert-labels # labellers who CAN check the category
python toy_alignment.py --beta 0.1 # looser leash for the sampled RLHF run and DPO

What each command does

  1. python toy_alignment.py

    Runs the whole story on one support ticket in about 2 seconds. It goes from the starting model to imitation of good agents (SFT), the reward model, RLHF at 8 leash settings, then DPO. Look for section 4’s agreement figure (71%) and section 5’s table: the average rating keeps rising as β shrinks, while the average true value peaks at β 0.50, then falls.

  2. python toy_alignment.py --expert-labels

    The same run with labellers who can check the category, so correctness is worth +3 to them too. The reward model now sees correctness and agrees with them 78% of the time. Look for every row of section 5 beating the imitation step’s average true value (+3.89, the E[true value] on the section 2 line).

  3. python toy_alignment.py --beta 0.1

    Loosens the leash from the default 0.5 to 0.1 for the two training runs: the sampled RLHF run (5b) and DPO (6). The sweep table stays the same. Look for the 5b line: its KL from SFT (0.94) stops well short of the exact optimum KL (1.81). So its 4,000 sampled steps have not yet reached the wrong-category answer the sweep predicts at β 0.10, and 5b still favours the right one. DPO (6) ends close to 5b (KL from SFT 0.99).

How to read it

The first table is Day 11’s. Today’s has one row per leash setting β, from tight to loose: E[RM] is the scorer’s average rating, E[true] the average real value, KL how far the model drifted from its start. The rating climbs row by row; the real value peaks, then falls: reward hacking (your run shows 8 rows, the lesson 3).

method trainable share adam_bytes n=16 n=48 n=200
full 4,096 100.0% 49,152 0.3600 0.1084 0.0002
lora r=1 128 3.1% 1,536 0.3937 0.1951 0.1321 ← rank too small: plateaus
lora r=2 256 6.2% 3,072 0.4024 0.0996 0.0000 ← = full, at 6% of the trainables
5 · RLHF: where the policy converges as the KL leash loosens
β E[RM] E[true] KL most likely answer
0.50 +1.69 +4.30 0.34 JSON + friendly line, right
0.10 +1.96 +3.45 1.81 short friendly JSON, WRONG category ← reward hacking
0.02 +2.11 +2.15 3.90 short friendly JSON, WRONG category

The reward model’s score (E[RM]) keeps rising as the leash loosens, while true value peaks and then falls. That is Goodhart’s law in one table. With --expert-labels, every row improves on SFT.

3 questions. Say your answer out loud, then tap to check it.

The scorer (the reward model) agrees with the people who labelled its training pairs 71% of the time. Yet the model trained against it still learns to game it (reward hacking). Why?Show answerHide

In plain words

Agreeing with the labellers means sharing their blind spot. They could not check the ticket’s category, so the scorer never learned to care about it, and the trained model exploits exactly that.

Picture it

A trainee copies a senior colleague’s marking well, 71 times in 100. But the senior never checks the sums, so neither does the trainee. Students who notice hand in wrong sums in neat handwriting and still get top marks.

With real numberslesson 16’s toy_alignment.py and README

  • The labellers score what they can see: +1.5 for valid JSON, +0.8 for friendly, minus 1.5 x length, and 0 for the right category (the true value gives it +3).
  • So they prefer short friendly JSON with the wrong category, 1.5 - (1.5 x 0.10) + 0.8 = 2.15, to JSON plus a friendly line with the right one, 1.5 - (1.5 x 0.35) + 0.8 = 1.775.
  • The reward model matches their choices 71% of the time, blind spot included.
  • The labellers’ favourite is only 2% of the imitation model’s answers, so only 17 of the 302 pairs include it: 71% agreement says little about it.
  • Loose leash, β 0.02: the scorer’s average rating climbs from +1.69 to +2.11, while the average true value sinks from +4.30 to +2.15. The two use different scales, so compare each with its own start.

Words to know

Reward model
A model trained on people’s comparisons to give any answer one score. Example: the toy’s scorer.
Labeller
A person who picks the better of two answers to make training data. Example: the toy’s fast labellers, who cannot check the category.
Imitation model (SFT model)
The model after it copied the agents’ examples. Example: the toy’s preference pairs are drawn from it.
Reward hacking
Raising the score through the scorer’s blind spots instead of getting better. Example: wrong category, top rating.
Go deeper: the engineer version

The kit's question

The reward model agrees with its labellers 71% of the time, yet RLHF still hacks it. Why?

The kit's answer

It learned the labellers’ blind spot. Agreement on the pairs it saw says nothing about answers it rarely saw.

More detail: The agreement is measured on the training pairs, and those are sampled from the SFT model, so they cover the answers it already writes. The reward model is a Bradley-Terry fit on VIS, the features the labellers can see; with fast labellers VIS = F[:, 1:] drops the correctness column, so no amount of data teaches it correctness. RLHF then moves probability to the answer the reward model rates highest (RM +2.11), which here has the wrong category. Two answers differ only in category (concise JSON, right and wrong, both RM +1.44), so every pair of them counts as a miss: part of the missing 29%. The labellers are noisy too (Bradley-Terry on LABEL_U), so even a perfect copy of their taste would not reach 100%.

DPO (direct preference optimization: training straight on the chosen and rejected pairs, with no scorer) changed the model less than RLHF with a loose leash did. Is that a good thing?Show answerHide

In plain words

Both. DPO only changes how often the model gives answers that appear in the pairs. So it chases a rarely seen loophole less far than loose RLHF does, and ignores a loophole the pairs never show. For the same reason, it cannot discover a better answer the pairs never show.

Picture it

A chef who only tweaks dishes that diners compared side by side never serves an untested gimmick. He never discovers a new favourite either. A chef who experiments freely finds both.

With real numberslesson 16’s toy_alignment.py (your default run) and README

  • The pairs come from the imitation-trained model: 400 draws of two answers, minus the 98 where both were the same, leaves 302.
  • That model gives the labellers’ favourite (short friendly JSON, wrong category) only 2% of the time, so only 17 of the 302 pairs include it.
  • Loose RLHF (β 0.10 in the lesson’s table) makes that answer the most likely; the average true value falls from +4.30 at β 0.50 to +3.45.
  • DPO at the default β 0.5 ends 0.51 away from the imitation-trained model (KL from SFT), against 1.81 for RLHF at β 0.10 and 3.90 at β 0.02.
  • It still copies the labellers’ bias where the pairs show it: the favourite rises from 0.02 to 0.18, so DPO’s average true value, +3.94, barely beats the imitation step’s +3.89.

Words to know

DPO (direct preference optimization)
Training straight on chosen and rejected pairs with one loss, no scorer and no sampling loop. Example: step 6 of the toy.
Preference pair
Two answers to the same prompt, one chosen and one rejected. Example: 302 in the toy (400 draws, identical ones dropped).
Loose leash
A small β, so the model may drift far from where it started. Example: β 0.10 or 0.02.
KL (drift)
How far a model’s choices have moved from another model’s; 0 means identical. Example: KL from SFT 0.51 for DPO, its drift from the imitation-trained (SFT) model.
Go deeper: the engineer version

The kit's question

DPO moved less than loose RLHF. Is that good?

The kit's answer

Both. It can’t exploit answers missing from the pairs, but it can’t discover better ones either.

More detail: DPO’s gradient only touches the logits of answers that appear as chosen or rejected (the two np.add.at lines), in proportion to how often. It inherits the labellers’ bias from the same pairs but follows it less far than the loose sweep rows. RLHF samples from its current policy and can push probability onto any answer the reward model rates highly, rare ones included. So moving less protects you from a flawed reward model and limits you when the reward model is good. Two cautions from the run: the comparison is with loose RLHF, not with RLHF at the same β (at β 0.5 the sampled run 5b drifts KL 0.34 and keeps +4.30, against DPO’s 0.51 and +3.94); and with 302 pairs every one of the six answers appears, so “never in the pairs” is the real-world case the toy only hints at.

What stops reward hacking (a model gaming its scorer) for good?Show answerHide

In plain words

Better judgements to learn from: experts who can check correctness, or a program that checks it. Then a well-set leash, and tests that measure real quality instead of the scorer’s rating.

Picture it

If students game a lazy grader, a stricter rule helps a little. What fixes it is a grader who checks the facts, or an answer key that marks the sums automatically.

With real numberslesson 16’s toy_alignment.py and README, and the RLHF and Training Techniques videos

  • Fast labellers: 0 points for the right category. Expert labellers (--expert-labels): +3, the same as the true value.
  • With experts, every leash setting improves on the imitation step: average true value +4.08 to +4.78, against +3.89.
  • A program grader, as in Fireworks RFT, scores each answer from 0 to 1, for example tests passing or an exact match (the RLHF and Training Techniques videos).
  • The leash alone limits the damage but cannot fix bad labels (the script’s own summary).

Words to know

RFT (reinforcement fine-tuning)
Fireworks’ reinforcement training that scores answers with a program instead of a reward model. Example: a grader checking the category.
Grader
Anything that scores an answer: a program, a judge model or a person. Example: a program that gives 1 if the JSON is valid and the category matches, else 0.
Expert labels
Choices made by people who can check correctness. Example: --expert-labels in the toy.
Eval
A repeatable test of a model on a fixed set of prompts. Example: pref_eval.py in step 2.
Go deeper: the engineer version

The kit's question

What fixes reward hacking for good?

The kit's answer

Better labels, meaning experts or a program that checks correctness, which is what RFT graders are, plus a tuned KL leash and evals on true quality.

More detail: In the toy, --expert-labels sets LABEL_U = TRUE_U and lets the reward model see the correctness feature (VIS = F), so the proxy and the goal line up. In production: fix the labels (expert labellers, or a programmatic grader as in Fireworks RFT), tune β on a held-out eval that measures true quality, and keep that eval running after training, because a rising reward alone proves nothing.

Question We trained a reward model on our labellers’ choices. Why not optimise against it as hard as we can?

One clear answer

A reward model is a proxy for what the customer wants. Optimise it too hard and it stops tracking the goal, so I keep a KL leash, and where correctness can be checked by a program I’d rather use a grader, which is what reinforcement fine-tuning on Fireworks does.

What this means

  • “A reward model is a proxy for what the customer wants”: The scorer is a stand-in, trained on labellers’ choices. It rates what the labellers could judge, not everything the customer cares about.
  • “Optimise it too hard and it stops tracking the goal”: Push the model to maximise that score and it finds loopholes. In the lesson the average rating rose from +1.69 to +2.11 while the average true value fell from +4.30 to +2.15.
  • “so I keep a KL leash”: I charge the model for drifting from its starting point, so it can improve but not run off after loopholes. β sets how long the leash is.
  • “where correctness can be checked by a program”: For a task such as “valid JSON with the right category”, code can mark each answer right or wrong.
  • “I’d rather use a grader, which is what reinforcement fine-tuning on Fireworks does”: Then I score answers with that program instead of a learned scorer. Fireworks sells this as RFT: a grader scores each answer from 0 to 1.

Your numbersSaved on this device and collected in the Day 12 wrap-up.

Hint: Section 4 of python toy_alignment.py, at the end of the 4 · REWARD MODEL line: “agrees with its labellers on ...% of pairs”.

Hint: Section 5: the row with the highest E[true]. The line under the table, “TRUE value peaks near β = ...”, gives it too.

Hint: The E[true] column of section 5, those two rows. Write both, for example “+4.30 / +2.15”.

Section 5 of your run shows the scorer’s average rating (E[RM]) rising as β shrinks, while the average true value (E[true]) peaks and then falls. In the --expert-labels run, every row beats the imitation step.

ModuleNotFoundError: No module named 'numpy'
The Python environment is off. From the 03-labs folder run source .venv/bin/activate, then run the script again. If it still fails, run pip install numpy.
--beta 0, as on the RLHF video’s command card, prints the same as the default.
The script treats 0 as its default, 0.5, because one of its lines divides by β. For a loose leash use a small number instead, such as python toy_alignment.py --beta 0.02.
Your section 5 table has 8 rows, not 3.
That is expected: the script tries 8 leash settings, from 4 down to 0.01. The lesson shows 3 of them (0.50, 0.10 and 0.02).
At --beta 0.1, 5b still favours the right-category answer, but the sweep says β 0.10 ends on the wrong one.
Expected: the sampled run stops short of the optimum. Compare the two KL numbers on the 5b line: 0.94 reached, 1.81 for the optimum. The script’s closing note that 5b lands on the optimum holds at the default β 0.5 (0.34 and 0.34), not here.
toy_alignment.py Python · 183 lines
"""
Lesson 16 · part 2 — the whole alignment story on one prompt, in numpy (seconds).
base model → SFT → reward model → RLHF (with / without the KL leash) → DPO
The "model" is a probability distribution over six possible answers to one support
ticket ("Card declined twice, launch is tomorrow"). That is a toy, but every step
below is the real algorithm, just on 6 outputs instead of all possible text:
SFT cross-entropy on demonstrations (videos 15)
RM Bradley-Terry fit on preference pairs (video 17)
RLHF policy gradient on RM reward − β·KL (video 17, PPO's objective)
DPO the DPO loss on the same pairs (video 18)
The twist, which is realistic: the people labelling preferences are fast, not
experts. They judge what they can see (valid JSON? short? friendly?) but cannot
check whether the ticket category is right. The reward model learns THEIR taste.
Without a KL leash, RLHF pushes toward the answer that scores best on that proxy:
short, friendly JSON with the WRONG category. That is reward hacking.
--expert-labels switches to labellers who can check correctness, which is the real fix.
python toy_alignment.py
python toy_alignment.py --beta 0.1 # looser leash for the sampled RLHF run and DPO
python toy_alignment.py --expert-labels # labellers who can verify the category
"""
import argparse
import numpy as np
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--beta", type=float, default=0.5, help="KL leash for the sampled RLHF run and for DPO")
ap.add_argument("--pairs", type=int, default=400, help="preference pairs collected")
ap.add_argument("--seed", type=int, default=1)
ap.add_argument("--expert-labels", action="store_true", help="labellers can check correctness")
a = ap.parse_args()
rng = np.random.default_rng(a.seed)
# ── the six possible answers ─────────────────────────────────────────────────
# name correct json length friendly
ANSWERS = [("concise JSON, right category", 1, 1, 0.15, 0),
("JSON + friendly line, right", 1, 1, 0.35, 1),
("concise JSON, WRONG category", 0, 1, 0.15, 0),
("friendly explanation, right, no JSON", 1, 0, 0.70, 1),
("short friendly JSON, WRONG category", 0, 1, 0.10, 1),
("'billing?' terse, no JSON", 1, 0, 0.05, 0)]
names = [x[0] for x in ANSWERS]
F = np.array([x[1:] for x in ANSWERS], float) # features: correct, json, length, friendly
K = len(ANSWERS)
# What people actually value (hidden from the reward model): right answer, parseable, not bloated.
true_w = np.array([3.0, 1.5, -1.5, 0.8])
TRUE_U = F @ true_w
# What fast, non-expert labellers reward: everything visible, but NOT correctness.
LABEL_U = TRUE_U if a.expert_labels else F @ np.array([0.0, 1.5, -1.5, 0.8])
def softmax(z):
z = z - z.max()
e = np.exp(z)
return e / e.sum()
def kl(p, q):
return float(np.sum(p * (np.log(p + 1e-12) - np.log(q + 1e-12))))
def show(title, p, extra=""):
top = int(np.argmax(p))
print(f"\n{title} E[true value] = {p @ TRUE_U:+.2f}{extra}")
for i in range(K):
bar = "█" * int(round(p[i] * 40))
print(f" {names[i]:38s} {p[i]:5.2f} {bar}{' ◀' if i == top else ''}")
# ── 1. base model: knows many ways to answer, no idea which one we want ──────
base_logits = np.log(np.array([0.10, 0.12, 0.18, 0.20, 0.25, 0.15]))
p_base = softmax(base_logits)
show("1 · BASE MODEL", p_base)
# ── 2. SFT: imitate demonstrations written by good support agents ───────────
demos = rng.choice(K, size=300, p=[0.38, 0.30, 0.05, 0.16, 0.03, 0.08]) # agents slip up sometimes # what agents actually wrote
logits = base_logits.copy()
for _ in range(400): # gradient descent on cross-entropy
p = softmax(logits)
grad = p - np.bincount(demos, minlength=K) / len(demos) # d(CE)/d(logits) for a softmax
logits -= 0.5 * grad
p_sft = softmax(logits)
show("2 · AFTER SFT (imitation)", p_sft, f" KL from base {kl(p_sft, p_base):.2f}")
# ── 3. preference pairs: sample two answers from the SFT model, a person picks one ──
i = rng.choice(K, size=a.pairs, p=p_sft)
j = rng.choice(K, size=a.pairs, p=p_sft)
keep = i != j
i, j = i[keep], j[keep]
prefer_i = rng.random(len(i)) < 1 / (1 + np.exp(-(LABEL_U[i] - LABEL_U[j]))) # labellers ~ Bradley-Terry on what THEY value
chosen = np.where(prefer_i, i, j)
rejected = np.where(prefer_i, j, i)
who = "expert labellers (can check the category)" if a.expert_labels else "fast labellers (cannot check the category)"
print(f"\n3 · PREFERENCE DATA {len(chosen)} pairs sampled from the SFT model, labelled by {who}")
# ── 4. reward model: Bradley-Terry on what it CAN see (json, length, confident) ──
VIS = F if a.expert_labels else F[:, 1:] # a fast labeller's RM can't see correctness
w_rm = np.zeros(VIS.shape[1])
for _ in range(3000):
margin = (VIS[chosen] - VIS[rejected]) @ w_rm
g = -((1 - 1 / (1 + np.exp(-margin)))[:, None] * (VIS[chosen] - VIS[rejected])).mean(0)
w_rm -= 0.5 * (g + 1e-3 * w_rm)
RM = VIS @ w_rm
acc = np.mean(RM[chosen] > RM[rejected])
labels = ["correct", "json", "length", "friendly"][-VIS.shape[1]:]
print("\n4 · REWARD MODEL learned weights " + " ".join(f"{n} {w:+.2f}" for n, w in zip(labels, w_rm))
+ f" (agrees with its labellers on {acc:.0%} of pairs)")
print(f" Its favourite answer: '{names[int(np.argmax(RM))]}' (true value {TRUE_U[int(np.argmax(RM))]:+.2f})")
for k in range(K):
print(f" {names[k]:38s} RM {RM[k]:+.2f} true {TRUE_U[k]:+.2f}")
# ── 5. RLHF: policy gradient on E[RM] − β·KL(π ‖ π_sft) (PPO optimises this objective) ──
def rlhf(beta, steps=4000, lr=0.2, batch=64):
th = np.log(p_sft).copy()
for _ in range(steps):
p = softmax(th)
y = rng.choice(K, size=batch, p=p) # sample answers from the current policy
adv = RM[y] - beta * (np.log(p[y]) - np.log(p_sft[y])) # reward minus the KL "leash" charge
adv = adv - adv.mean() # baseline, reduces variance
g = np.zeros(K)
for yy, aa in zip(y, adv): # REINFORCE: raise log-prob of above-average answers
g += aa * (np.eye(K)[yy] - p)
th += lr * g / batch
return softmax(th)
# Where does RLHF end up? The KL-regularised objective has an exact optimum:
# π*(y) ∝ π_sft(y) · exp(reward(y) / β)
# PPO is a sampling-based way of approaching it. Sweep the leash β from tight to loose:
print("\n5 · RLHF: where the policy converges as the KL leash loosens (π* ∝ π_sft · exp(RM/β))")
print(f" {'β':>6s} {'E[RM]':>7s} {'E[true]':>8s} {'KL':>5s} most likely answer")
sweep = []
for beta in (4, 1, 0.5, 0.2, 0.1, 0.05, 0.02, 0.01):
pi = softmax(np.log(p_sft) + RM / beta)
sweep.append((beta, pi))
print(f" {beta:6.2f} {pi @ RM:+7.2f} {pi @ TRUE_U:+8.2f} {kl(pi, p_sft):5.2f} {names[int(np.argmax(pi))]}")
best = max(sweep, key=lambda t: t[1] @ TRUE_U)
worst = sweep[-1]
if worst[1] @ TRUE_U < best[1] @ TRUE_U - 0.2:
print(f" ↑ the proxy (E[RM]) keeps rising as β shrinks, but TRUE value peaks near β = {best[0]} and then falls:"
"\n that is reward hacking (Goodhart's law: optimise a proxy hard enough and it stops tracking the goal).")
b = a.beta if a.beta > 0 else 0.5
p_rlhf = rlhf(b)
closed = softmax(np.log(p_sft) + RM / b)
show(f"5b · RLHF by policy gradient (β = {b}), the PPO-style sampled version", p_rlhf,
f" E[RM] {p_rlhf @ RM:+.2f} KL from SFT {kl(p_rlhf, p_sft):.2f} (exact optimum KL {kl(closed, p_sft):.2f})")
# ── 6. DPO: same pairs, no reward model, no sampling ─────────────────────────
th = np.log(p_sft).copy()
ref = np.log(p_sft)
for _ in range(3000):
lp = th - np.log(np.sum(np.exp(th))) # log π(y)
m = b * ((lp[chosen] - ref[chosen]) - (lp[rejected] - ref[rejected]))
s = 1 - 1 / (1 + np.exp(-m)) # = σ(−m): how wrong each pair still is
p = softmax(th)
g = np.zeros(K)
np.add.at(g, chosen, -b * s) # push chosen up …
np.add.at(g, rejected, b * s) # … rejected down
# (the softmax normaliser cancels inside each pair's margin, so these are the exact gradients)
th -= 0.5 * g / len(chosen)
p_dpo = softmax(th)
show(f"6 · DPO (β = {b}, same pairs, no reward model)", p_dpo, f" KL from SFT {kl(p_dpo, p_sft):.2f}")
print(f"""
What to notice
• SFT moved probability onto the kinds of answers agents write. Imitation got most of the way.
• The reward model learned its labellers' taste, blind spot included. It is a proxy for what you want.
• Loosening the leash (smaller β) always raises the proxy E[RM]. True value rises, peaks, then falls.
That is reward hacking, and it is why production RLHF always keeps a KL term and tunes β.
• The sampled policy-gradient run (5b) lands on the same answer as the exact optimum: that is PPO's job.
• The leash limits the damage by keeping the policy close to SFT. It cannot fix bad labels.
• DPO learns from the same pairs, so it inherits the same bias. It stays closer to SFT because it only
moves answers that appear in the pairs.
• Re-run with --expert-labels: every method now improves on SFT. Better preferences beat a better
algorithm, and a program that checks correctness is what RFT graders are for.
""")