About 10 minutes, free, and it works on your phone. Lock in today before you move on.
Done when
Your table shows, before and after DPO: valid JSON (up), the right category (not down) and the margin between chosen and rejected answers (grown). And you can explain, with the toy’s numbers, why a scorer that cannot check correctness gets gamed.
Fields you filled in on the step pages are already here. Nothing leaves this browser.
Step 1 · Training toys, part 2
High agreement, yet the scorer still gets gamed: it copied the labellers’ blind spot along with their taste.
Example: 71, as in the lesson (the script uses a fixed random seed, so your run should match)
%
Looser than this, the model starts gaming the scorer: the rating still rises but the real value falls. In a real project this is the setting you tune on an eval.
Example: 0.5: +4.30, the highest of all 8 rows; the script’s line under the table says the same.
β
How much real value the model’s answers carry on average, on a firm leash and on a very loose one. The drop is reward hacking in one pair of numbers.
Example: +4.30 / +2.15 (the lesson’s rows)
true value
Step 2 · Preference tuning with DPO
The share of the 100 test tickets answered with JSON that fits the ticket format. DPO should push it up.
Example: about 80 / 98, as the lesson expects
%
The correctness check: it must not drop. If it does, taste has cost correctness.
Example: about 70, then the same or higher, as the lesson expects
%
On average over the 100 checking pairs, how much more likely the model finds the chosen answer than the rejected one, on a log scale: +1 is about 2.7 times as likely, +10 about 22,000 times (e^1 = 2.72, e^10 = 22,026). It should grow a lot after DPO.
The scorer (the reward model) agrees with the people who labelled its training pairs 71% of the time. Yet the model trained against it still learns to game it (reward hacking). Why?
Show answerHide answer
In plain words
Agreeing with the labellers means sharing their blind spot. They could not check the ticket’s category, so the scorer never learned to care about it, and the trained model exploits exactly that.
Picture it
A trainee copies a senior colleague’s marking well, 71 times in 100. But the senior never checks the sums, so neither does the trainee. Students who notice hand in wrong sums in neat handwriting and still get top marks.
With real numberslesson 16’s toy_alignment.py and README
The labellers score what they can see: +1.5 for valid JSON, +0.8 for friendly, minus 1.5 x length, and 0 for the right category (the true value gives it +3).
So they prefer short friendly JSON with the wrong category, 1.5 - (1.5 x 0.10) + 0.8 = 2.15, to JSON plus a friendly line with the right one, 1.5 - (1.5 x 0.35) + 0.8 = 1.775.
The reward model matches their choices 71% of the time, blind spot included.
The labellers’ favourite is only 2% of the imitation model’s answers, so only 17 of the 302 pairs include it: 71% agreement says little about it.
Loose leash, β 0.02: the scorer’s average rating climbs from +1.69 to +2.11, while the average true value sinks from +4.30 to +2.15. The two use different scales, so compare each with its own start.
Words to know
Reward model
A model trained on people’s comparisons to give any answer one score. Example: the toy’s scorer.
Labeller
A person who picks the better of two answers to make training data. Example: the toy’s fast labellers, who cannot check the category.
Imitation model (SFT model)
The model after it copied the agents’ examples. Example: the toy’s preference pairs are drawn from it.
Reward hacking
Raising the score through the scorer’s blind spots instead of getting better. Example: wrong category, top rating.
Go deeper: the engineer version
The kit's question
The reward model agrees with its labellers 71% of the time, yet RLHF still hacks it. Why?
The kit's answer
It learned the labellers’ blind spot. Agreement on the pairs it saw says nothing about answers it rarely saw.
More detail: The agreement is measured on the training pairs, and those are sampled from the SFT model, so they cover the answers it already writes. The reward model is a Bradley-Terry fit on VIS, the features the labellers can see; with fast labellers VIS = F[:, 1:] drops the correctness column, so no amount of data teaches it correctness. RLHF then moves probability to the answer the reward model rates highest (RM +2.11), which here has the wrong category. Two answers differ only in category (concise JSON, right and wrong, both RM +1.44), so every pair of them counts as a miss: part of the missing 29%. The labellers are noisy too (Bradley-Terry on LABEL_U), so even a perfect copy of their taste would not reach 100%.
DPO (direct preference optimization: training straight on the chosen and rejected pairs, with no scorer) changed the model less than RLHF with a loose leash did. Is that a good thing?
Show answerHide answer
In plain words
Both. DPO only changes how often the model gives answers that appear in the pairs. So it chases a rarely seen loophole less far than loose RLHF does, and ignores a loophole the pairs never show. For the same reason, it cannot discover a better answer the pairs never show.
Picture it
A chef who only tweaks dishes that diners compared side by side never serves an untested gimmick. He never discovers a new favourite either. A chef who experiments freely finds both.
With real numberslesson 16’s toy_alignment.py (your default run) and README
The pairs come from the imitation-trained model: 400 draws of two answers, minus the 98 where both were the same, leaves 302.
That model gives the labellers’ favourite (short friendly JSON, wrong category) only 2% of the time, so only 17 of the 302 pairs include it.
Loose RLHF (β 0.10 in the lesson’s table) makes that answer the most likely; the average true value falls from +4.30 at β 0.50 to +3.45.
DPO at the default β 0.5 ends 0.51 away from the imitation-trained model (KL from SFT), against 1.81 for RLHF at β 0.10 and 3.90 at β 0.02.
It still copies the labellers’ bias where the pairs show it: the favourite rises from 0.02 to 0.18, so DPO’s average true value, +3.94, barely beats the imitation step’s +3.89.
Words to know
DPO (direct preference optimization)
Training straight on chosen and rejected pairs with one loss, no scorer and no sampling loop. Example: step 6 of the toy.
Preference pair
Two answers to the same prompt, one chosen and one rejected. Example: 302 in the toy (400 draws, identical ones dropped).
Loose leash
A small β, so the model may drift far from where it started. Example: β 0.10 or 0.02.
KL (drift)
How far a model’s choices have moved from another model’s; 0 means identical. Example: KL from SFT 0.51 for DPO, its drift from the imitation-trained (SFT) model.
Go deeper: the engineer version
The kit's question
DPO moved less than loose RLHF. Is that good?
The kit's answer
Both. It can’t exploit answers missing from the pairs, but it can’t discover better ones either.
More detail: DPO’s gradient only touches the logits of answers that appear as chosen or rejected (the two np.add.at lines), in proportion to how often. It inherits the labellers’ bias from the same pairs but follows it less far than the loose sweep rows. RLHF samples from its current policy and can push probability onto any answer the reward model rates highly, rare ones included. So moving less protects you from a flawed reward model and limits you when the reward model is good. Two cautions from the run: the comparison is with loose RLHF, not with RLHF at the same β (at β 0.5 the sampled run 5b drifts KL 0.34 and keeps +4.30, against DPO’s 0.51 and +3.94); and with 302 pairs every one of the six answers appears, so “never in the pairs” is the real-world case the toy only hints at.
What stops reward hacking (a model gaming its scorer) for good?
Show answerHide answer
In plain words
Better judgements to learn from: experts who can check correctness, or a program that checks it. Then a well-set leash, and tests that measure real quality instead of the scorer’s rating.
Picture it
If students game a lazy grader, a stricter rule helps a little. What fixes it is a grader who checks the facts, or an answer key that marks the sums automatically.
With real numberslesson 16’s toy_alignment.py and README, and the RLHF and Training Techniques videos
Fast labellers: 0 points for the right category. Expert labellers (--expert-labels): +3, the same as the true value.
With experts, every leash setting improves on the imitation step: average true value +4.08 to +4.78, against +3.89.
A program grader, as in Fireworks RFT, scores each answer from 0 to 1, for example tests passing or an exact match (the RLHF and Training Techniques videos).
The leash alone limits the damage but cannot fix bad labels (the script’s own summary).
Words to know
RFT (reinforcement fine-tuning)
Fireworks’ reinforcement training that scores answers with a program instead of a reward model. Example: a grader checking the category.
Grader
Anything that scores an answer: a program, a judge model or a person. Example: a program that gives 1 if the JSON is valid and the category matches, else 0.
Expert labels
Choices made by people who can check correctness. Example: --expert-labels in the toy.
Eval
A repeatable test of a model on a fixed set of prompts. Example: pref_eval.py in step 2.
Go deeper: the engineer version
The kit's question
What fixes reward hacking for good?
The kit's answer
Better labels, meaning experts or a program that checks correctness, which is what RFT graders are, plus a tuned KL leash and evals on true quality.
More detail: In the toy, --expert-labels sets LABEL_U = TRUE_U and lets the reward model see the correctness feature (VIS = F), so the proxy and the goal line up. In production: fix the labels (expert labellers, or a programmatic grader as in Fireworks RFT), tune β on a held-out eval that measures true quality, and keep that eval running after training, because a rising reward alone proves nothing.
Why merge Day 11’s small add-on (the SFT adapter, trained on examples) into the model first, instead of running DPO on the plain instruct model?
Show answerHide answer
In plain words
DPO polishes a model that already does the task. The plain model gets only about 40% of ticket categories right, so the approved answers are still unlikely for it, and DPO mostly learns to avoid the rejected text.
Picture it
Coaching someone’s handwriting helps once they can form the letters. Show a beginner pairs of neat and messy pages and they mostly learn what to avoid, not how to write.
With real numberslesson 17’s README (Day 11), dpo_mlx.sh and the RLHF and DPO videos
Day 11’s expected table: before its training, the model gets about 40% of ticket categories right; after training, all three methods score well above that.
dpo_mlx.sh merges Day 11’s LoRA into sft_model, then runs DPO from there and uses it as the frozen reference too.
Without it, DPO starts from the plain instruct model, the one at about 40%.
The RLHF video’s rule: “You cannot reinforce behaviour the model never produces.” The DPO video’s: SFT first, then DPO.
Words to know
SFT (supervised fine-tuning)
Training on examples of the answers you want. Example: Day 11’s LoRA on 800 support tickets.
Adapter (LoRA)
A small add-on of extra learned numbers trained on top of a frozen model. Example: Day 11’s, in adapters/lora.
Fuse
Bake an adapter’s numbers into the model so it becomes one self-contained model. Example: sft_model.
Instruct model
A model already trained to follow instructions and chat. Example: Qwen2.5-0.5B-Instruct.
Go deeper: the engineer version
The kit's question
Why fuse the SFT adapter first rather than running DPO on the raw instruct model?
The kit's answer
DPO refines a model that already does the task. With nothing to refine, it mostly learns “not the rejected text”.
More detail: DPO only reshapes probability between responses the model can already produce. Starting from a model that already emits valid triage JSON, the pairs refine format and correctness at the margin. From the raw instruct model the chosen answers are unlikely, so much of the loss is reduced by pushing the rejected text down. The same model is the frozen reference (--reference-model-path "$START"), so the reference also knows the task.
What does β (beta, the leash setting, 0.1 in dpo_mlx.sh) actually do?
Show answerHide answer
In plain words
It sets how far the model may drift from its frozen starting copy before training stops pushing it further. A smaller β lets it move further.
Picture it
Think of β as how short the leash is. On a short leash (large β), the dog gets a few steps away and the leash goes taut. On a long leash (small β), it wanders much further before it is pulled back.
With real numbersthe DPO video’s worked step, recomputed at other β values
Margin: the chosen answer +0.8, the rejected -0.4, both against the frozen copy: 0.8 - (-0.4) = 1.2.
β 0.1: 0.1 x 1.2 = 0.12. The sigmoid σ turns any number into a value between 0 and 1: σ(0.12) = 0.53. Loss = -log(0.53) = 0.63, down from 0.69 when the margin was 0.
β 0.05: the same margin earns only 0.06, loss 0.66, so training keeps pulling the two apart for longer.
β 0.2: 0.24, loss 0.58, so less drift is needed before the pull fades.
Words to know
β (beta)
The leash: how far the model may move from the reference. Example: 0.1 in dpo_mlx.sh.
Margin
How much more the model favours the chosen answer than the rejected one. Example: 1.2 in the DPO video.
Log-probability
The natural log of how likely the model is to write an answer; DPO works in these. Example: +0.8 against the reference.
Loss
The single number training tries to shrink. Example: 0.69 at the start, 0.63 after one step.
Go deeper: the engineer version
The kit's question
What does β do here, concretely?
The kit's answer
It scales the margin in the loss. A smaller β lets the policy move further from the reference before the loss stops rewarding it.
More detail: L = −log σ(β · [(log π(y_c) − log π_ref(y_c)) − (log π(y_r) − log π_ref(y_r))]). Its gradient is scaled by σ(−β · margin), which shrinks as β · margin grows, so a larger β saturates at a smaller margin (a tighter leash). β is the same KL coefficient as in RLHF: DPO is derived from that KL-regularised objective, with the leash built into the loss. Set it with BETA=0.2 bash dpo_mlx.sh.
Suppose half your rejected answers are chatty (the right JSON wrapped in chat), and none have the wrong category (wrong_cat). What happens to category accuracy after DPO?
Show answerHide answer
In plain words
The model learns to fix its format, not to get categories right. Category accuracy may stay flat or even slip, because nothing in the pairs rewards the right category.
Picture it
If a supervisor only ever corrects greetings and sign-offs, the new agent writes polite emails that may still send customers to the wrong team.
With real numberslesson 18’s make_pairs.py and README
The lesson’s 800 pairs: each rejected answer is one of four mistakes, about 200 each (800 ÷ 4).
Only the wrong-category kind differs from the chosen answer in its category. Chatty, prose and wrong-severity answers all keep the right one.
The question’s mix: 400 chatty (half of 800) and 0 wrong-category, so no pair teaches the category.
Watch category_%: the lesson expects it to hold at about 70% or rise. A drop means taste has cost correctness.
Words to know
Chatty answer
The right JSON wrapped in friendly chat, so a program cannot read it. Example: “Sure! Here’s the JSON you asked for: ...” (chatty_% also counts prose answers.)
wrong_cat pair
A pair whose rejected answer is valid JSON with the wrong category. Example: 213 of the 800.
Category accuracy
The share of test tickets given the right category. Example: category_%, about 70% before.
Preference pair
Two answers to the same prompt, one chosen and one rejected. Example: 800 in this lesson.
Go deeper: the engineer version
The kit's question
Your pairs are 50% “chatty”. What happens to accuracy if none of them are wrong_cat?
The kit's answer
The model learns format, not correctness, and category accuracy may even slip.
More detail: DPO raises the log-probability gap along whatever separates chosen from rejected. If no pair separates on category, the gradient carries no category signal, and drift from the reference (bounded by β) can move it either way. The README’s fixes: a larger β, fewer iterations, or more wrong_cat pairs, so correctness is part of what “preferred” means. make_pairs.py’s own rule: pairs should differ in the thing you care about.
When would you use ORPO (odds-ratio preference optimization, a relative of DPO) instead of DPO? Fireworks runs it with --loss-method ORPO.
Show answerHide answer
In plain words
When you want to teach the task and the preference in one training run, without keeping a second, frozen copy of the model in memory.
Picture it
DPO is coaching after a course: first the course, then a coach who compares you with your old self. ORPO is one class that teaches the skill and the taste together, with no old self to compare against.
With real numbersthe DPO and RLHF videos, dpo_mlx.sh and dpo_fireworks.sh
RLHF with PPO: 4 models in memory. DPO: 2, the model being trained and a frozen reference.
ORPO: no reference copy, so only the model being trained, and no separate example-training step first.
On your Mac, each copy of the 0.5-billion-number model is 0.5 billion x 2 bytes = about 1 GB.
For the Qwen2.5 7B model dpo_fireworks.sh uses (7.6 billion numbers), a frozen copy at 2 bytes each is 7.6 x 2 = about 15 GB.
Words to know
ORPO (odds-ratio preference optimization)
A preference method that does example training and preference training in one run, with no reference model. Example: --loss-method ORPO.
Reference model
A frozen copy of the starting model that training measures drift against. Example: one of DPO’s 2 models.
SFT (supervised fine-tuning)
Training on examples of the answers you want. Example: Day 11’s LoRA.
bf16
A 16-bit number format, 2 bytes per number, used for training. Example: the 0.5B model is about 1 GB in bf16.
Go deeper: the engineer version
The kit's question
When would you use ORPO instead (--loss-method ORPO on Fireworks)?
The kit's answer
When you want SFT and preference tuning in one step without holding a reference model in memory.
More detail: ORPO adds an odds-ratio term, −λ · log σ(log(odds(chosen) / odds(rejected))) with odds(y) = P(y) ÷ (1 − P(y)), to the ordinary SFT loss on the chosen answer, so it needs neither a reference model nor a separate SFT stage. With no reference there is no explicit leash to a known-good model, so evaluate accuracy as carefully as with DPO. On Fireworks: firectl dpo-job create --loss-method ORPO with --orpo-lambda (the comment at the top of dpo_fireworks.sh).
Under 20 seconds each. Record yourself once and listen back.
Step 1 · Training toys, part 2
QuestionWe trained a reward model on our labellers’ choices. Why not optimise against it as hard as we can?
One clear answer
A reward model is a proxy for what the customer wants. Optimise it too hard and it stops tracking the goal, so I keep a KL leash, and where correctness can be checked by a program I’d rather use a grader, which is what reinforcement fine-tuning on Fireworks does.
What this means
“A reward model is a proxy for what the customer wants”: The scorer is a stand-in, trained on labellers’ choices. It rates what the labellers could judge, not everything the customer cares about.
“Optimise it too hard and it stops tracking the goal”: Push the model to maximise that score and it finds loopholes. In the lesson the average rating rose from +1.69 to +2.11 while the average true value fell from +4.30 to +2.15.
“so I keep a KL leash”: I charge the model for drifting from its starting point, so it can improve but not run off after loopholes. β sets how long the leash is.
“where correctness can be checked by a program”: For a task such as “valid JSON with the right category”, code can mark each answer right or wrong.
“I’d rather use a grader, which is what reinforcement fine-tuning on Fireworks does”: Then I score answers with that program instead of a learned scorer. Fireworks sells this as RFT: a grader scores each answer from 0 to 1.
Step 2 · Preference tuning with DPO
QuestionOur fine-tuned support model still wraps its JSON in chat and breaks our parser. We have no budget for labellers. What would you do?
One clear answer
Their agents already edit the model’s drafts. Every edit is a preference pair. Run DPO on an SFT’d model and the formatting failures disappear. I’d watch category accuracy so taste doesn’t cost correctness, and where correctness can be checked by code, I’d use RFT with a grader instead.
What this means
“Their agents already edit the model’s drafts”: Support staff already fix the model’s replies as part of their day.
“Every edit is a preference pair”: Each fix gives two answers to the same ticket: the edited one (chosen) and the original draft (rejected). That is free training data.
“Run DPO on an SFT’d model”: Train on those pairs with DPO, starting from the model already fine-tuned on examples (SFT). Here that is Day 11’s LoRA, merged in.
“the formatting failures disappear”: As the lesson expects: valid JSON from about 80% to about 98%, chatty replies from about 20% to about 1%.
“I’d watch category accuracy so taste doesn’t cost correctness”: I compare the right-category rate before and after, because pairs that differ only in style can quietly make answers less accurate.
“where correctness can be checked by code, I’d use RFT with a grader instead”: If a program can mark an answer right or wrong, I train with RFT (reinforcement fine-tuning), which scores each answer with that program instead of preferences.
2 worked examples from today’s videos. Work each on paper, then reveal.
From the video:RLHF4:10
Whiteboard 1 of 2
A customer complains that an outage cost them a day. A reward model (a scorer trained on people’s choices) rates answer A, an apology with the cause, a fix and a credit, at 2.1. It rates answer B, accurate but with no apology, at -0.4. How likely is a person to prefer A?
Given, in plain words
The reward model turns each answer into one number. The chance a person prefers A depends only on the gap between the two numbers, passed through the sigmoid (σ): a curve that turns any number into a chance between 0% and 100%. A gap of 0 gives 50%.
Reveal the answerHide the answer
Answer · in plain words
About 92%: a bit more than 9 times in 10. A gap of 2.5 points on the scorer’s scale is a strong preference, but never a certainty.
Picture it
Like chess ratings: the bigger the gap between two players’ ratings, the more surely the stronger one wins, but an upset is always possible.
Careful
The reward model knows only what its labellers could judge. In step 1’s toy they could not check the ticket’s category, and the trained model exploited exactly that.
Worked answer, step by step
Gap between the two scores: 2.1 - (-0.4) = 2.1 + 0.4 = 2.5.
Sigmoid of the gap: σ(2.5) = 1 ÷ (1 + e^-2.5), where e is about 2.718 and e^-2.5 means 1 divided by e to the power 2.5.
e^2.5 is about 12.18, so e^-2.5 = 1 ÷ 12.18 = 0.082.
1 ÷ (1 + 0.082) = 1 ÷ 1.082 = 0.924, about 92%.
The other side: the chance of preferring B is 1 - 0.924 = 0.076, about 8%.
Training the reward model means adjusting its scores until chances like these match tens of thousands of real human choices (the video’s figure).
Go deeper: the engineer version
The kit's question
Worked example · the reward model’s job · prompt: the outage complaint
The kit's answer
answer A · apology, cause, fix, credit: score 2.1. answer B · accurate, no apology: score −0.4. P(person prefers A) = σ(2.1 + 0.4): ≈ 92%. training data: tens of thousands of comparisons. The reward model turns human judgement into a number an optimizer can chase.
More detail: This is the Bradley-Terry model: P(A preferred to B) = σ(r_A − r_B), and the reward model is fitted by minimising −log σ(r_chosen − r_rejected) over the comparisons. Only differences matter, so adding a constant to every score changes nothing. toy_alignment.py fits exactly this in its step 4, on the features the labellers could see. The DPO loss has the same −log σ shape, with the reward replaced by β times the log-probability ratio to the reference.
Words to know
Reward model
A model trained on people’s comparisons to give any answer one score. Example: 2.1 and -0.4 here.
Sigmoid (σ)
A curve that turns any number into a value between 0 and 1; 0 gives 0.5. Example: σ(2.5) = 0.92.
Bradley-Terry model
The rule that turns two scores into the chance a person prefers one: the sigmoid of the gap. Example: this whiteboard.
Comparison
A person’s pick of the better of two answers to the same prompt. Example: A over B.
From the video:Direct Preference Optimization3:50
Whiteboard 2 of 2
In one DPO training step, the model has made the chosen answer 0.8 more likely and the rejected answer 0.4 less likely, both in log terms against a frozen copy of the starting model. β is 0.1. What is the loss, and what was it when training started?
Given, in plain words
DPO compares log-probabilities (the natural log of how likely the model is to write an answer), each minus the frozen reference copy’s. +0.8 in log terms means about 2.2 times as likely as the frozen copy thinks (e^0.8 = 2.23); -0.4 means about two thirds as likely (e^-0.4 = 0.67). The margin is the chosen change minus the rejected change. β, the leash setting, scales the margin. The loss, the number training tries to shrink, is -log σ(β x margin).
Reveal the answerHide the answer
Answer · in plain words
About 0.63, down from 0.69 at the start. The further the chosen and rejected answers pull apart, the lower the loss.
Picture it
A see-saw: DPO pushes the chosen end up and the rejected end down. The loss says how level it still is: level at the start (0.69), a little tipped after this step (0.63).
Careful
This margin is measured against the frozen copy. pref_eval.py --pairwise in step 2 reports a different margin: the raw gap in log-probability between chosen and rejected, with no reference subtracted. Its growth from before to after is what tracks DPO’s progress.
At the start both answers match the reference: margin 0, σ(0) = 0.5, loss = -log(0.5) = 0.69.
For comparison, a margin of 5 at the same β gives σ(0.5) = 0.62 and a loss of 0.47: more separation, lower loss.
Go deeper: the engineer version
The kit's question
Worked example · one DPO step · β = 0.1 · log-probabilities relative to the reference
The kit's answer
chosen: log π − log π_ref: +0.8. rejected: log π − log π_ref: −0.4. β × margin: 0.1 × 1.2 = 0.12. loss = −log σ(0.12): 0.63 (0.69 at start). Read it: The loss falls as chosen and rejected pull apart.
More detail: L = −log σ(β[(log π(y_c) − log π_ref(y_c)) − (log π(y_r) − log π_ref(y_r))]). e^−0.12 = 0.8869, σ(0.12) = 0.5300, −ln 0.5300 = 0.6349. The gradient of the loss with respect to the margin is −β · σ(−β · margin); its size is 0.1 x 0.470 = 0.047 here, against 0.05 at the start: it shrinks as the pair separates, which is the leash at work. toy_alignment.py step 6 computes σ(−β · margin) as s (0.470 here) and applies b * s as the gradient (the two np.add.at lines).
Words to know
DPO (direct preference optimization)
Training straight on chosen and rejected pairs with one loss, no scorer and no sampling loop.
Log-probability
The natural log of how likely the model is to write an answer. Example: +0.8 against the reference.
Reference model
A frozen copy of the starting model that DPO measures drift against.
Loss
The single number training tries to shrink. Example: 0.69 falling to 0.63.
How you will use this
When a team wants better answers, first ask what feedback it has. Written examples can teach a desired format or task; pairs of better and worse answers can teach a preference; a program that checks each answer can provide another training signal. Today you tried the preference route and saw why any scorer must still be checked against what customers actually need.