Day 12 · Preference tuning
Part 2 · Day 12 of 12
0 of 12 days done
About 1 h 55 Free (Fireworks DPO optional, paid) Mac and Fireworks
Today: Watch a model learn to game its scorer in a small simulation (reward hacking). Then teach a real model on your Mac to prefer clean answers with DPO (direct preference optimization), and check it did not get less accurate.
Example
Imagine a support bot rewarded for sounding helpful, while nobody checks whether it sent the ticket to the right team. It learns to give polished answers that earn higher scores even as its real help gets worse. Today’s small simulation makes that failure visible. Then you will train a real model to prefer better answers and check that it still chooses the right category.
By the end you’ll have
Before-and-after numbers for three things: how often the model answers in valid JSON, how often it picks the right category, and how clearly it prefers the approved answers.
Expected: Roughly, as in the lesson: valid JSON from about 80% to about 98%, the right category about 70% before and the same or higher after, and the margin (how strongly the model favours the approved answer, on a log scale) from about +1 to about +10.
The 6 numbers you write down
How often the reward model agrees with its labellers (%)
High agreement, yet the scorer still gets gamed: it copied the labellers’ blind spot along with their taste.
Example: 71, as in the lesson (the script uses a fixed random seed, so your run should match)
Leash setting where the average true value peaks (β)
Looser than this, the model starts gaming the scorer: the rating still rises but the real value falls. In a real project this is the setting you tune on an eval.
Example: 0.5: +4.30, the highest of all 8 rows; the script’s line under the table says the same.
Average true value with a firm and a very loose leash (
E[true]at β 0.50 and 0.02)How much real value the model’s answers carry on average, on a firm leash and on a very loose one. The drop is reward hacking in one pair of numbers.
Example: +4.30 / +2.15 (the lesson’s rows)
Valid JSON before and after DPO (%)
The share of the 100 test tickets answered with JSON that fits the ticket format. DPO should push it up.
Example: about 80 / 98, as the lesson expects
Right category before and after DPO (%)
The correctness check: it must not drop. If it does, taste has cost correctness.
Example: about 70, then the same or higher, as the lesson expects
Preference margin before and after DPO
On average over the 100 checking pairs, how much more likely the model finds the chosen answer than the rejected one, on a log scale: +1 is about 2.7 times as likely, +10 about 22,000 times (e^1 = 2.72, e^10 = 22,026). It should grow a lot after DPO.
Example: about +1 / +10, as the lesson expects
How you will use this: When a team wants better answers, first ask what feedback it has. Written examples can teach a desired format or task; pairs of better and worse answers can teach a preference; a program that checks each answer can provide another training signal. Today you tried the preference route and saw why any scorer must still be checked against what customers actually need.
Before you start
Day 11’s LoRA adapter, if you trained it.
Step 2 first merges the small add-on you trained on Day 11 (a LoRA adapter) into the model: train on examples first, then DPO. Without it, DPO still runs, but it polishes a model that gets only about 40% of ticket categories right. Quick check, from the
03-labsfolder:ls lessons/17-full-vs-peft-mlx/adapters/loraA list of files means it is there. “No such file or directory” means Day 11’s training has not finished:
VARIANTS=lora bash lessons/17-full-vs-peft-mlx/run_variants.shtrains only the LoRA one, in 5 to 15 minutes.The Python environment from Day 1, and internet for the first run.
Run everything from the
03-labsfolder aftersource .venv/bin/activate. The first DPO run installsmlx-lm-lora, the tool that trains DPO with MLX. If Day 11 did not already download it, it also fetches the 0.5-billion-number Qwen model, about 1 GB.Only for the optional Fireworks part: your key and the spend cap from Day 1.
dpo_fireworks.shtrains the same job on Fireworks for cents, then offers to deploy it on a dedicated GPU, which bills by the hour. Skip it and the day stays free.
Warm-up from Day 11
Section titled “Warm-up from Day 11”About 2 minutes. Say your answer out loud, then tap to check it. From Day 11 · Training methods in miniature.
In Day 11’s toy (one small layer of 4,096 numbers), the rank-1 add-on never catches up with full fine-tuning, even with 200 examples. Why? (The rank is how many separate kinds of change the add-on can make at once.)Show answerHide
In plain words
The change this task needs has two separate kinds (two directions), and a rank-1 add-on can make only one. More examples cannot supply the missing one; only a bigger rank can.
Picture it
Treasure lies to the north-east, and you may only move along a straight track running east. However many tries you get, you stop at the closest point on the track, never at the treasure. Add a second track running north and you reach it exactly.
With real numberslesson 16’s toy_finetune.py, the lesson’s sample run
- The task needs a rank-2 change: two separate kinds of change (the script’s
--true-rankis 2 by default). - Rank 1: 64 + 64 = 128 trained numbers, 3.1% of the layer. Error 0.3937, 0.1951 and 0.1321 at 16, 48 and 200 examples (0 is perfect).
- Rank 2: 256 trained numbers, 6.2%. Error at 200 examples: 0.0000, as good as full fine-tuning (0.0002).
- So rank 1 trains half as many numbers as rank 2 but misses one whole kind of change: at 200 examples its error is still 0.1321, while rank 2’s is 0.0000.
Words to know
- Rank
- How many separate kinds of change an add-on can make at once. Example: the toy’s task needs 2.
- Held-out error
- How wrong the trained layer is on new examples it never saw; lower is better. Example: 0.1321 for rank 1 at 200 examples.
- Adapter (add-on)
- The small set of extra numbers LoRA trains while the model stays frozen. Example: 128 numbers at rank 1.
- Frozen
- Left unchanged during training. Example: the toy’s pretrained layer
W0.
Go deeper: the engineer version
The kit's question
Why does LoRA r=1 never catch up, even with 200 examples?
The kit's answer
The change the task needs is rank 2, and a rank-1 adapter cannot represent it.
More detail: LoRA’s update is B·A with B of shape 64 x r and A of shape r x 64, so its rank is at most r. The toy’s target shift dW_true is built from two random directions, so it has rank 2 (toy_finetune.py line 38). A rank-1 B·A can capture at best the strongest single direction of dW_true (its top singular direction); the remaining direction stays as error however many examples you add. With --true-rank 16 every rank the script tries (up to 8) falls short the same way.
Day 11’s real training run starts from a model stored at 16 bits per number (bf16). On Day 8 you trained an add-on for a 7-billion-number model stored at 4 bits. Why must this comparison of full fine-tuning and LoRA use a 16-bit model, not a 4-bit one?Show answerHide
In plain words
Full fine-tuning changes every stored number, and 4-bit numbers are too coarse to take the tiny nudges training makes. Day 8’s QLoRA worked only because it never changed the 4-bit numbers: it trained a separate add-on instead.
Picture it
A 4-bit number is like a ruler marked only in whole centimetres. Training wants to move each mark by a hair’s width, and on that ruler the nudge rounds away to nothing. QLoRA leaves the coarse ruler alone and writes the fine corrections on a separate note.
With real numberslesson 17, and Day 8’s lesson 11
- 4 bits can hold only 2 x 2 x 2 x 2 = 16 different values; 16 bits can hold 65,536.
- Day 11’s model (Qwen2.5-0.5B-Instruct) at 16 bits: 490 million numbers x 2 bytes = 0.98 GB, small enough to train fully on a 16 GB Mac (about 8 to 10 GB at peak).
- Day 8’s QLoRA: a 7B model frozen at 4 bits plus a small 16-bit add-on, about 6 to 10 GB to train. A full fine-tune of the same model needs about 120 GB (it has 7.6 billion numbers: 7.6 x 16 = 122 GB).
Words to know
- bf16
- A 16-bit number format used for training, 2 bytes per number. Example: Day 11’s model is 0.98 GB in bf16.
- QLoRA
- LoRA on a model stored in 4 bits: the model stays frozen, only the add-on trains. Example: about 6 to 10 GB for a 7B model on Day 8.
- Frozen
- Left unchanged during training. Example: the 4-bit model in QLoRA.
- Gradient
- The correction worked out for each trainable number at each training step. Example: full fine-tuning keeps one for every number.
Go deeper: the engineer version
The kit's question
Why must the base be bf16 here, and not the 4-bit model from lesson 11?
The kit's answer
Full fine-tuning updates the weights, and you can’t train 4-bit weights directly. QLoRA only works because the base stays frozen.
More detail: Quantized weights sit on a coarse grid, and rounding to that grid gives no useful gradient, so a training step cannot move them reliably. QLoRA keeps the 4-bit base frozen (only read, and turned back into 16-bit numbers for the math) and sends gradients only into the 16-bit adapter. run_variants.sh uses the bf16 base for all three runs so that full, LoRA and DoRA compare fairly.
Day 11’s memory calculator (peft_params.py --model 8b) says LoRA needs about 16 GB to train, but full fine-tuning needs 128 GB. Where do LoRA’s 16 GB go?Show answerHide
In plain words
Almost all of it is the model itself, frozen and stored at 2 bytes per number. The add-on and its training state are tiny. Store the frozen model in 4 bits instead (QLoRA) and it drops to about 5 GB.
Picture it
You hire a moving truck to deliver one small box: nearly all the cost is the truck, not the box. In LoRA the truck is the frozen model and the box is the add-on. QLoRA books a smaller truck: the same model squeezed into 4 bits.
With real numberspeft_params.py --model 8b --rank 16, lesson 17
- Frozen model: 8 billion numbers x 2 bytes = 16 GB.
- LoRA add-on at rank 16: 13.6 million numbers (0.17% of the model) x 16 bytes = 0.22 GB.
- Total: 16 + 0.22 = 16.2 GB. Full: 8 billion x 16 bytes = 128 GB, about 8 times as much.
- QLoRA: the frozen model at 4 bits is half a byte per number, plus a little for shared scale numbers: 0.5625 bytes. 8 billion x 0.5625 = 4.5 GB, plus 0.22 = 4.7 GB, about 5 GB.
Words to know
- Frozen model
- The original model, kept unchanged and only read during LoRA training. Example: 16 GB for an 8B model at 2 bytes per number.
- Adapter (add-on)
- The small set of extra numbers LoRA trains and saves. Example: 13.6 million numbers for an 8B model at rank 16.
- Optimizer state
- The running averages the optimizer (Adam) keeps per trained number to size each nudge. Example: 8 of the 16 bytes per number.
- QLoRA
- LoRA on a model stored in 4 bits. Example: 4.7 GB for an 8B model instead of 16.2 GB.
Go deeper: the engineer version
The kit's question
peft_params.py --model 8b says LoRA needs about 16 GB, but full needs 128. Where does LoRA’s 16 GB go?
The kit's answer
Almost all of it is the frozen bf16 base, 2 bytes × 8B. The adapter’s optimizer state is tiny. Switch to QLoRA and it drops to about 5 GB.
More detail: peft_params.py counts LoRA as 2 B x base parameters (frozen bf16 weights, no gradients or optimizer state) + 16 B x adapter parameters (bf16 weight and gradient, fp32 Adam m and v, fp32 master copy). QLoRA’s 0.5625 B per base weight is 4 bits plus the shared scales of each small block. Activations come on top of every figure and grow with batch x sequence length; --grad-checkpoint trims them.
Today’s steps
Section titled “Today’s steps”2 steps, then the wrap-up.
-
Training toys, part 2
You watch a model learn to game its scorer (reward hacking). The scorer learned from people who could not check the ticket’s category, so the model trained against it stops getting the category right.
Video: RLHF 4:10
-
Preference tuning with DPO
DPO (direct preference optimization) teaches a model which of two answers is better, straight from pairs. You pair 800 correct ticket answers, each with one of four realistic mistakes, train on your Mac, and check the format improves without costing accuracy.
Video: Direct Preference Optimization 3:50
- Wrap-up and drill Put your numbers in one table, answer 7 questions out loud, explain 2 results and work 2 examples.
Your progress is saved on this device.