Skip to content

Day 12 · Preference tuning

Part 2 · Day 12 of 12

0 of 12 days done

About 1 h 55 Free (Fireworks DPO optional, paid) Mac and Fireworks

Today: Watch a model learn to game its scorer in a small simulation (reward hacking). Then teach a real model on your Mac to prefer clean answers with DPO (direct preference optimization), and check it did not get less accurate.

Example

Imagine a support bot rewarded for sounding helpful, while nobody checks whether it sent the ticket to the right team. It learns to give polished answers that earn higher scores even as its real help gets worse. Today’s small simulation makes that failure visible. Then you will train a real model to prefer better answers and check that it still chooses the right category.

By the end you’ll have

Before-and-after numbers for three things: how often the model answers in valid JSON, how often it picks the right category, and how clearly it prefers the approved answers.

Expected: Roughly, as in the lesson: valid JSON from about 80% to about 98%, the right category about 70% before and the same or higher after, and the margin (how strongly the model favours the approved answer, on a log scale) from about +1 to about +10.

The 6 numbers you write down

  1. How often the reward model agrees with its labellers (%)

    High agreement, yet the scorer still gets gamed: it copied the labellers’ blind spot along with their taste.

    Example: 71, as in the lesson (the script uses a fixed random seed, so your run should match)

  2. Leash setting where the average true value peaks (β)

    Looser than this, the model starts gaming the scorer: the rating still rises but the real value falls. In a real project this is the setting you tune on an eval.

    Example: 0.5: +4.30, the highest of all 8 rows; the script’s line under the table says the same.

  3. Average true value with a firm and a very loose leash (E[true] at β 0.50 and 0.02)

    How much real value the model’s answers carry on average, on a firm leash and on a very loose one. The drop is reward hacking in one pair of numbers.

    Example: +4.30 / +2.15 (the lesson’s rows)

  4. Valid JSON before and after DPO (%)

    The share of the 100 test tickets answered with JSON that fits the ticket format. DPO should push it up.

    Example: about 80 / 98, as the lesson expects

  5. Right category before and after DPO (%)

    The correctness check: it must not drop. If it does, taste has cost correctness.

    Example: about 70, then the same or higher, as the lesson expects

  6. Preference margin before and after DPO

    On average over the 100 checking pairs, how much more likely the model finds the chosen answer than the rejected one, on a log scale: +1 is about 2.7 times as likely, +10 about 22,000 times (e^1 = 2.72, e^10 = 22,026). It should grow a lot after DPO.

    Example: about +1 / +10, as the lesson expects

How you will use this: When a team wants better answers, first ask what feedback it has. Written examples can teach a desired format or task; pairs of better and worse answers can teach a preference; a program that checks each answer can provide another training signal. Today you tried the preference route and saw why any scorer must still be checked against what customers actually need.

Before you start

  1. Day 11’s LoRA adapter, if you trained it.

    Step 2 first merges the small add-on you trained on Day 11 (a LoRA adapter) into the model: train on examples first, then DPO. Without it, DPO still runs, but it polishes a model that gets only about 40% of ticket categories right. Quick check, from the 03-labs folder:

    ls lessons/17-full-vs-peft-mlx/adapters/lora

    A list of files means it is there. “No such file or directory” means Day 11’s training has not finished: VARIANTS=lora bash lessons/17-full-vs-peft-mlx/run_variants.sh trains only the LoRA one, in 5 to 15 minutes.

  2. The Python environment from Day 1, and internet for the first run.

    Run everything from the 03-labs folder after source .venv/bin/activate. The first DPO run installs mlx-lm-lora, the tool that trains DPO with MLX. If Day 11 did not already download it, it also fetches the 0.5-billion-number Qwen model, about 1 GB.

  3. Only for the optional Fireworks part: your key and the spend cap from Day 1.

    dpo_fireworks.sh trains the same job on Fireworks for cents, then offers to deploy it on a dedicated GPU, which bills by the hour. Skip it and the day stays free.

About 2 minutes. Say your answer out loud, then tap to check it. From Day 11 · Training methods in miniature.

In Day 11’s toy (one small layer of 4,096 numbers), the rank-1 add-on never catches up with full fine-tuning, even with 200 examples. Why? (The rank is how many separate kinds of change the add-on can make at once.)Show answerHide

In plain words

The change this task needs has two separate kinds (two directions), and a rank-1 add-on can make only one. More examples cannot supply the missing one; only a bigger rank can.

Picture it

Treasure lies to the north-east, and you may only move along a straight track running east. However many tries you get, you stop at the closest point on the track, never at the treasure. Add a second track running north and you reach it exactly.

With real numberslesson 16’s toy_finetune.py, the lesson’s sample run

  • The task needs a rank-2 change: two separate kinds of change (the script’s --true-rank is 2 by default).
  • Rank 1: 64 + 64 = 128 trained numbers, 3.1% of the layer. Error 0.3937, 0.1951 and 0.1321 at 16, 48 and 200 examples (0 is perfect).
  • Rank 2: 256 trained numbers, 6.2%. Error at 200 examples: 0.0000, as good as full fine-tuning (0.0002).
  • So rank 1 trains half as many numbers as rank 2 but misses one whole kind of change: at 200 examples its error is still 0.1321, while rank 2’s is 0.0000.

Words to know

Rank
How many separate kinds of change an add-on can make at once. Example: the toy’s task needs 2.
Held-out error
How wrong the trained layer is on new examples it never saw; lower is better. Example: 0.1321 for rank 1 at 200 examples.
Adapter (add-on)
The small set of extra numbers LoRA trains while the model stays frozen. Example: 128 numbers at rank 1.
Frozen
Left unchanged during training. Example: the toy’s pretrained layer W0.
Go deeper: the engineer version

The kit's question

Why does LoRA r=1 never catch up, even with 200 examples?

The kit's answer

The change the task needs is rank 2, and a rank-1 adapter cannot represent it.

More detail: LoRA’s update is B·A with B of shape 64 x r and A of shape r x 64, so its rank is at most r. The toy’s target shift dW_true is built from two random directions, so it has rank 2 (toy_finetune.py line 38). A rank-1 B·A can capture at best the strongest single direction of dW_true (its top singular direction); the remaining direction stays as error however many examples you add. With --true-rank 16 every rank the script tries (up to 8) falls short the same way.

Day 11’s real training run starts from a model stored at 16 bits per number (bf16). On Day 8 you trained an add-on for a 7-billion-number model stored at 4 bits. Why must this comparison of full fine-tuning and LoRA use a 16-bit model, not a 4-bit one?Show answerHide

In plain words

Full fine-tuning changes every stored number, and 4-bit numbers are too coarse to take the tiny nudges training makes. Day 8’s QLoRA worked only because it never changed the 4-bit numbers: it trained a separate add-on instead.

Picture it

A 4-bit number is like a ruler marked only in whole centimetres. Training wants to move each mark by a hair’s width, and on that ruler the nudge rounds away to nothing. QLoRA leaves the coarse ruler alone and writes the fine corrections on a separate note.

With real numberslesson 17, and Day 8’s lesson 11

  • 4 bits can hold only 2 x 2 x 2 x 2 = 16 different values; 16 bits can hold 65,536.
  • Day 11’s model (Qwen2.5-0.5B-Instruct) at 16 bits: 490 million numbers x 2 bytes = 0.98 GB, small enough to train fully on a 16 GB Mac (about 8 to 10 GB at peak).
  • Day 8’s QLoRA: a 7B model frozen at 4 bits plus a small 16-bit add-on, about 6 to 10 GB to train. A full fine-tune of the same model needs about 120 GB (it has 7.6 billion numbers: 7.6 x 16 = 122 GB).

Words to know

bf16
A 16-bit number format used for training, 2 bytes per number. Example: Day 11’s model is 0.98 GB in bf16.
QLoRA
LoRA on a model stored in 4 bits: the model stays frozen, only the add-on trains. Example: about 6 to 10 GB for a 7B model on Day 8.
Frozen
Left unchanged during training. Example: the 4-bit model in QLoRA.
Gradient
The correction worked out for each trainable number at each training step. Example: full fine-tuning keeps one for every number.
Go deeper: the engineer version

The kit's question

Why must the base be bf16 here, and not the 4-bit model from lesson 11?

The kit's answer

Full fine-tuning updates the weights, and you can’t train 4-bit weights directly. QLoRA only works because the base stays frozen.

More detail: Quantized weights sit on a coarse grid, and rounding to that grid gives no useful gradient, so a training step cannot move them reliably. QLoRA keeps the 4-bit base frozen (only read, and turned back into 16-bit numbers for the math) and sends gradients only into the 16-bit adapter. run_variants.sh uses the bf16 base for all three runs so that full, LoRA and DoRA compare fairly.

Day 11’s memory calculator (peft_params.py --model 8b) says LoRA needs about 16 GB to train, but full fine-tuning needs 128 GB. Where do LoRA’s 16 GB go?Show answerHide

In plain words

Almost all of it is the model itself, frozen and stored at 2 bytes per number. The add-on and its training state are tiny. Store the frozen model in 4 bits instead (QLoRA) and it drops to about 5 GB.

Picture it

You hire a moving truck to deliver one small box: nearly all the cost is the truck, not the box. In LoRA the truck is the frozen model and the box is the add-on. QLoRA books a smaller truck: the same model squeezed into 4 bits.

With real numberspeft_params.py --model 8b --rank 16, lesson 17

  • Frozen model: 8 billion numbers x 2 bytes = 16 GB.
  • LoRA add-on at rank 16: 13.6 million numbers (0.17% of the model) x 16 bytes = 0.22 GB.
  • Total: 16 + 0.22 = 16.2 GB. Full: 8 billion x 16 bytes = 128 GB, about 8 times as much.
  • QLoRA: the frozen model at 4 bits is half a byte per number, plus a little for shared scale numbers: 0.5625 bytes. 8 billion x 0.5625 = 4.5 GB, plus 0.22 = 4.7 GB, about 5 GB.

Words to know

Frozen model
The original model, kept unchanged and only read during LoRA training. Example: 16 GB for an 8B model at 2 bytes per number.
Adapter (add-on)
The small set of extra numbers LoRA trains and saves. Example: 13.6 million numbers for an 8B model at rank 16.
Optimizer state
The running averages the optimizer (Adam) keeps per trained number to size each nudge. Example: 8 of the 16 bytes per number.
QLoRA
LoRA on a model stored in 4 bits. Example: 4.7 GB for an 8B model instead of 16.2 GB.
Go deeper: the engineer version

The kit's question

peft_params.py --model 8b says LoRA needs about 16 GB, but full needs 128. Where does LoRA’s 16 GB go?

The kit's answer

Almost all of it is the frozen bf16 base, 2 bytes × 8B. The adapter’s optimizer state is tiny. Switch to QLoRA and it drops to about 5 GB.

More detail: peft_params.py counts LoRA as 2 B x base parameters (frozen bf16 weights, no gradients or optimizer state) + 16 B x adapter parameters (bf16 weight and gradient, fp32 Adam m and v, fp32 master copy). QLoRA’s 0.5625 B per base weight is 4 bits plus the shared scales of each small block. Activations come on top of every figure and grow with batch x sequence length; --grad-checkpoint trims them.

Start step 1: Training toys, part 2 (25 min)

Your progress is saved on this device.