Skip to content

Preference tuning with DPO

course 22 of 22

lesson 18 · lessons/18-preference-dpo

1 h 20 Free locally, Fireworks optional (paid) Mac, then Fireworks

What you will do and why

DPO (direct preference optimization) teaches a model which of two answers is better, straight from pairs. You pair 800 correct ticket answers, each with one of four realistic mistakes, train on your Mac, and check the format improves without costing accuracy.

Why it matters: A support-ticket system needs a reply it can read and the correct destination team. After training, count both on 100 tickets the model has never seen. In the lesson’s sample, readable replies improve from about 80 to 98, while correct categories must stay at least as common as before. A prettier reply is not an improvement if more tickets go to the wrong team.

You are done when: pref_eval.py has printed your before and after-dpo rows and both pairwise rows. Valid JSON went up, chatty went down, category_% did not drop, and the margin grew.

Start this first

Run it from the 03-labs folder with the Python environment on (source .venv/bin/activate). It builds the 800 pairs, installs the DPO training tool (mlx-lm-lora) the first time, merges Day 11’s LoRA add-on if it finds it, then trains for 10 to 20 minutes. Watch the video while it runs. For the commands below, open a second terminal, go to 03-labs, run source .venv/bin/activate, then cd lessons/18-preference-dpo.

Terminal window
cd lessons/18-preference-dpo && bash dpo_mlx.sh

Direct Preference Optimization · 3:50

Download mp4 (14.7 MB)

Chapters

In this video How DPO gets RLHF’s result from the same pairs with one simple loss, what β does, and how your own product already produces preference pairs.

3 key points

  1. DPO learns straight from chosen and rejected pairs with one simple loss: no reward model, no trial-and-error loop.

    Each step makes the chosen answer more likely and the rejected one less, compared with a frozen copy. In the video’s step the gap reaches 1.2, and with β (the leash setting) at 0.1 the loss falls from 0.69 to 0.63.

  2. It needs 2 models in memory instead of RLHF’s 4, and trains almost as steadily as ordinary fine-tuning.

    The 2 are the model being trained and the frozen reference copy. Pairs often come free from your product: an agent’s edit of a draft is chosen, the original draft rejected.

  3. Fine-tune on examples first, then DPO, and measure both win rate (how often the new answer is preferred to the old) and task accuracy.

    DPO refines a model that already does the task. Start β around 0.1; lower lets the model move further. Improving taste can quietly cost correctness, so check accuracy every time.

DPO trains a model on pairs of answers to the same question: one chosen, one rejected. It learns to make the chosen kind more likely and the rejected kind less. There is no separate scorer and no trial-and-error loop, only a loss (a score of how wrong the model still is, which training shrinks) computed from the pairs.

Picture it

A new support agent watches a supervisor edit drafts: this version, not that one. After hundreds of before-and-after edits, the agent drafts like the supervisor. Edits that only fix tone teach tone, not facts, which is why accuracy is checked separately.

With real numberslesson 18’s scripts and README, Qwen2.5-0.5B-Instruct on your Mac

  • Pairs: 800 for training and 100 for checking, one per support ticket.
  • Each rejected answer is one of four mistakes picked at random, so about 800 ÷ 4 = 200 of each: chatty (the right JSON wrapped in chat), prose (a paragraph, no JSON), wrong category, wrong severity.
  • Training: 200 steps x 2 pairs = 400 pairs seen, half of the 800, in 10 to 20 minutes.
  • Memory: two copies of the model, the trained one and a frozen reference. 0.5 billion numbers x 2 bytes = about 1 GB each.
  • Roughly, as the lesson expects: valid JSON about 80% to about 98%, chatty about 20% to about 1%, right category about 70% to the same or higher.

Words to know

DPO (direct preference optimization)
Training straight on chosen and rejected pairs with one loss, no scorer and no sampling loop. Example: dpo_mlx.sh.
Preference pair
Two answers to the same prompt, one chosen and one rejected. Example: clean JSON against the same JSON wrapped in chat.
Reference model
A frozen copy of the starting model that training measures drift against. Example: --reference-model-path in dpo_mlx.sh.
Loss
The single number training tries to shrink; for DPO it falls as the chosen and rejected answers pull apart. Example: 0.69 at the start, 0.63 after the video’s step.

From the lesson

lessons/18-preference-dpo/README.md

SFT teaches what to answer. DPO teaches which of two answers is better using pairs of chosen and rejected answers, with no reward model and no reinforcement loop. Our triage model sometimes wraps its JSON in chat (“Sure! Here’s the JSON…”), writes prose, or gets the category or severity wrong. Each of those is a rejected answer, and the clean, correct JSON is chosen.

prompt "Card declined twice, launch is tomorrow."
chosen {"category": "billing", "severity": 4, "next_action": "route to billing; verify charge"}
rejected Sure! Here's the JSON you asked for: {…} Let me know if you need anything else!

In a real product, these pairs come for free: an agent’s edit (chosen) of a model draft (rejected), a thumbs-down, a regenerate.

step file runs on does
1 make_pairs.py anywhere 800 train and 100 valid pairs, with four realistic failure kinds; written in mlx-lm-lora and Fireworks formats
2 dpo_mlx.sh Mac fuses Day 11’s LoRA into an SFT model if present (SFT first, then DPO), then mlx_lm_lora.train --train-mode dpo --beta 0.1
3 dpo_fireworks.sh Mac → FW firectl dpo-job create --loss-method DPO; optional deployment and eval, with auto-delete
4 pref_eval.py any target valid JSON %, chatty %, category and severity accuracy, answer length; with --pairwise, DPO’s implicit reward (chosen vs rejected log-prob)
Terminal window
cd lessons/18-preference-dpo
python make_pairs.py
python pref_eval.py --target mock --label mock # rehearse the metrics offline
bash dpo_mlx.sh # 10–20 min
# two terminals, as the script prints:
# mlx_lm.server --model <START> --port 8081
# mlx_lm.server --model <START> --adapter-path adapters/dpo --port 8082
MLX_MODEL=<START> python pref_eval.py --target mlx --label before
MLX_URL=http://localhost:8082/v1 MLX_MODEL=<START> python pref_eval.py --target mlx --label after-dpo
python pref_eval.py --pairwise <START> --label pairs-before
python pref_eval.py --pairwise <START> --adapter adapters/dpo --label pairs-after
bash dpo_fireworks.sh # optional, paid

What each command does

  1. python make_pairs.py

    Turns each support ticket into a pair: the correct JSON as chosen, one realistic mistake as rejected, in a Mac format and a Fireworks format. Look for train: 800 pairs, valid: 100 pairs, the four mistake kinds at about 200 each (191 to 213), and one example pair. If dpo_mlx.sh is already running from before the video, it built these: read its output and skip this command.

  2. python pref_eval.py --target mock --label mock

    First start the practice server: from 03-labs, run make mock &. Then this rehearses against it: it sends the 100 test tickets and prints the five measures you compare later. The numbers mean nothing; it proves the scoring works.

  3. bash dpo_mlx.sh

    Trains DPO on your Mac for 10 to 20 minutes, as a small add-on saved in adapters/dpo. It first prints DPO starting from: and a model: sft_model if Day 11’s LoRA was merged in, else the plain instruct model. That model is <START> below. At the end it prints two server commands: run each in its own terminal and leave them running. If you started it before the video, let that run finish.

  4. MLX_MODEL=<START> python pref_eval.py --target mlx --label before

    Scores the model before DPO on the 100 test tickets, through the server on port 8081 (no add-on). Replace <START> with the model dpo_mlx.sh printed. Look for valid_% near 80 and chatty_% near 20, roughly as the lesson expects; yours will differ.

  5. MLX_URL=http://localhost:8082/v1 MLX_MODEL=<START> python pref_eval.py --target mlx --label after-dpo

    The same 100 tickets through the server on port 8082, which adds the DPO add-on. Look for valid_% near 98, chatty_% near 1, and category_% the same as before or higher.

  6. python pref_eval.py --pairwise <START> --label pairs-before

    No server needed: it loads the model itself. For each of the 100 checking pairs it compares how likely the model is to write the chosen answer against the rejected one. Look for pref_acc_% near 60 and a margin near +1. The row prints as pairwise base: the script names pairwise rows itself.

  7. python pref_eval.py --pairwise <START> --adapter adapters/dpo --label pairs-after

    The same with the DPO add-on loaded. Look for pref_acc_% near 95 and a margin near +10: the model now clearly favours the chosen answers. The row prints as pairwise +dpo. 3 of the 100 checking pairs are identical (the kit’s wrong-severity mistake cannot push a severity-5 ticket higher, or a severity-1 ticket lower), so 97% is the ceiling.

  8. bash dpo_fireworks.sh

    Optional and paid. Runs the same DPO job on Fireworks, on Qwen2.5 7B Instruct by default. Training costs cents: the lab book lists $1.00 per million training tokens for LoRA DPO up to 16B. It asks before deploying, because a dedicated deployment bills by the hour ($8 an hour for an H100 in the lab book); the script deletes it on exit. Afterwards run make fw-check: it should list nothing.

<START> is whichever model dpo_mlx.sh printed: sft_model/ if you trained Day 11’s LoRA first, otherwise the instruct model.

How to read it

Only the pattern matters; your numbers will differ. The first table scores 100 test tickets before and after DPO: valid_% JSON that fits the ticket format, chatty_% chat or prose, category_% and severity_% right answers, avg_chars average length (≈ means about the same, the up arrow higher). In the second, pref_acc_% is how often the model favours the chosen answer and margin by how much; format should jump while category holds.

label valid_% chatty_% category_% severity_% avg_chars
before ~80 ~20 ~70 ~60 ~110
after-dpo ~98 ~1 ≈ or ↑ ≈ or ↑ ~85 ← format fixed, no correctness lost
label pref_acc_% margin
pairwise base ~60 +1.x
pairwise +dpo ~95 +10.x ← the implicit reward grew

If category_% drops after DPO, taste has cost correctness (the Direct Preference Optimization video at the top of this step). Try a larger β, fewer iterations, or more wrong_cat pairs so correctness is part of what “preferred” means.

4 questions. Say your answer out loud, then tap to check it.

Why merge Day 11’s small add-on (the SFT adapter, trained on examples) into the model first, instead of running DPO on the plain instruct model?Show answerHide

In plain words

DPO polishes a model that already does the task. The plain model gets only about 40% of ticket categories right, so the approved answers are still unlikely for it, and DPO mostly learns to avoid the rejected text.

Picture it

Coaching someone’s handwriting helps once they can form the letters. Show a beginner pairs of neat and messy pages and they mostly learn what to avoid, not how to write.

With real numberslesson 17’s README (Day 11), dpo_mlx.sh and the RLHF and DPO videos

  • Day 11’s expected table: before its training, the model gets about 40% of ticket categories right; after training, all three methods score well above that.
  • dpo_mlx.sh merges Day 11’s LoRA into sft_model, then runs DPO from there and uses it as the frozen reference too.
  • Without it, DPO starts from the plain instruct model, the one at about 40%.
  • The RLHF video’s rule: “You cannot reinforce behaviour the model never produces.” The DPO video’s: SFT first, then DPO.

Words to know

SFT (supervised fine-tuning)
Training on examples of the answers you want. Example: Day 11’s LoRA on 800 support tickets.
Adapter (LoRA)
A small add-on of extra learned numbers trained on top of a frozen model. Example: Day 11’s, in adapters/lora.
Fuse
Bake an adapter’s numbers into the model so it becomes one self-contained model. Example: sft_model.
Instruct model
A model already trained to follow instructions and chat. Example: Qwen2.5-0.5B-Instruct.
Go deeper: the engineer version

The kit's question

Why fuse the SFT adapter first rather than running DPO on the raw instruct model?

The kit's answer

DPO refines a model that already does the task. With nothing to refine, it mostly learns “not the rejected text”.

More detail: DPO only reshapes probability between responses the model can already produce. Starting from a model that already emits valid triage JSON, the pairs refine format and correctness at the margin. From the raw instruct model the chosen answers are unlikely, so much of the loss is reduced by pushing the rejected text down. The same model is the frozen reference (--reference-model-path "$START"), so the reference also knows the task.

What does β (beta, the leash setting, 0.1 in dpo_mlx.sh) actually do?Show answerHide

In plain words

It sets how far the model may drift from its frozen starting copy before training stops pushing it further. A smaller β lets it move further.

Picture it

Think of β as how short the leash is. On a short leash (large β), the dog gets a few steps away and the leash goes taut. On a long leash (small β), it wanders much further before it is pulled back.

With real numbersthe DPO video’s worked step, recomputed at other β values

  • Margin: the chosen answer +0.8, the rejected -0.4, both against the frozen copy: 0.8 - (-0.4) = 1.2.
  • β 0.1: 0.1 x 1.2 = 0.12. The sigmoid σ turns any number into a value between 0 and 1: σ(0.12) = 0.53. Loss = -log(0.53) = 0.63, down from 0.69 when the margin was 0.
  • β 0.05: the same margin earns only 0.06, loss 0.66, so training keeps pulling the two apart for longer.
  • β 0.2: 0.24, loss 0.58, so less drift is needed before the pull fades.

Words to know

β (beta)
The leash: how far the model may move from the reference. Example: 0.1 in dpo_mlx.sh.
Margin
How much more the model favours the chosen answer than the rejected one. Example: 1.2 in the DPO video.
Log-probability
The natural log of how likely the model is to write an answer; DPO works in these. Example: +0.8 against the reference.
Loss
The single number training tries to shrink. Example: 0.69 at the start, 0.63 after one step.
Go deeper: the engineer version

The kit's question

What does β do here, concretely?

The kit's answer

It scales the margin in the loss. A smaller β lets the policy move further from the reference before the loss stops rewarding it.

More detail: L = −log σ(β · [(log π(y_c) − log π_ref(y_c)) − (log π(y_r) − log π_ref(y_r))]). Its gradient is scaled by σ(−β · margin), which shrinks as β · margin grows, so a larger β saturates at a smaller margin (a tighter leash). β is the same KL coefficient as in RLHF: DPO is derived from that KL-regularised objective, with the leash built into the loss. Set it with BETA=0.2 bash dpo_mlx.sh.

Suppose half your rejected answers are chatty (the right JSON wrapped in chat), and none have the wrong category (wrong_cat). What happens to category accuracy after DPO?Show answerHide

In plain words

The model learns to fix its format, not to get categories right. Category accuracy may stay flat or even slip, because nothing in the pairs rewards the right category.

Picture it

If a supervisor only ever corrects greetings and sign-offs, the new agent writes polite emails that may still send customers to the wrong team.

With real numberslesson 18’s make_pairs.py and README

  • The lesson’s 800 pairs: each rejected answer is one of four mistakes, about 200 each (800 ÷ 4).
  • Only the wrong-category kind differs from the chosen answer in its category. Chatty, prose and wrong-severity answers all keep the right one.
  • The question’s mix: 400 chatty (half of 800) and 0 wrong-category, so no pair teaches the category.
  • Watch category_%: the lesson expects it to hold at about 70% or rise. A drop means taste has cost correctness.

Words to know

Chatty answer
The right JSON wrapped in friendly chat, so a program cannot read it. Example: “Sure! Here’s the JSON you asked for: ...” (chatty_% also counts prose answers.)
wrong_cat pair
A pair whose rejected answer is valid JSON with the wrong category. Example: 213 of the 800.
Category accuracy
The share of test tickets given the right category. Example: category_%, about 70% before.
Preference pair
Two answers to the same prompt, one chosen and one rejected. Example: 800 in this lesson.
Go deeper: the engineer version

The kit's question

Your pairs are 50% “chatty”. What happens to accuracy if none of them are wrong_cat?

The kit's answer

The model learns format, not correctness, and category accuracy may even slip.

More detail: DPO raises the log-probability gap along whatever separates chosen from rejected. If no pair separates on category, the gradient carries no category signal, and drift from the reference (bounded by β) can move it either way. The README’s fixes: a larger β, fewer iterations, or more wrong_cat pairs, so correctness is part of what “preferred” means. make_pairs.py’s own rule: pairs should differ in the thing you care about.

When would you use ORPO (odds-ratio preference optimization, a relative of DPO) instead of DPO? Fireworks runs it with --loss-method ORPO.Show answerHide

In plain words

When you want to teach the task and the preference in one training run, without keeping a second, frozen copy of the model in memory.

Picture it

DPO is coaching after a course: first the course, then a coach who compares you with your old self. ORPO is one class that teaches the skill and the taste together, with no old self to compare against.

With real numbersthe DPO and RLHF videos, dpo_mlx.sh and dpo_fireworks.sh

  • RLHF with PPO: 4 models in memory. DPO: 2, the model being trained and a frozen reference.
  • ORPO: no reference copy, so only the model being trained, and no separate example-training step first.
  • On your Mac, each copy of the 0.5-billion-number model is 0.5 billion x 2 bytes = about 1 GB.
  • For the Qwen2.5 7B model dpo_fireworks.sh uses (7.6 billion numbers), a frozen copy at 2 bytes each is 7.6 x 2 = about 15 GB.

Words to know

ORPO (odds-ratio preference optimization)
A preference method that does example training and preference training in one run, with no reference model. Example: --loss-method ORPO.
Reference model
A frozen copy of the starting model that training measures drift against. Example: one of DPO’s 2 models.
SFT (supervised fine-tuning)
Training on examples of the answers you want. Example: Day 11’s LoRA.
bf16
A 16-bit number format, 2 bytes per number, used for training. Example: the 0.5B model is about 1 GB in bf16.
Go deeper: the engineer version

The kit's question

When would you use ORPO instead (--loss-method ORPO on Fireworks)?

The kit's answer

When you want SFT and preference tuning in one step without holding a reference model in memory.

More detail: ORPO adds an odds-ratio term, −λ · log σ(log(odds(chosen) / odds(rejected))) with odds(y) = P(y) ÷ (1 − P(y)), to the ordinary SFT loss on the chosen answer, so it needs neither a reference model nor a separate SFT stage. With no reference there is no explicit leash to a known-good model, so evaluate accuracy as carefully as with DPO. On Fireworks: firectl dpo-job create --loss-method ORPO with --orpo-lambda (the comment at the top of dpo_fireworks.sh).

Question Our fine-tuned support model still wraps its JSON in chat and breaks our parser. We have no budget for labellers. What would you do?

One clear answer

Their agents already edit the model’s drafts. Every edit is a preference pair. Run DPO on an SFT’d model and the formatting failures disappear. I’d watch category accuracy so taste doesn’t cost correctness, and where correctness can be checked by code, I’d use RFT with a grader instead.

What this means

  • “Their agents already edit the model’s drafts”: Support staff already fix the model’s replies as part of their day.
  • “Every edit is a preference pair”: Each fix gives two answers to the same ticket: the edited one (chosen) and the original draft (rejected). That is free training data.
  • “Run DPO on an SFT’d model”: Train on those pairs with DPO, starting from the model already fine-tuned on examples (SFT). Here that is Day 11’s LoRA, merged in.
  • “the formatting failures disappear”: As the lesson expects: valid JSON from about 80% to about 98%, chatty replies from about 20% to about 1%.
  • “I’d watch category accuracy so taste doesn’t cost correctness”: I compare the right-category rate before and after, because pairs that differ only in style can quietly make answers less accurate.
  • “where correctness can be checked by code, I’d use RFT with a grader instead”: If a program can mark an answer right or wrong, I train with RFT (reinforcement fine-tuning), which scores each answer with that program instead of preferences.

Your numbersSaved on this device and collected in the Day 12 wrap-up.

Hint: valid_% in the before and after-dpo rows of pref_eval.py. Write both, for example “80 / 98”.

Hint: category_%, same two rows. Write both as before / after, for example “70 / 70”.

Hint: margin in the pairwise base and pairwise +dpo rows. Write both as before / after; the lesson shows about +1 and about +10.

pref_eval.py has printed your before and after-dpo rows and both pairwise rows. Valid JSON went up, chatty went down, category_% did not drop, and the margin grew.

pref_eval.py --target mock says Connection refused.
The practice server is not running. From the 03-labs folder run make mock &, then try again.
pref_eval.py --target mlx says Connection refused.
The two model servers are not running. dpo_mlx.sh printed their commands at the end. In each new terminal, first go to 03-labs and run source .venv/bin/activate, then cd lessons/18-preference-dpo. Start mlx_lm.server --model <START> --port 8081 in one terminal and the same with --adapter-path adapters/dpo --port 8082 in another. When both are up, run the check again.
You do not know what to type for <START>.
Use the model dpo_mlx.sh printed on its DPO starting from: line, such as the full path to sft_model. It is also saved in the file .state in the lesson folder.
dpo_mlx.sh starts from the instruct model, not sft_model.
It did not find Day 11’s LoRA in lessons/17-full-vs-peft-mlx/adapters/lora. DPO still runs, but on a model that gets only about 40% of categories right (check yourself 1). From 03-labs, run VARIANTS=lora bash lessons/17-full-vs-peft-mlx/run_variants.sh (5 to 15 minutes), then bash lessons/18-preference-dpo/dpo_mlx.sh again, and restart both servers with the new <START> it prints (sft_model).
mlx_lm_lora.train stops with an unknown option or argument error.
The tool’s options change between versions, and the script says its --help wins. Run mlx_lm_lora.train --help, match the options at the top of dpo_mlx.sh to it, and run it again.
category_% dropped after DPO.
Taste has cost correctness. The lesson’s fixes: a larger β (BETA=0.2 bash dpo_mlx.sh), fewer steps (ITERS=100 bash dpo_mlx.sh), or more wrong-category pairs, so correctness is part of what “preferred” means.
The pairwise rows are not in results/18-pref.csv.
They have different columns, so the kit’s recorder saves each one in a new file named results/18-pref-<time>.csv. Copy the numbers from the terminal or from those files.
dpo_fireworks.sh says the base model is not found or not DPO-enabled.
Model ids change. Pick one the Fireworks model library marks as DPO-enabled and run BASE_MODEL=accounts/fireworks/models/<id> bash dpo_fireworks.sh. Afterwards run make fw-check: it should list nothing.
dpo_fireworks.sh Bash · 39 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 18 · step 3 — the same DPO job on Fireworks. ⏱ PAID (training + a deployment)
#
# firectl dpo-job create --loss-method DPO (or ORPO with --orpo-lambda; see video 18)
#
# The base model must be one the Fireworks model library marks as DPO-enabled. Override
# with BASE_MODEL=… Training is billed per token (cents for 800 short pairs). Evaluating
# needs a dedicated deployment (⏱ per hour), which a trap deletes on exit, even on Ctrl-C.
# Flags follow the Fireworks docs at time of writing; `firectl dpo-job create --help` wins.
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
cd "$(dirname "$0")"
BASE_MODEL=${BASE_MODEL:-accounts/fireworks/models/qwen2p5-7b-instruct}
ACCOUNT=${FIREWORKS_ACCOUNT_ID:-$(firectl whoami 2>/dev/null | grep -oE 'accounts/[a-z0-9-]+' | head -1 | cut -d/ -f2)}
STAMP=$(date +%m%d%H%M)
DATASET="triage-dpo-$STAMP"; OUT="triage-dpo-$STAMP"
[[ -f data/fireworks/train.jsonl ]] || python make_pairs.py
firectl dataset create "$DATASET" data/fireworks/train.jsonl
out=$(firectl dpo-job create --loss-method DPO --base-model "$BASE_MODEL" \
--dataset "accounts/$ACCOUNT/datasets/$DATASET" --output-model "$OUT")
echo "$out"
JOB=$(echo "$out" | grep -oE 'dpoJobs/[A-Za-z0-9-]+' | head -1 | cut -d/ -f2)
echo "▸ job $JOB — polling every 60 s (Ctrl-C is safe; the job keeps running)"
until firectl dpo-job get "$JOB" | grep -qE "COMPLETED"; do
firectl dpo-job get "$JOB" | grep -qE "FAILED|CANCELLED" && { firectl dpo-job get "$JOB"; exit 1; }
sleep 60; echo -n "."
done
echo " trained: accounts/$ACCOUNT/models/$OUT"
read -r -p "Deploy for evaluation now? This starts per-hour billing. [y/N] " yn
[[ "$yn" =~ ^[Yy]$ ]] || { echo "Skipped. Deploy later with lesson 10's 3_deploy.sh pattern."; exit 0; }
dep=$(firectl deployment create "accounts/$ACCOUNT/models/$OUT" --deployment-shape default)
DEP=$(echo "$dep" | grep -oE 'deployments/[A-Za-z0-9-]+' | head -1 | cut -d/ -f2)
trap 'echo "▸ deleting $DEP"; firectl deployment delete "$DEP"; firectl deployment list' EXIT
until firectl deployment get "$DEP" | grep -qiE "state.*READY"; do sleep 20; echo -n "."; done; echo " ready"
python pref_eval.py --target fireworks --model "$BASE_MODEL" --label fw-base
python pref_eval.py --target fireworks --model "accounts/$ACCOUNT/models/$OUT#accounts/$ACCOUNT/deployments/$DEP" --label fw-dpo
dpo_mlx.sh Bash · 46 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 18 · step 2 — DPO on your Mac with mlx-lm-lora (video 18).
#
# Rule from the video: SFT first, then DPO. If lesson 17's LoRA adapter exists we fuse it
# into an SFT model and run DPO on top of that; otherwise we start from the instruct model.
# The same model is used as the frozen REFERENCE, which is what β measures distance from.
#
# --train-mode dpo the DPO loss (also: orpo, cpo, grpo, ppo … see --help)
# --beta 0.1 the leash (video 18: start around 0.1)
# --dpo-cpo-loss-type sigmoid the original DPO loss: −log σ(β · margin)
# --reference-model-path the frozen copy the policy is compared against
#
# Flags follow the mlx-lm-lora README at time of writing; `mlx_lm_lora.train --help` wins.
# ~10–20 min on M-series for 0.5B. Two models in memory (policy + reference).
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
cd "$(dirname "$0")"
BASE=${MODEL:-mlx-community/Qwen2.5-0.5B-Instruct-bf16}
L17=../17-full-vs-peft-mlx/adapters/lora
python -c "import mlx_lm_lora" 2>/dev/null || pip install -U mlx-lm-lora
[[ -f data/mlx/train.jsonl ]] || python make_pairs.py
if [[ -d "$L17" && ! -d sft_model ]]; then
echo "▸ fusing lesson 17's LoRA into an SFT model (SFT first, then DPO)"
mlx_lm.fuse --model "$BASE" --adapter-path "$L17" --save-path sft_model
fi
START=$([[ -d sft_model ]] && echo "$(pwd)/sft_model" || echo "$BASE")
echo "▸ DPO starting from: $START"
mlx_lm_lora.train --model "$START" --reference-model-path "$START" \
--train --train-mode dpo --data data/mlx \
--beta ${BETA:-0.1} --dpo-cpo-loss-type sigmoid \
--iters ${ITERS:-200} --batch-size 2 --learning-rate 5e-6 \
--adapter-path adapters/dpo 2>&1 | tee dpo_train.log
echo "START_MODEL=\"$START\"" > .state
cat <<EOT
▸ serve before and after (two terminals), then compare:
mlx_lm.server --model "$START" --port 8081
mlx_lm.server --model "$START" --adapter-path adapters/dpo --port 8082
MLX_MODEL="$START" python pref_eval.py --label before --target mlx
MLX_URL=http://localhost:8082/v1 MLX_MODEL="$START" python pref_eval.py --label after-dpo --target mlx
EOT
make_pairs.py Python · 86 lines
"""
Lesson 18 · step 1 — build preference pairs from the ticket data (video 18).
Each training ticket becomes one pair with the SAME prompt:
chosen the correct, concise JSON answer (what a reviewer approves)
rejected one realistic failure (what a reviewer would edit or thumbs-down):
chatty "Sure! Here's the JSON: …" → breaks the parser
prose a helpful paragraph, no JSON → breaks the parser
wrong_cat valid JSON, wrong category → wrong answer
wrong_sev valid JSON, severity off by 2 → wrong judgement
Pairs should differ in the thing you care about. Here that is "parseable AND right".
In a real product, these come from agent edits, thumbs-down and regenerate clicks.
Writes two formats of the same pairs:
data/mlx/{train,valid}.jsonl {"system", "prompt", "chosen", "rejected"} (mlx-lm-lora)
data/fireworks/train.jsonl {"input": {"messages": [...]}, "preferred_output": [...],
"non_preferred_output": [...]} (firectl dpo-job)
python make_pairs.py
"""
import json
import random
from pathlib import Path
HERE = Path(__file__).parent
SRC = HERE.parents[1] / "data" / "tickets"
CATS = ["billing", "outage", "how-to", "abuse"]
rng = random.Random(18)
def rejected_for(ticket: str, ans: dict) -> tuple[str, str]:
kind = rng.choice(["chatty", "prose", "wrong_cat", "wrong_sev"])
if kind == "chatty":
return kind, "Sure! Here's the JSON you asked for:\n" + json.dumps(ans) + "\nLet me know if you need anything else!"
if kind == "prose":
art = "an" if ans["category"][0] in "aeiou" else "a"
return kind, (f"This looks like {art} {ans['category']} issue. I'd rate it around {ans['severity']} out of 5, "
f"and the next step would be to {ans['next_action']}.")
if kind == "wrong_cat":
wrong = dict(ans, category=rng.choice([c for c in CATS if c != ans["category"]]))
return kind, json.dumps(wrong)
wrong = dict(ans, severity=max(1, min(5, ans["severity"] + rng.choice([-2, 2]))))
return kind, json.dumps(wrong)
def build(split: str) -> list[dict]:
pairs = []
for line in (SRC / f"{split}.jsonl").read_text().splitlines():
msgs = json.loads(line)["messages"]
system, user, answer = msgs[0]["content"], msgs[1]["content"], msgs[2]["content"]
kind, bad = rejected_for(user, json.loads(answer))
pairs.append({"system": system, "prompt": user, "chosen": answer, "rejected": bad, "kind": kind})
return pairs
def main() -> None:
if not (SRC / "train.jsonl").exists():
raise SystemExit("run: python data/make_tickets.py")
(HERE / "data" / "mlx").mkdir(parents=True, exist_ok=True)
(HERE / "data" / "fireworks").mkdir(parents=True, exist_ok=True)
kinds = {}
for split in ("train", "valid"):
pairs = build(split)
with (HERE / "data" / "mlx" / f"{split}.jsonl").open("w") as f:
for p in pairs:
f.write(json.dumps({k: p[k] for k in ("system", "prompt", "chosen", "rejected")}) + "\n")
kinds[p["kind"]] = kinds.get(p["kind"], 0) + (split == "train")
if split == "train":
with (HERE / "data" / "fireworks" / "train.jsonl").open("w") as f:
for p in pairs:
f.write(json.dumps({
"input": {"messages": [{"role": "system", "content": p["system"]},
{"role": "user", "content": p["prompt"]}]},
"preferred_output": [{"role": "assistant", "content": p["chosen"]}],
"non_preferred_output": [{"role": "assistant", "content": p["rejected"]}],
}) + "\n")
print(f"{split}: {len(pairs)} pairs")
print("rejected kinds (train): " + ", ".join(f"{k} {v}" for k, v in sorted(kinds.items())))
print(f"→ {HERE / 'data'}")
ex = build("valid")[0]
print(f"\nexample\n prompt {ex['prompt']}\n chosen {ex['chosen']}\n rejected {ex['rejected']} ({ex['kind']})")
if __name__ == "__main__":
main()
pref_eval.py Python · 93 lines
"""
Lesson 18 · step 4 — did preference tuning move what we wanted, and break anything?
Behaviour on the 100 held-out tickets (any --target: mock, mlx, fireworks …):
valid_% strict JSON that matches the schema ← DPO should push this up
chatty_% answers wrapped in chat ("Sure! …") or prose ← and this down
category_% right category ← must not drop (taste can cost correctness)
severity_% right severity
avg_chars answer length ← watch for length drift
Pairwise, on the validation pairs (Apple silicon, --pairwise MODEL [--adapter PATH]):
pref_acc_% how often the model finds CHOSEN more likely than REJECTED (log-prob sum)
margin mean log-prob gap: DPO's implicit reward. It should grow after training.
python pref_eval.py --target mock --label mock
python pref_eval.py --target mlx --label before
python pref_eval.py --pairwise mlx-community/Qwen2.5-0.5B-Instruct-bf16 --adapter adapters/dpo
"""
import argparse
import json
from pathlib import Path
from felab import add_target_args, banner, client, record, resolve, table
from felab.tickets import SYSTEM, grade, load_test, parse
HERE = Path(__file__).parent
def behaviour(a) -> dict:
t = resolve(a)
banner(t)
cli = client(t)
rows = load_test()
g, lengths, chatty = [], [], 0
for r in rows:
resp = cli.chat.completions.create(model=t.model, temperature=0, max_tokens=120,
messages=[{"role": "system", "content": SYSTEM},
{"role": "user", "content": r["ticket"]}])
out = (resp.choices[0].message.content or "").strip()
g.append(grade(out, r))
lengths.append(len(out))
chatty += parse(out) is None # not parseable as a bare JSON object → chat or prose
n = len(rows)
return {"label": a.label, "target": t.name,
"valid_%": 100 * sum(x["valid"] for x in g) / n, "chatty_%": 100 * chatty / n,
"category_%": 100 * sum(x["category_ok"] for x in g) / n,
"severity_%": 100 * sum(x["severity_ok"] for x in g) / n,
"avg_chars": sum(lengths) / n}
def pairwise(model_path: str, adapter: str | None, limit: int) -> dict:
"""Sum of token log-probs of each response given the prompt: DPO's own currency."""
import mlx.core as mx
from mlx_lm import load
model, tok = load(model_path, adapter_path=adapter)
def logprob(system: str, prompt: str, response: str) -> float:
p_txt = tok.apply_chat_template([{"role": "system", "content": system}, {"role": "user", "content": prompt}],
add_generation_prompt=True, tokenize=False)
p_ids = tok.encode(p_txt)
r_ids = tok.encode(response, add_special_tokens=False)
ids = p_ids + r_ids
logits = model(mx.array([ids[:-1]]))[0]
logp = logits - mx.logsumexp(logits, axis=-1, keepdims=True)
tgt = mx.array(ids[1:])
tok_lp = mx.take_along_axis(logp, tgt[:, None], axis=-1)[:, 0]
return tok_lp[len(p_ids) - 1:].sum().item() # only the response tokens
pairs = [json.loads(line) for line in (HERE / "data/mlx/valid.jsonl").read_text().splitlines()][:limit]
margins = [logprob(p["system"], p["prompt"], p["chosen"]) - logprob(p["system"], p["prompt"], p["rejected"])
for p in pairs]
return {"label": f"pairwise {'+' + Path(adapter).name if adapter else 'base'}",
"pref_acc_%": 100 * sum(m > 0 for m in margins) / len(margins),
"margin": sum(margins) / len(margins), "pairs": len(margins)}
def main() -> None:
ap = add_target_args(argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter))
ap.add_argument("--label", default="run")
ap.add_argument("--pairwise", metavar="MODEL", help="local MLX model path/repo for log-prob scoring")
ap.add_argument("--adapter", help="adapter path for --pairwise (e.g. adapters/dpo)")
ap.add_argument("--limit", type=int, default=100)
a = ap.parse_args()
row = pairwise(a.pairwise, a.adapter, a.limit) if a.pairwise else behaviour(a)
print("\n" + table([row]))
record("18-pref", row)
if not a.pairwise:
print("\nRun it before and after DPO with different --label values; results/18-pref.csv keeps both.")
if __name__ == "__main__":
main()