Preference tuning with DPO
course 22 of 22
lesson 18 · lessons/18-preference-dpo
1 h 20 Free locally, Fireworks optional (paid) Mac, then Fireworks
What you will do and why
DPO (direct preference optimization) teaches a model which of two answers is better, straight from pairs. You pair 800 correct ticket answers, each with one of four realistic mistakes, train on your Mac, and check the format improves without costing accuracy.
Why it matters: A support-ticket system needs a reply it can read and the correct destination team. After training, count both on 100 tickets the model has never seen. In the lesson’s sample, readable replies improve from about 80 to 98, while correct categories must stay at least as common as before. A prettier reply is not an improvement if more tickets go to the wrong team.
You are done when: pref_eval.py has printed your before and after-dpo rows and both pairwise rows. Valid JSON went up, chatty went down, category_% did not drop, and the margin grew.
Start this first
Run it from the 03-labs folder with the Python environment on (source .venv/bin/activate). It builds the 800 pairs, installs the DPO training tool (mlx-lm-lora) the first time, merges Day 11’s LoRA add-on if it finds it, then trains for 10 to 20 minutes. Watch the video while it runs. For the commands below, open a second terminal, go to 03-labs, run source .venv/bin/activate, then cd lessons/18-preference-dpo.
cd lessons/18-preference-dpo && bash dpo_mlx.shDirect Preference Optimization · 3:50
Download mp4 (14.7 MB)Chapters
In this video How DPO gets RLHF’s result from the same pairs with one simple loss, what β does, and how your own product already produces preference pairs.
3 key points
DPO learns straight from chosen and rejected pairs with one simple loss: no reward model, no trial-and-error loop.
Each step makes the chosen answer more likely and the rejected one less, compared with a frozen copy. In the video’s step the gap reaches 1.2, and with β (the leash setting) at 0.1 the loss falls from 0.69 to 0.63.
It needs 2 models in memory instead of RLHF’s 4, and trains almost as steadily as ordinary fine-tuning.
The 2 are the model being trained and the frozen reference copy. Pairs often come free from your product: an agent’s edit of a draft is chosen, the original draft rejected.
Fine-tune on examples first, then DPO, and measure both win rate (how often the new answer is preferred to the old) and task accuracy.
DPO refines a model that already does the task. Start β around 0.1; lower lets the model move further. Improving taste can quietly cost correctness, so check accuracy every time.
In plain words
Section titled “In plain words”DPO trains a model on pairs of answers to the same question: one chosen, one rejected. It learns to make the chosen kind more likely and the rejected kind less. There is no separate scorer and no trial-and-error loop, only a loss (a score of how wrong the model still is, which training shrinks) computed from the pairs.
Picture it
A new support agent watches a supervisor edit drafts: this version, not that one. After hundreds of before-and-after edits, the agent drafts like the supervisor. Edits that only fix tone teach tone, not facts, which is why accuracy is checked separately.
With real numberslesson 18’s scripts and README, Qwen2.5-0.5B-Instruct on your Mac
- Pairs: 800 for training and 100 for checking, one per support ticket.
- Each rejected answer is one of four mistakes picked at random, so about 800 ÷ 4 = 200 of each: chatty (the right JSON wrapped in chat), prose (a paragraph, no JSON), wrong category, wrong severity.
- Training: 200 steps x 2 pairs = 400 pairs seen, half of the 800, in 10 to 20 minutes.
- Memory: two copies of the model, the trained one and a frozen reference. 0.5 billion numbers x 2 bytes = about 1 GB each.
- Roughly, as the lesson expects: valid JSON about 80% to about 98%, chatty about 20% to about 1%, right category about 70% to the same or higher.
Words to know
- DPO (direct preference optimization)
- Training straight on chosen and rejected pairs with one loss, no scorer and no sampling loop. Example:
dpo_mlx.sh. - Preference pair
- Two answers to the same prompt, one chosen and one rejected. Example: clean JSON against the same JSON wrapped in chat.
- Reference model
- A frozen copy of the starting model that training measures drift against. Example:
--reference-model-pathindpo_mlx.sh. - Loss
- The single number training tries to shrink; for DPO it falls as the chosen and rejected answers pull apart. Example: 0.69 at the start, 0.63 after the video’s step.
From the lesson
lessons/18-preference-dpo/README.md
What and why
Section titled “What and why”SFT teaches what to answer. DPO teaches which of two answers is better using pairs of chosen and rejected answers, with no reward model and no reinforcement loop. Our triage model sometimes wraps its JSON in chat (“Sure! Here’s the JSON…”), writes prose, or gets the category or severity wrong. Each of those is a rejected answer, and the clean, correct JSON is chosen.
prompt "Card declined twice, launch is tomorrow."chosen {"category": "billing", "severity": 4, "next_action": "route to billing; verify charge"}rejected Sure! Here's the JSON you asked for: {…} Let me know if you need anything else!In a real product, these pairs come for free: an agent’s edit (chosen) of a model draft (rejected), a thumbs-down, a regenerate.
The files, in order
Section titled “The files, in order”| step | file | runs on | does |
|---|---|---|---|
| 1 | make_pairs.py |
anywhere | 800 train and 100 valid pairs, with four realistic failure kinds; written in mlx-lm-lora and Fireworks formats |
| 2 | dpo_mlx.sh |
Mac | fuses Day 11’s LoRA into an SFT model if present (SFT first, then DPO), then mlx_lm_lora.train --train-mode dpo --beta 0.1 |
| 3 | dpo_fireworks.sh |
Mac → FW | firectl dpo-job create --loss-method DPO; optional deployment and eval, with auto-delete |
| 4 | pref_eval.py |
any target | valid JSON %, chatty %, category and severity accuracy, answer length; with --pairwise, DPO’s implicit reward (chosen vs rejected log-prob) |
cd lessons/18-preference-dpopython make_pairs.pypython pref_eval.py --target mock --label mock # rehearse the metrics offline
bash dpo_mlx.sh # 10–20 min# two terminals, as the script prints:# mlx_lm.server --model <START> --port 8081# mlx_lm.server --model <START> --adapter-path adapters/dpo --port 8082MLX_MODEL=<START> python pref_eval.py --target mlx --label beforeMLX_URL=http://localhost:8082/v1 MLX_MODEL=<START> python pref_eval.py --target mlx --label after-dpo
python pref_eval.py --pairwise <START> --label pairs-beforepython pref_eval.py --pairwise <START> --adapter adapters/dpo --label pairs-after
bash dpo_fireworks.sh # optional, paidWhat each command does
python make_pairs.pyTurns each support ticket into a pair: the correct JSON as chosen, one realistic mistake as rejected, in a Mac format and a Fireworks format. Look for
train: 800 pairs,valid: 100 pairs, the four mistake kinds at about 200 each (191 to 213), and one example pair. Ifdpo_mlx.shis already running from before the video, it built these: read its output and skip this command.python pref_eval.py --target mock --label mockFirst start the practice server: from
03-labs, runmake mock &. Then this rehearses against it: it sends the 100 test tickets and prints the five measures you compare later. The numbers mean nothing; it proves the scoring works.bash dpo_mlx.shTrains DPO on your Mac for 10 to 20 minutes, as a small add-on saved in
adapters/dpo. It first printsDPO starting from:and a model:sft_modelif Day 11’s LoRA was merged in, else the plain instruct model. That model is<START>below. At the end it prints two server commands: run each in its own terminal and leave them running. If you started it before the video, let that run finish.MLX_MODEL=<START> python pref_eval.py --target mlx --label beforeScores the model before DPO on the 100 test tickets, through the server on port 8081 (no add-on). Replace
<START>with the modeldpo_mlx.shprinted. Look forvalid_%near 80 andchatty_%near 20, roughly as the lesson expects; yours will differ.MLX_URL=http://localhost:8082/v1 MLX_MODEL=<START> python pref_eval.py --target mlx --label after-dpoThe same 100 tickets through the server on port 8082, which adds the DPO add-on. Look for
valid_%near 98,chatty_%near 1, andcategory_%the same as before or higher.python pref_eval.py --pairwise <START> --label pairs-beforeNo server needed: it loads the model itself. For each of the 100 checking pairs it compares how likely the model is to write the chosen answer against the rejected one. Look for
pref_acc_%near 60 and amarginnear +1. The row prints aspairwise base: the script names pairwise rows itself.python pref_eval.py --pairwise <START> --adapter adapters/dpo --label pairs-afterThe same with the DPO add-on loaded. Look for
pref_acc_%near 95 and amarginnear +10: the model now clearly favours the chosen answers. The row prints aspairwise +dpo. 3 of the 100 checking pairs are identical (the kit’s wrong-severity mistake cannot push a severity-5 ticket higher, or a severity-1 ticket lower), so 97% is the ceiling.bash dpo_fireworks.shOptional and paid. Runs the same DPO job on Fireworks, on Qwen2.5 7B Instruct by default. Training costs cents: the lab book lists $1.00 per million training tokens for LoRA DPO up to 16B. It asks before deploying, because a dedicated deployment bills by the hour ($8 an hour for an H100 in the lab book); the script deletes it on exit. Afterwards run
make fw-check: it should list nothing.
<START> is whichever model dpo_mlx.sh printed: sft_model/ if you trained Day 11’s LoRA first,
otherwise the instruct model.
What you should see (shape)
Section titled “What you should see (shape)”How to read it
Only the pattern matters; your numbers will differ. The first table scores 100 test tickets before and after DPO: valid_% JSON that fits the ticket format, chatty_% chat or prose, category_% and severity_% right answers, avg_chars average length (≈ means about the same, the up arrow higher). In the second, pref_acc_% is how often the model favours the chosen answer and margin by how much; format should jump while category holds.
label valid_% chatty_% category_% severity_% avg_charsbefore ~80 ~20 ~70 ~60 ~110after-dpo ~98 ~1 ≈ or ↑ ≈ or ↑ ~85 ← format fixed, no correctness lost
label pref_acc_% marginpairwise base ~60 +1.xpairwise +dpo ~95 +10.x ← the implicit reward grewIf category_% drops after DPO, taste has cost correctness (the Direct Preference Optimization video at the top of this step). Try a larger β,
fewer iterations, or more wrong_cat pairs so correctness is part of what “preferred” means.
Check yourself
Section titled “Check yourself”4 questions. Say your answer out loud, then tap to check it.
Why merge Day 11’s small add-on (the SFT adapter, trained on examples) into the model first, instead of running DPO on the plain instruct model?Show answerHide
In plain words
DPO polishes a model that already does the task. The plain model gets only about 40% of ticket categories right, so the approved answers are still unlikely for it, and DPO mostly learns to avoid the rejected text.
Picture it
Coaching someone’s handwriting helps once they can form the letters. Show a beginner pairs of neat and messy pages and they mostly learn what to avoid, not how to write.
With real numberslesson 17’s README (Day 11), dpo_mlx.sh and the RLHF and DPO videos
- Day 11’s expected table: before its training, the model gets about 40% of ticket categories right; after training, all three methods score well above that.
dpo_mlx.shmerges Day 11’s LoRA intosft_model, then runs DPO from there and uses it as the frozen reference too.- Without it, DPO starts from the plain instruct model, the one at about 40%.
- The RLHF video’s rule: “You cannot reinforce behaviour the model never produces.” The DPO video’s: SFT first, then DPO.
Words to know
- SFT (supervised fine-tuning)
- Training on examples of the answers you want. Example: Day 11’s LoRA on 800 support tickets.
- Adapter (LoRA)
- A small add-on of extra learned numbers trained on top of a frozen model. Example: Day 11’s, in
adapters/lora. - Fuse
- Bake an adapter’s numbers into the model so it becomes one self-contained model. Example:
sft_model. - Instruct model
- A model already trained to follow instructions and chat. Example: Qwen2.5-0.5B-Instruct.
Go deeper: the engineer version
The kit's question
Why fuse the SFT adapter first rather than running DPO on the raw instruct model?
The kit's answer
DPO refines a model that already does the task. With nothing to refine, it mostly learns “not the rejected text”.
More detail: DPO only reshapes probability between responses the model can already produce. Starting from a model that already emits valid triage JSON, the pairs refine format and correctness at the margin. From the raw instruct model the chosen answers are unlikely, so much of the loss is reduced by pushing the rejected text down. The same model is the frozen reference (--reference-model-path "$START"), so the reference also knows the task.
What does β (beta, the leash setting, 0.1 in dpo_mlx.sh) actually do?Show answerHide
In plain words
It sets how far the model may drift from its frozen starting copy before training stops pushing it further. A smaller β lets it move further.
Picture it
Think of β as how short the leash is. On a short leash (large β), the dog gets a few steps away and the leash goes taut. On a long leash (small β), it wanders much further before it is pulled back.
With real numbersthe DPO video’s worked step, recomputed at other β values
- Margin: the chosen answer +0.8, the rejected -0.4, both against the frozen copy: 0.8 - (-0.4) = 1.2.
- β 0.1: 0.1 x 1.2 = 0.12. The sigmoid σ turns any number into a value between 0 and 1: σ(0.12) = 0.53. Loss = -log(0.53) = 0.63, down from 0.69 when the margin was 0.
- β 0.05: the same margin earns only 0.06, loss 0.66, so training keeps pulling the two apart for longer.
- β 0.2: 0.24, loss 0.58, so less drift is needed before the pull fades.
Words to know
- β (beta)
- The leash: how far the model may move from the reference. Example: 0.1 in
dpo_mlx.sh. - Margin
- How much more the model favours the chosen answer than the rejected one. Example: 1.2 in the DPO video.
- Log-probability
- The natural log of how likely the model is to write an answer; DPO works in these. Example: +0.8 against the reference.
- Loss
- The single number training tries to shrink. Example: 0.69 at the start, 0.63 after one step.
Go deeper: the engineer version
The kit's question
What does β do here, concretely?
The kit's answer
It scales the margin in the loss. A smaller β lets the policy move further from the reference before the loss stops rewarding it.
More detail: L = −log σ(β · [(log π(y_c) − log π_ref(y_c)) − (log π(y_r) − log π_ref(y_r))]). Its gradient is scaled by σ(−β · margin), which shrinks as β · margin grows, so a larger β saturates at a smaller margin (a tighter leash). β is the same KL coefficient as in RLHF: DPO is derived from that KL-regularised objective, with the leash built into the loss. Set it with BETA=0.2 bash dpo_mlx.sh.
Suppose half your rejected answers are chatty (the right JSON wrapped in chat), and none have the wrong category (wrong_cat). What happens to category accuracy after DPO?Show answerHide
In plain words
The model learns to fix its format, not to get categories right. Category accuracy may stay flat or even slip, because nothing in the pairs rewards the right category.
Picture it
If a supervisor only ever corrects greetings and sign-offs, the new agent writes polite emails that may still send customers to the wrong team.
With real numberslesson 18’s make_pairs.py and README
- The lesson’s 800 pairs: each rejected answer is one of four mistakes, about 200 each (800 ÷ 4).
- Only the wrong-category kind differs from the chosen answer in its category. Chatty, prose and wrong-severity answers all keep the right one.
- The question’s mix: 400 chatty (half of 800) and 0 wrong-category, so no pair teaches the category.
- Watch
category_%: the lesson expects it to hold at about 70% or rise. A drop means taste has cost correctness.
Words to know
- Chatty answer
- The right JSON wrapped in friendly chat, so a program cannot read it. Example: “Sure! Here’s the JSON you asked for: ...” (
chatty_%also counts prose answers.) - wrong_cat pair
- A pair whose rejected answer is valid JSON with the wrong category. Example: 213 of the 800.
- Category accuracy
- The share of test tickets given the right category. Example:
category_%, about 70% before. - Preference pair
- Two answers to the same prompt, one chosen and one rejected. Example: 800 in this lesson.
Go deeper: the engineer version
The kit's question
Your pairs are 50% “chatty”. What happens to accuracy if none of them are wrong_cat?
The kit's answer
The model learns format, not correctness, and category accuracy may even slip.
More detail: DPO raises the log-probability gap along whatever separates chosen from rejected. If no pair separates on category, the gradient carries no category signal, and drift from the reference (bounded by β) can move it either way. The README’s fixes: a larger β, fewer iterations, or more wrong_cat pairs, so correctness is part of what “preferred” means. make_pairs.py’s own rule: pairs should differ in the thing you care about.
When would you use ORPO (odds-ratio preference optimization, a relative of DPO) instead of DPO? Fireworks runs it with --loss-method ORPO.Show answerHide
In plain words
When you want to teach the task and the preference in one training run, without keeping a second, frozen copy of the model in memory.
Picture it
DPO is coaching after a course: first the course, then a coach who compares you with your old self. ORPO is one class that teaches the skill and the taste together, with no old self to compare against.
With real numbersthe DPO and RLHF videos, dpo_mlx.sh and dpo_fireworks.sh
- RLHF with PPO: 4 models in memory. DPO: 2, the model being trained and a frozen reference.
- ORPO: no reference copy, so only the model being trained, and no separate example-training step first.
- On your Mac, each copy of the 0.5-billion-number model is 0.5 billion x 2 bytes = about 1 GB.
- For the Qwen2.5 7B model
dpo_fireworks.shuses (7.6 billion numbers), a frozen copy at 2 bytes each is 7.6 x 2 = about 15 GB.
Words to know
- ORPO (odds-ratio preference optimization)
- A preference method that does example training and preference training in one run, with no reference model. Example:
--loss-method ORPO. - Reference model
- A frozen copy of the starting model that training measures drift against. Example: one of DPO’s 2 models.
- SFT (supervised fine-tuning)
- Training on examples of the answers you want. Example: Day 11’s LoRA.
- bf16
- A 16-bit number format, 2 bytes per number, used for training. Example: the 0.5B model is about 1 GB in bf16.
Go deeper: the engineer version
The kit's question
When would you use ORPO instead (--loss-method ORPO on Fireworks)?
The kit's answer
When you want SFT and preference tuning in one step without holding a reference model in memory.
More detail: ORPO adds an odds-ratio term, −λ · log σ(log(odds(chosen) / odds(rejected))) with odds(y) = P(y) ÷ (1 − P(y)), to the ordinary SFT loss on the chosen answer, so it needs neither a reference model nor a separate SFT stage. With no reference there is no explicit leash to a known-good model, so evaluate accuracy as carefully as with DPO. On Fireworks: firectl dpo-job create --loss-method ORPO with --orpo-lambda (the comment at the top of dpo_fireworks.sh).
Explain what you learned
Section titled “Explain what you learned”Question Our fine-tuned support model still wraps its JSON in chat and breaks our parser. We have no budget for labellers. What would you do?
One clear answer
Their agents already edit the model’s drafts. Every edit is a preference pair. Run DPO on an SFT’d model and the formatting failures disappear. I’d watch category accuracy so taste doesn’t cost correctness, and where correctness can be checked by code, I’d use RFT with a grader instead.
What this means
- “Their agents already edit the model’s drafts”: Support staff already fix the model’s replies as part of their day.
- “Every edit is a preference pair”: Each fix gives two answers to the same ticket: the edited one (chosen) and the original draft (rejected). That is free training data.
- “Run DPO on an SFT’d model”: Train on those pairs with DPO, starting from the model already fine-tuned on examples (SFT). Here that is Day 11’s LoRA, merged in.
- “the formatting failures disappear”: As the lesson expects: valid JSON from about 80% to about 98%, chatty replies from about 20% to about 1%.
- “I’d watch category accuracy so taste doesn’t cost correctness”: I compare the right-category rate before and after, because pairs that differ only in style can quietly make answers less accurate.
- “where correctness can be checked by code, I’d use RFT with a grader instead”: If a program can mark an answer right or wrong, I train with RFT (reinforcement fine-tuning), which scores each answer with that program instead of preferences.
Your numbersSaved on this device and collected in the Day 12 wrap-up.
Hint: valid_% in the before and after-dpo rows of pref_eval.py. Write both, for example “80 / 98”.
Hint: category_%, same two rows. Write both as before / after, for example “70 / 70”.
Hint: margin in the pairwise base and pairwise +dpo rows. Write both as before / after; the lesson shows about +1 and about +10.
Done when
Section titled “Done when”pref_eval.py has printed your before and after-dpo rows and both pairwise rows. Valid JSON went up, chatty went down, category_% did not drop, and the margin grew.
Stuck?
Section titled “Stuck?”pref_eval.py --target mocksaysConnection refused.- The practice server is not running. From the
03-labsfolder runmake mock &, then try again. pref_eval.py --target mlxsaysConnection refused.- The two model servers are not running.
dpo_mlx.shprinted their commands at the end. In each new terminal, first go to03-labsand runsource .venv/bin/activate, thencd lessons/18-preference-dpo. Startmlx_lm.server --model <START> --port 8081in one terminal and the same with--adapter-path adapters/dpo --port 8082in another. When both are up, run the check again. - You do not know what to type for
<START>. - Use the model
dpo_mlx.shprinted on itsDPO starting from:line, such as the full path tosft_model. It is also saved in the file.statein the lesson folder. dpo_mlx.shstarts from the instruct model, notsft_model.- It did not find Day 11’s LoRA in
lessons/17-full-vs-peft-mlx/adapters/lora. DPO still runs, but on a model that gets only about 40% of categories right (check yourself 1). From03-labs, runVARIANTS=lora bash lessons/17-full-vs-peft-mlx/run_variants.sh(5 to 15 minutes), thenbash lessons/18-preference-dpo/dpo_mlx.shagain, and restart both servers with the new<START>it prints (sft_model). mlx_lm_lora.trainstops with an unknown option or argument error.- The tool’s options change between versions, and the script says its
--helpwins. Runmlx_lm_lora.train --help, match the options at the top ofdpo_mlx.shto it, and run it again. category_%dropped after DPO.- Taste has cost correctness. The lesson’s fixes: a larger β (
BETA=0.2 bash dpo_mlx.sh), fewer steps (ITERS=100 bash dpo_mlx.sh), or more wrong-category pairs, so correctness is part of what “preferred” means. - The pairwise rows are not in
results/18-pref.csv. - They have different columns, so the kit’s recorder saves each one in a new file named
results/18-pref-<time>.csv. Copy the numbers from the terminal or from those files. dpo_fireworks.shsays the base model is not found or not DPO-enabled.- Model ids change. Pick one the Fireworks model library marks as DPO-enabled and run
BASE_MODEL=accounts/fireworks/models/<id> bash dpo_fireworks.sh. Afterwards runmake fw-check: it should list nothing.
Go deeper: the lab book's training lab for this step All fixes
Code in this step
Section titled “Code in this step”dpo_fireworks.sh Bash · 39 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 18 · step 3 — the same DPO job on Fireworks. ⏱ PAID (training + a deployment)## firectl dpo-job create --loss-method DPO (or ORPO with --orpo-lambda; see video 18)## The base model must be one the Fireworks model library marks as DPO-enabled. Override# with BASE_MODEL=… Training is billed per token (cents for 800 short pairs). Evaluating# needs a dedicated deployment (⏱ per hour), which a trap deletes on exit, even on Ctrl-C.# Flags follow the Fireworks docs at time of writing; `firectl dpo-job create --help` wins.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"BASE_MODEL=${BASE_MODEL:-accounts/fireworks/models/qwen2p5-7b-instruct}ACCOUNT=${FIREWORKS_ACCOUNT_ID:-$(firectl whoami 2>/dev/null | grep -oE 'accounts/[a-z0-9-]+' | head -1 | cut -d/ -f2)}STAMP=$(date +%m%d%H%M)DATASET="triage-dpo-$STAMP"; OUT="triage-dpo-$STAMP"
[[ -f data/fireworks/train.jsonl ]] || python make_pairs.pyfirectl dataset create "$DATASET" data/fireworks/train.jsonlout=$(firectl dpo-job create --loss-method DPO --base-model "$BASE_MODEL" \ --dataset "accounts/$ACCOUNT/datasets/$DATASET" --output-model "$OUT")echo "$out"JOB=$(echo "$out" | grep -oE 'dpoJobs/[A-Za-z0-9-]+' | head -1 | cut -d/ -f2)echo "▸ job $JOB — polling every 60 s (Ctrl-C is safe; the job keeps running)"until firectl dpo-job get "$JOB" | grep -qE "COMPLETED"; do firectl dpo-job get "$JOB" | grep -qE "FAILED|CANCELLED" && { firectl dpo-job get "$JOB"; exit 1; } sleep 60; echo -n "."doneecho " trained: accounts/$ACCOUNT/models/$OUT"
read -r -p "Deploy for evaluation now? This starts per-hour billing. [y/N] " yn[[ "$yn" =~ ^[Yy]$ ]] || { echo "Skipped. Deploy later with lesson 10's 3_deploy.sh pattern."; exit 0; }dep=$(firectl deployment create "accounts/$ACCOUNT/models/$OUT" --deployment-shape default)DEP=$(echo "$dep" | grep -oE 'deployments/[A-Za-z0-9-]+' | head -1 | cut -d/ -f2)trap 'echo "▸ deleting $DEP"; firectl deployment delete "$DEP"; firectl deployment list' EXITuntil firectl deployment get "$DEP" | grep -qiE "state.*READY"; do sleep 20; echo -n "."; done; echo " ready"python pref_eval.py --target fireworks --model "$BASE_MODEL" --label fw-basepython pref_eval.py --target fireworks --model "accounts/$ACCOUNT/models/$OUT#accounts/$ACCOUNT/deployments/$DEP" --label fw-dpodpo_mlx.sh Bash · 46 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 18 · step 2 — DPO on your Mac with mlx-lm-lora (video 18).## Rule from the video: SFT first, then DPO. If lesson 17's LoRA adapter exists we fuse it# into an SFT model and run DPO on top of that; otherwise we start from the instruct model.# The same model is used as the frozen REFERENCE, which is what β measures distance from.## --train-mode dpo the DPO loss (also: orpo, cpo, grpo, ppo … see --help)# --beta 0.1 the leash (video 18: start around 0.1)# --dpo-cpo-loss-type sigmoid the original DPO loss: −log σ(β · margin)# --reference-model-path the frozen copy the policy is compared against## Flags follow the mlx-lm-lora README at time of writing; `mlx_lm_lora.train --help` wins.# ~10–20 min on M-series for 0.5B. Two models in memory (policy + reference).# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"BASE=${MODEL:-mlx-community/Qwen2.5-0.5B-Instruct-bf16}L17=../17-full-vs-peft-mlx/adapters/lora
python -c "import mlx_lm_lora" 2>/dev/null || pip install -U mlx-lm-lora[[ -f data/mlx/train.jsonl ]] || python make_pairs.py
if [[ -d "$L17" && ! -d sft_model ]]; then echo "▸ fusing lesson 17's LoRA into an SFT model (SFT first, then DPO)" mlx_lm.fuse --model "$BASE" --adapter-path "$L17" --save-path sft_modelfiSTART=$([[ -d sft_model ]] && echo "$(pwd)/sft_model" || echo "$BASE")echo "▸ DPO starting from: $START"
mlx_lm_lora.train --model "$START" --reference-model-path "$START" \ --train --train-mode dpo --data data/mlx \ --beta ${BETA:-0.1} --dpo-cpo-loss-type sigmoid \ --iters ${ITERS:-200} --batch-size 2 --learning-rate 5e-6 \ --adapter-path adapters/dpo 2>&1 | tee dpo_train.log
echo "START_MODEL=\"$START\"" > .statecat <<EOT
▸ serve before and after (two terminals), then compare: mlx_lm.server --model "$START" --port 8081 mlx_lm.server --model "$START" --adapter-path adapters/dpo --port 8082 MLX_MODEL="$START" python pref_eval.py --label before --target mlx MLX_URL=http://localhost:8082/v1 MLX_MODEL="$START" python pref_eval.py --label after-dpo --target mlxEOTmake_pairs.py Python · 86 lines
"""Lesson 18 · step 1 — build preference pairs from the ticket data (video 18).
Each training ticket becomes one pair with the SAME prompt: chosen the correct, concise JSON answer (what a reviewer approves) rejected one realistic failure (what a reviewer would edit or thumbs-down): chatty "Sure! Here's the JSON: …" → breaks the parser prose a helpful paragraph, no JSON → breaks the parser wrong_cat valid JSON, wrong category → wrong answer wrong_sev valid JSON, severity off by 2 → wrong judgement
Pairs should differ in the thing you care about. Here that is "parseable AND right".In a real product, these come from agent edits, thumbs-down and regenerate clicks.
Writes two formats of the same pairs: data/mlx/{train,valid}.jsonl {"system", "prompt", "chosen", "rejected"} (mlx-lm-lora) data/fireworks/train.jsonl {"input": {"messages": [...]}, "preferred_output": [...], "non_preferred_output": [...]} (firectl dpo-job)
python make_pairs.py"""import jsonimport randomfrom pathlib import Path
HERE = Path(__file__).parentSRC = HERE.parents[1] / "data" / "tickets"CATS = ["billing", "outage", "how-to", "abuse"]rng = random.Random(18)
def rejected_for(ticket: str, ans: dict) -> tuple[str, str]: kind = rng.choice(["chatty", "prose", "wrong_cat", "wrong_sev"]) if kind == "chatty": return kind, "Sure! Here's the JSON you asked for:\n" + json.dumps(ans) + "\nLet me know if you need anything else!" if kind == "prose": art = "an" if ans["category"][0] in "aeiou" else "a" return kind, (f"This looks like {art} {ans['category']} issue. I'd rate it around {ans['severity']} out of 5, " f"and the next step would be to {ans['next_action']}.") if kind == "wrong_cat": wrong = dict(ans, category=rng.choice([c for c in CATS if c != ans["category"]])) return kind, json.dumps(wrong) wrong = dict(ans, severity=max(1, min(5, ans["severity"] + rng.choice([-2, 2])))) return kind, json.dumps(wrong)
def build(split: str) -> list[dict]: pairs = [] for line in (SRC / f"{split}.jsonl").read_text().splitlines(): msgs = json.loads(line)["messages"] system, user, answer = msgs[0]["content"], msgs[1]["content"], msgs[2]["content"] kind, bad = rejected_for(user, json.loads(answer)) pairs.append({"system": system, "prompt": user, "chosen": answer, "rejected": bad, "kind": kind}) return pairs
def main() -> None: if not (SRC / "train.jsonl").exists(): raise SystemExit("run: python data/make_tickets.py") (HERE / "data" / "mlx").mkdir(parents=True, exist_ok=True) (HERE / "data" / "fireworks").mkdir(parents=True, exist_ok=True) kinds = {} for split in ("train", "valid"): pairs = build(split) with (HERE / "data" / "mlx" / f"{split}.jsonl").open("w") as f: for p in pairs: f.write(json.dumps({k: p[k] for k in ("system", "prompt", "chosen", "rejected")}) + "\n") kinds[p["kind"]] = kinds.get(p["kind"], 0) + (split == "train") if split == "train": with (HERE / "data" / "fireworks" / "train.jsonl").open("w") as f: for p in pairs: f.write(json.dumps({ "input": {"messages": [{"role": "system", "content": p["system"]}, {"role": "user", "content": p["prompt"]}]}, "preferred_output": [{"role": "assistant", "content": p["chosen"]}], "non_preferred_output": [{"role": "assistant", "content": p["rejected"]}], }) + "\n") print(f"{split}: {len(pairs)} pairs") print("rejected kinds (train): " + ", ".join(f"{k} {v}" for k, v in sorted(kinds.items()))) print(f"→ {HERE / 'data'}") ex = build("valid")[0] print(f"\nexample\n prompt {ex['prompt']}\n chosen {ex['chosen']}\n rejected {ex['rejected']} ({ex['kind']})")
if __name__ == "__main__": main()pref_eval.py Python · 93 lines
"""Lesson 18 · step 4 — did preference tuning move what we wanted, and break anything?
Behaviour on the 100 held-out tickets (any --target: mock, mlx, fireworks …): valid_% strict JSON that matches the schema ← DPO should push this up chatty_% answers wrapped in chat ("Sure! …") or prose ← and this down category_% right category ← must not drop (taste can cost correctness) severity_% right severity avg_chars answer length ← watch for length drift
Pairwise, on the validation pairs (Apple silicon, --pairwise MODEL [--adapter PATH]): pref_acc_% how often the model finds CHOSEN more likely than REJECTED (log-prob sum) margin mean log-prob gap: DPO's implicit reward. It should grow after training.
python pref_eval.py --target mock --label mock python pref_eval.py --target mlx --label before python pref_eval.py --pairwise mlx-community/Qwen2.5-0.5B-Instruct-bf16 --adapter adapters/dpo"""import argparseimport jsonfrom pathlib import Path
from felab import add_target_args, banner, client, record, resolve, tablefrom felab.tickets import SYSTEM, grade, load_test, parse
HERE = Path(__file__).parent
def behaviour(a) -> dict: t = resolve(a) banner(t) cli = client(t) rows = load_test() g, lengths, chatty = [], [], 0 for r in rows: resp = cli.chat.completions.create(model=t.model, temperature=0, max_tokens=120, messages=[{"role": "system", "content": SYSTEM}, {"role": "user", "content": r["ticket"]}]) out = (resp.choices[0].message.content or "").strip() g.append(grade(out, r)) lengths.append(len(out)) chatty += parse(out) is None # not parseable as a bare JSON object → chat or prose n = len(rows) return {"label": a.label, "target": t.name, "valid_%": 100 * sum(x["valid"] for x in g) / n, "chatty_%": 100 * chatty / n, "category_%": 100 * sum(x["category_ok"] for x in g) / n, "severity_%": 100 * sum(x["severity_ok"] for x in g) / n, "avg_chars": sum(lengths) / n}
def pairwise(model_path: str, adapter: str | None, limit: int) -> dict: """Sum of token log-probs of each response given the prompt: DPO's own currency.""" import mlx.core as mx from mlx_lm import load model, tok = load(model_path, adapter_path=adapter)
def logprob(system: str, prompt: str, response: str) -> float: p_txt = tok.apply_chat_template([{"role": "system", "content": system}, {"role": "user", "content": prompt}], add_generation_prompt=True, tokenize=False) p_ids = tok.encode(p_txt) r_ids = tok.encode(response, add_special_tokens=False) ids = p_ids + r_ids logits = model(mx.array([ids[:-1]]))[0] logp = logits - mx.logsumexp(logits, axis=-1, keepdims=True) tgt = mx.array(ids[1:]) tok_lp = mx.take_along_axis(logp, tgt[:, None], axis=-1)[:, 0] return tok_lp[len(p_ids) - 1:].sum().item() # only the response tokens
pairs = [json.loads(line) for line in (HERE / "data/mlx/valid.jsonl").read_text().splitlines()][:limit] margins = [logprob(p["system"], p["prompt"], p["chosen"]) - logprob(p["system"], p["prompt"], p["rejected"]) for p in pairs] return {"label": f"pairwise {'+' + Path(adapter).name if adapter else 'base'}", "pref_acc_%": 100 * sum(m > 0 for m in margins) / len(margins), "margin": sum(margins) / len(margins), "pairs": len(margins)}
def main() -> None: ap = add_target_args(argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)) ap.add_argument("--label", default="run") ap.add_argument("--pairwise", metavar="MODEL", help="local MLX model path/repo for log-prob scoring") ap.add_argument("--adapter", help="adapter path for --pairwise (e.g. adapters/dpo)") ap.add_argument("--limit", type=int, default=100) a = ap.parse_args() row = pairwise(a.pairwise, a.adapter, a.limit) if a.pairwise else behaviour(a) print("\n" + table([row])) record("18-pref", row) if not a.pairwise: print("\nRun it before and after DPO with different --label values; results/18-pref.csv keeps both.")
if __name__ == "__main__": main()