Skip to content

Full vs LoRA vs DoRA

course 20 of 22

lesson 17 · lessons/17-full-vs-peft-mlx

1 h 05 Free Mac

What you will do and why

You train one real model three ways on the same 800 support tickets: change every number, or train a small add-on (LoRA, and its variant DoRA). One table then shows what each cost, what it gained and what it broke.

Why it matters: Lesson example: the fully trained model saves a file of about 990 MB; the LoRA add-on saves about 9 MB, 110 times smaller, with accuracy close to full.

You are done when: compare_variants.py has printed one table with rows for the untrained model and for full, LoRA and DoRA, showing peak memory, file size, task accuracy and the general quiz.

Start this first

Run it from the 03-labs folder with the Python environment on (source .venv/bin/activate). It checks the data, downloads the model the first time (about 1 GB) and trains it three times: 15 to 45 minutes in all. Leave it running while you watch. For the other commands, open a second terminal window, go to 03-labs, run source .venv/bin/activate, then cd lessons/17-full-vs-peft-mlx.

Terminal window
cd lessons/17-full-vs-peft-mlx && bash run_variants.sh

Supervised Fine-Tuning · 4:02

Download mp4 (15.2 MB)

Chapters

In this video What training-by-example data looks like, why only the answer is graded, how much data you need, and how to read the training curves.

3 key points

  1. Supervised fine-tuning (SFT) teaches by example: a prompt in, the ideal answer out, and only the answer is graded.

    Each ticket example holds a system message (the standing instructions), the customer’s ticket and the ideal JSON answer. The trainer skips the first two when scoring (masking the prompt); run_variants.sh turns this on with --mask-prompt.

  2. A few hundred clean examples beat many noisy ones.

    Typical ranges for ticket triage with LoRA on a 7B model: 50 examples mostly fix the format, 500 add 8 to 12 points of accuracy, 2,000 add 12 to 16, then it flattens. 20,000 noisy ones often do worse than 2,000 clean. Today’s data: 800 training tickets.

  3. Read the two loss curves together: training loss (how wrong it is on the examples it learns from) and validation loss (how wrong on examples held back).

    Both falling: it is learning. Training falling while validation rises: it is memorising, so stop earlier or add data. Both flat and high: check the data, or raise the learning rate (the size of each nudge). The video’s usual starting values: about 0.0001 (written 1e-4) for LoRA, 10 times smaller for full.

Today you train one small model three ways on the same 800 support tickets. Full fine-tuning changes every number in it. LoRA freezes it and trains a small add-on; DoRA is a LoRA variant. One table then shows what each cost (memory, time, file size), what it gained (accuracy) and what it broke (general knowledge).

Picture it

Full fine-tuning reprints the whole textbook with your changes, and while it works it needs about 8 times the book’s room for drafts and notes. LoRA sticks a pad of notes on the pages that matter: cheap, easy to swap, and the original text stays readable underneath. With LoRA almost all the weight you carry is the book itself.

With real numbersQwen2.5-0.5B-Instruct, lesson 17’s calculator and expected table

  • The model: 490 million numbers, stored at 2 bytes each (bf16, a 16-bit format): 0.98 GB, about 1 GB.
  • Full fine-tuning: 490 million x 16 bytes = 7.8 GB, 8 times the model. The working numbers for each batch of examples (activations) come on top: the lesson expects a peak of about 8 to 10 GB.
  • LoRA at rank 16 (the add-on’s size setting, in the calculator’s example): 2,162,688 trained numbers, 0.44% of the model. Memory: the 0.98 GB frozen model plus 2.16 million x 16 bytes (0.03 GB), about 1.0 GB. With activations, about 2 to 3 GB at peak.
  • Saved file, from the lesson’s expected table: about 990 MB for full (a whole new model) against about 9 MB for LoRA, 110 times smaller. Your own run’s sizes are the real figures.
  • The same sums for an 8B model: 128 GB to train fully, 16.2 GB with LoRA, 4.7 GB with QLoRA (LoRA on a 4-bit copy of the model).

Words to know

Full fine-tuning
Training that changes every number in the model. Example: 7.8 GB of training memory for the 0.5B model.
LoRA and DoRA
Freeze the model and train a small add-on beside it. DoRA also trains one extra number per row of each grid, setting how strong that row is; it often scores 1 to 2 accuracy points higher and trains a bit slower. Example: under 1% of the numbers trained.
Checkpoint
The file a training run saves: the whole changed model for full fine-tuning, only the add-on for LoRA. Example: about 990 MB against about 9 MB in the lesson’s expected table.
Forgetting
Getting better at the trained task and worse at everything else. Example: a drop in the 12-question general quiz.

From the lesson

lessons/17-full-vs-peft-mlx/README.md

Today’s three videos (Full Fine-Tuning, Supervised Fine-Tuning and Parameter-Efficient Fine-Tuning) make three claims. This lesson tests all three on a real model:

  1. Full fine-tuning costs about 16 bytes per parameter to train, against about 2 to serve.
  2. LoRA and DoRA train under 1% of the weights and land close to full fine-tuning on a narrow task.
  3. Full fine-tuning risks more forgetting, meaning getting better at your task and worse at everything else.

You train Qwen2.5-0.5B-Instruct three ways on the same ticket data. Then you read cost (memory, minutes, checkpoint size), gain (task accuracy) and damage (a general quiz) from one table. The model is small enough that full fine-tuning fits on a 16 GB Mac.

step file runs on does
0 sft_data_check.py anywhere checks the data: JSON, roles, ends-with-assistant, duplicates, test leakage, length, label balance
1 peft_params.py anywhere trainable parameters and training memory for full, LoRA, DoRA and QLoRA, from real model shapes
2 run_variants.sh Mac mlx_lm.lora --fine-tune-type full / lora / dora, same data and settings
3 compare_variants.py Mac one table: trainable %, peak memory, val loss, minutes, checkpoint MB, task accuracy, general quiz

This checker earned its place while the course was being built. On its first run it found that the synthetic dataset had 301 training tickets identical to test tickets, which would have inflated every fine-tune score. data/make_tickets.py now dedupes, keeps the splits disjoint and keeps one phrasing per category for the test set only.

Terminal window
cd lessons/17-full-vs-peft-mlx
python sft_data_check.py
python peft_params.py --model 0.5b # what you are about to train
python peft_params.py --model 8b --rank 16 # the same maths at a size customers use
python peft_params.py --model 70b --rank 64 --mlp
bash run_variants.sh # ~15–45 min for all three
python compare_variants.py

What each command does

  1. python sft_data_check.py

    Checks the 800 training tickets before any training. It looks for unreadable lines, unknown speakers (each message must be marked as the instructions, the customer, the model or a tool), examples that do not end with the answer, duplicates and over-long examples. It also catches any ticket that also sits in the 100-ticket test set (a leak). Look for rows 800 and no problems found.

  2. python peft_params.py --model 0.5b

    Works out, from the model’s real shape, how many numbers each method trains and how much memory training needs for today’s 0.5B model. Look for 7.8 GB for full against about 1.0 GB for LoRA, which trains 2,162,688 numbers (0.441%). Its LoRA file size says about 4 MB, not the lesson’s expected 9 MB: the calculator assumes rank 16 and 2 bytes per saved number, while run_variants.sh uses the trainer’s own defaults. Your own run’s files give the real size.

  3. python peft_params.py --model 8b --rank 16

    The same sums for Llama 3.1 8B, a size customers use. Look for 128.0 GB for full, 16.2 GB for LoRA and 4.7 GB for QLoRA (LoRA on a 4-bit copy of the model); the saved file is 16.0 GB for full against 27 MB for LoRA. The Full Fine-Tuning video’s rough figures are about 20 million numbers, 20 GB and 40 MB; the calculator counts rank 16 on the attention grids exactly.

  4. python peft_params.py --model 70b --rank 64 --mlp

    A 70B model at a higher rank, also adapting the MLP grids (the other big block in every layer). Look for 1129.6 GB for full, about 1.1 TB, against 154.5 GB for LoRA and 53.0 GB for QLoRA.

  5. bash run_variants.sh

    Checks the data again, then trains the model three ways, full, LoRA and DoRA, with the same tickets and settings: 300 steps of 4 examples (1,200 examples, 1.5 passes over the 800). Full uses a learning rate (the size of each nudge) 10 times smaller: 0.00001 against 0.0001, printed 1e-5 and 1e-4. Expect 5 to 15 minutes per method. If you started it before the video, let that run finish instead of starting a second one. Look for Val loss (how wrong it is on 100 held-back tickets) falling every 100 steps, and a Peak mem figure in each run.

  6. python compare_variants.py

    Reads the three training logs, then tests the untrained model and each trained one on two things: the category of 100 held-out tickets (30 in wordings never seen in training), and Day 4’s 12 quick general questions. Look for task accuracy jumping from about 40% to high for all three, LoRA and DoRA close to full, and any quiz drop in the full row.

Tight on memory? Run VARIANTS="lora dora" bash run_variants.sh and skip full. Short on time? Use ITERS=150.

What you should see (shape, not exact numbers)

Section titled “What you should see (shape, not exact numbers)”

How to read it

Each row is one method. trainable_% is how much of the model it trained; peak_mem_GB the most memory it used; minutes how long it took; val_loss how wrong it still was on held-back tickets. ckpt_MB adds up everything in the run’s folder, including the trainer’s in-between copies, so it can read several times the saved file. task_acc_% is the share of test tickets sorted correctly, and general_quiz the 12-question check for forgetting. The lesson’s table is the expected shape, not exact figures, and your trainer version may count a little differently. Yours lists the untrained model first, then dora, full, lora (alphabetical). Read across a row: what it cost, what it bought, what it broke.

variant trainable_% peak_mem_GB val_loss minutes ckpt_MB task_acc_% general_quiz
base (no training) 0.0 nan nan 0.0 0.0 ~40 ~9/12
full 100.0 ~8–10 lowest slowest ~990 high may drop
lora ~0.4 ~2–3 close fast ~9 high ≈ base
dora ~0.45 ~3 close a bit slower ~9 high ≈ base
  • Peak memory: full is several times LoRA’s. That is claim 1.
  • Checkpoint: about 1 GB for full against single-digit MB. That is why multi-LoRA serving works.
  • Task accuracy: all three jump well above base, and LoRA and DoRA land close to full. That is claim 2. Look at the held-out phrasings too, since they test generalisation rather than memory.
  • General quiz: if any variant drops, it is usually full. That is claim 3. On 12 questions it is a smoke alarm, not proof.

3 questions. Say your answer out loud, then tap to check it.

Day 11’s real training run starts from a model stored at 16 bits per number (bf16). On Day 8 you trained an add-on for a 7-billion-number model stored at 4 bits. Why must this comparison of full fine-tuning and LoRA use a 16-bit model, not a 4-bit one?Show answerHide

In plain words

Full fine-tuning changes every stored number, and 4-bit numbers are too coarse to take the tiny nudges training makes. Day 8’s QLoRA worked only because it never changed the 4-bit numbers: it trained a separate add-on instead.

Picture it

A 4-bit number is like a ruler marked only in whole centimetres. Training wants to move each mark by a hair’s width, and on that ruler the nudge rounds away to nothing. QLoRA leaves the coarse ruler alone and writes the fine corrections on a separate note.

With real numberslesson 17, and Day 8’s lesson 11

  • 4 bits can hold only 2 x 2 x 2 x 2 = 16 different values; 16 bits can hold 65,536.
  • Day 11’s model (Qwen2.5-0.5B-Instruct) at 16 bits: 490 million numbers x 2 bytes = 0.98 GB, small enough to train fully on a 16 GB Mac (about 8 to 10 GB at peak).
  • Day 8’s QLoRA: a 7B model frozen at 4 bits plus a small 16-bit add-on, about 6 to 10 GB to train. A full fine-tune of the same model needs about 120 GB (it has 7.6 billion numbers: 7.6 x 16 = 122 GB).

Words to know

bf16
A 16-bit number format used for training, 2 bytes per number. Example: Day 11’s model is 0.98 GB in bf16.
QLoRA
LoRA on a model stored in 4 bits: the model stays frozen, only the add-on trains. Example: about 6 to 10 GB for a 7B model on Day 8.
Frozen
Left unchanged during training. Example: the 4-bit model in QLoRA.
Gradient
The correction worked out for each trainable number at each training step. Example: full fine-tuning keeps one for every number.
Go deeper: the engineer version

The kit's question

Why must the base be bf16 here, and not the 4-bit model from lesson 11?

The kit's answer

Full fine-tuning updates the weights, and you can’t train 4-bit weights directly. QLoRA only works because the base stays frozen.

More detail: Quantized weights sit on a coarse grid, and rounding to that grid gives no useful gradient, so a training step cannot move them reliably. QLoRA keeps the 4-bit base frozen (only read, and turned back into 16-bit numbers for the math) and sends gradients only into the 16-bit adapter. run_variants.sh uses the bf16 base for all three runs so that full, LoRA and DoRA compare fairly.

Day 11’s memory calculator (peft_params.py --model 8b) says LoRA needs about 16 GB to train, but full fine-tuning needs 128 GB. Where do LoRA’s 16 GB go?Show answerHide

In plain words

Almost all of it is the model itself, frozen and stored at 2 bytes per number. The add-on and its training state are tiny. Store the frozen model in 4 bits instead (QLoRA) and it drops to about 5 GB.

Picture it

You hire a moving truck to deliver one small box: nearly all the cost is the truck, not the box. In LoRA the truck is the frozen model and the box is the add-on. QLoRA books a smaller truck: the same model squeezed into 4 bits.

With real numberspeft_params.py --model 8b --rank 16, lesson 17

  • Frozen model: 8 billion numbers x 2 bytes = 16 GB.
  • LoRA add-on at rank 16: 13.6 million numbers (0.17% of the model) x 16 bytes = 0.22 GB.
  • Total: 16 + 0.22 = 16.2 GB. Full: 8 billion x 16 bytes = 128 GB, about 8 times as much.
  • QLoRA: the frozen model at 4 bits is half a byte per number, plus a little for shared scale numbers: 0.5625 bytes. 8 billion x 0.5625 = 4.5 GB, plus 0.22 = 4.7 GB, about 5 GB.

Words to know

Frozen model
The original model, kept unchanged and only read during LoRA training. Example: 16 GB for an 8B model at 2 bytes per number.
Adapter (add-on)
The small set of extra numbers LoRA trains and saves. Example: 13.6 million numbers for an 8B model at rank 16.
Optimizer state
The running averages the optimizer (Adam) keeps per trained number to size each nudge. Example: 8 of the 16 bytes per number.
QLoRA
LoRA on a model stored in 4 bits. Example: 4.7 GB for an 8B model instead of 16.2 GB.
Go deeper: the engineer version

The kit's question

peft_params.py --model 8b says LoRA needs about 16 GB, but full needs 128. Where does LoRA’s 16 GB go?

The kit's answer

Almost all of it is the frozen bf16 base, 2 bytes × 8B. The adapter’s optimizer state is tiny. Switch to QLoRA and it drops to about 5 GB.

More detail: peft_params.py counts LoRA as 2 B x base parameters (frozen bf16 weights, no gradients or optimizer state) + 16 B x adapter parameters (bf16 weight and gradient, fp32 Adam m and v, fp32 master copy). QLoRA’s 0.5625 B per base weight is 4 bits plus the shared scales of each small block. Activations come on top of every figure and grow with batch x sequence length; --grad-checkpoint trims them.

In today’s runs, full fine-tuning makes nudges 10 times smaller than LoRA’s (a learning rate of 0.00001 against 0.0001). Why?Show answerHide

In plain words

Full fine-tuning moves every number in the model, including the ones that hold its general skills. Big steps across all of them would damage what the model already knows.

Picture it

Restoring a painting: if you may touch every inch, you use a fine brush and light strokes, or you ruin the parts that were already right. Painting on a clear sheet laid over it (LoRA), bolder strokes are safe, because the original underneath is untouched.

With real numbersrun_variants.sh and the Supervised Fine-Tuning video

  • LoRA and DoRA: learning rate 1e-4 (0.0001). Full: 1e-5 (0.00001), 10 times smaller.
  • Full moves all 490 million numbers. LoRA at rank 16 (the calculator’s example) moves about 2.2 million (0.44%) and the rest stay frozen. Your training log’s Trainable parameters line gives today’s exact share.
  • The Supervised Fine-Tuning video gives the same pair: about 1e-4 for LoRA, about 10 times smaller for full.
  • The damage would show in the general quiz: the untrained model scores about 9 of 12, and a drop below that is usually in the full row.

Words to know

Learning rate
How big a nudge each training step gives each number. Example: 1e-4 (0.0001) for LoRA, 1e-5 (0.00001) for full.
Forgetting
Getting better at the trained task and worse at everything else. Example: a lower general quiz score.
General quiz
Day 4’s 12 quick questions (math, facts, format), reused here to spot forgetting. Example: about 9 of 12 for the untrained model.
Training step
One update: the model sees a few examples and adjusts. Example: 300 steps of 4 tickets.
Go deeper: the engineer version

The kit's question

Why is full fine-tuning’s learning rate 10× smaller?

The kit's answer

Every weight moves, so a big step damages what the model already knows.

More detail: An Adam update moves each weight by roughly the learning rate, so 1e-5 across every weight is still a large total change to the network. Adapter updates are confined to a low-rank correction that starts at zero (B = 0), so a larger rate is safe. Too large a rate in full fine-tuning shows up as a general-quiz drop, the forgetting that claim 3 of the lesson looks for.

Question Should we fine-tune the whole model or use LoRA?

One clear answer

I ran full, LoRA and DoRA on the same data. LoRA got within a couple of points of full at a fraction of the memory, with a 9 MB checkpoint instead of a gigabyte, and without denting general ability. For a narrow task, I start with LoRA and only go full when the eval gap is real.

What this means

  • “I ran full, LoRA and DoRA on the same data”: I trained the same model three ways on the same 800 tickets with the same settings, so the comparison is fair. Quote your own table’s numbers.
  • “LoRA got within a couple of points of full”: LoRA’s task accuracy landed within about 2 percentage points of full training. The Full Fine-Tuning video says the same: usually within 1 to 2 points on a narrow task.
  • “at a fraction of the memory”: About 2 to 3 GB at peak for LoRA against about 8 to 10 GB for full, on the 0.5B model. On an 8B model: 16.2 GB against 128 GB.
  • “with a 9 MB checkpoint instead of a gigabyte”: The lesson expects the saved LoRA add-on at about 9 MB, and the fully trained model as a whole new copy of about 990 MB. Say your own adapters.safetensors sizes from today’s run. Small files are why one server can hold many customers’ add-ons.
  • “and without denting general ability”: The 12-question general quiz stayed about where the untrained model was (about 9 of 12): no forgetting.
  • “For a narrow task, I start with LoRA”: For one well-defined job, like sorting tickets into 4 categories, LoRA is my first choice.
  • “and only go full when the eval gap is real”: I switch to full fine-tuning only if tests on the customer’s own examples show LoRA falling measurably short.

Your numbersSaved on this device and collected in the Day 11 wrap-up.

Hint: The peak_mem_GB column of compare_variants.py. The table lists dora first; write them as full / LoRA / DoRA, separated by slashes.

Hint: The size of each saved file: run ls -lh adapters/*/adapters.safetensors in the lesson folder (M means about a million bytes). The ckpt_MB column adds up the whole folder, including the trainer’s in-between copies, so it can read several times higher. Write them as full / LoRA / DoRA.

Hint: The task_acc_% column: tickets whose category the model got right, out of 100 held-out tickets. The table lists the untrained model, then dora, full, lora; write them in the label’s order.

Hint: The general_quiz column of compare_variants.py. The table lists the untrained model, then dora, full, lora; write them in the label’s order.

compare_variants.py has printed one table with rows for the untrained model and for full, LoRA and DoRA, showing peak memory, file size, task accuracy and the general quiz.

The full run stops with an out-of-memory error, or the Mac slows to a crawl.
Full training peaks at about 8 to 10 GB. Close big apps and try again, or skip it. To skip it, first remove what the failed run left behind in the lesson folder, or compare_variants.py will try to test it: rm -rf logs/full.log adapters/full. Then run VARIANTS="lora dora" bash run_variants.sh. Record that full training did not fit on your Mac: it is a useful measured limit for your comparison.
Training takes too long.
Run ITERS=150 bash run_variants.sh: 150 steps instead of 300, about half the time. Note it next to your numbers.
run_variants.sh stops with fix the data first.
sft_data_check.py found a problem in the training file and printed its type with the first line number. If you edited data/tickets/, rebuild it from the 03-labs folder with python data/make_tickets.py.
compare_variants.py says No logs yet: run bash run_variants.sh first.
It reads the logs/ folder inside lessons/17-full-vs-peft-mlx. Run bash run_variants.sh there first, then python compare_variants.py from the same folder.
compare_variants.py says the model evaluation was skipped.
The accuracy and quiz columns need mlx-lm on an Apple-silicon Mac. Turn the environment on (source .venv/bin/activate from 03-labs) and run it again. The memory, time and file-size columns come from the logs either way.
mlx_lm.lora: command not found
mlx-lm is missing from the environment. From 03-labs, run source .venv/bin/activate, then pip install -r requirements.txt (Day 1’s setup), then bash run_variants.sh again.
compare_variants.py Python · 99 lines
"""
Lesson 17 · step 3 — one table: what each method cost, what it bought, what it broke.
From the training logs (works anywhere):
trainable %, peak memory, final validation loss, wall time, checkpoint size on disk
From the models themselves (Apple silicon, needs mlx-lm):
task accuracy triage category on the 100 held-out tickets (30 in phrasings never trained on)
general quiz the 12 general questions from lesson 04, which checks for FORGETTING (video 14)
python compare_variants.py # logs + evaluation
python compare_variants.py --logs-only # skip the model evaluation
"""
import argparse
import importlib.util
import json
import re
from pathlib import Path
from felab import record, table
from felab.tickets import SCHEMA, SYSTEM, grade, load_test
HERE = Path(__file__).parent
MODEL = "mlx-community/Qwen2.5-0.5B-Instruct-bf16"
def parse_log(path: Path) -> dict:
txt = path.read_text(errors="ignore")
m = re.search(r"Trainable parameters:\s*([\d.]+)%\s*\(([\d.]+)M", txt)
peaks = [float(x) for x in re.findall(r"Peak mem\s*([\d.]+)\s*GB", txt)]
vals = re.findall(r"Val loss\s*([\d.]+)", txt)
wall = re.search(r"wall_seconds\s+(\d+)", txt)
return {
"trainable_%": float(m.group(1)) if m else float("nan"),
"trainable_M": float(m.group(2)) if m else float("nan"),
"peak_mem_GB": max(peaks) if peaks else float("nan"),
"val_loss": float(vals[-1]) if vals else float("nan"),
"minutes": int(wall.group(1)) / 60 if wall else float("nan"),
}
def dir_mb(p: Path) -> float:
return sum(f.stat().st_size for f in p.rglob("*") if f.is_file()) / 1e6 if p.exists() else float("nan")
def evaluate(adapter: str | None) -> tuple[float, int]:
"""Greedy-decode the test tickets and the general quiz with mlx-lm."""
from mlx_lm import generate, load
model, tok = load(MODEL, adapter_path=adapter)
def ask(messages, max_tokens):
prompt = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
return generate(model, tok, prompt=prompt, max_tokens=max_tokens, verbose=False)
rows = load_test()
ok = sum(grade(ask([{"role": "system", "content": SYSTEM}, {"role": "user", "content": r["ticket"]}], 60), r)["category_ok"]
for r in rows)
spec = importlib.util.spec_from_file_location("q", HERE.parent / "04-quantization" / "quality_check.py")
q = importlib.util.module_from_spec(spec)
spec.loader.exec_module(q) # safe: its main() only runs under __main__
quiz = sum(want.replace(" ", "").lower() in ask([{"role": "user", "content": qq}], 40).replace(" ", "").lower()
for qq, want in q.QA)
return 100 * ok / len(rows), quiz
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--logs-only", action="store_true")
a = ap.parse_args()
variants = [p.stem for p in sorted((HERE / "logs").glob("*.log"))]
if not variants:
raise SystemExit("No logs yet: run bash run_variants.sh first.")
rows = []
can_eval = not a.logs_only and importlib.util.find_spec("mlx_lm") is not None
if can_eval:
print("evaluating base model …", flush=True)
acc, quiz = evaluate(None)
rows.append({"variant": "base (no training)", "trainable_%": 0.0, "peak_mem_GB": float("nan"),
"val_loss": float("nan"), "minutes": 0.0, "ckpt_MB": 0.0, "task_acc_%": acc, "general_quiz": f"{quiz}/12"})
for v in variants:
r = {"variant": v, **parse_log(HERE / "logs" / f"{v}.log"), "ckpt_MB": dir_mb(HERE / "adapters" / v)}
r.pop("trainable_M")
if can_eval:
print(f"evaluating {v} …", flush=True)
acc, quiz = evaluate(str(HERE / "adapters" / v))
r.update({"task_acc_%": acc, "general_quiz": f"{quiz}/12"})
rows.append(r)
record("17-variants", r)
print("\n" + table(rows))
print("\nRead across a row: what it cost (memory, minutes, MB) → what it bought (task accuracy) →"
"\nwhat it broke (general quiz; a drop means forgetting). On a narrow task, LoRA and DoRA"
"\nshould land near full fine-tuning at a fraction of the memory and checkpoint size.")
if not can_eval:
print("\n(model evaluation skipped: needs mlx-lm on Apple silicon)")
if __name__ == "__main__":
main()
peft_params.py Python · 63 lines
"""
Lesson 17 · step 1 — how many parameters does each method train, and how much memory does it need?
Counts come from the model's real shapes (config.json), not rules of thumb.
For a weight matrix of shape (in × out), a LoRA adapter of rank r adds r·(in + out) parameters.
DoRA adds one magnitude value per output channel on top.
Memory to TRAIN (weights + gradients + Adam state; activations excluded):
full 16 B × all params (bf16 weights & grads, fp32 Adam m, v, master copy)
LoRA 2 B × base + 16 B × adapter params (bf16 frozen base)
QLoRA ≈0.56 B × base + 16 B × adapter params (4-bit frozen base)
python peft_params.py --model 8b --rank 16
python peft_params.py --model 70b --rank 64 --mlp
python peft_params.py --model 0.5b # what lesson 17 trains on your Mac
"""
import argparse
# hidden, layers, attention heads, kv heads, head_dim, mlp (intermediate) — from each config.json
MODELS = {
"0.5b": ("Qwen2.5-0.5B", 896, 24, 14, 2, 64, 4864, 0.49e9),
"7b": ("Qwen2.5-7B", 3584, 28, 28, 4, 128, 18944, 7.6e9),
"8b": ("Llama-3.1-8B", 4096, 32, 32, 8, 128, 14336, 8.0e9),
"70b": ("Llama-3.1-70B", 8192, 80, 64, 8, 128, 28672, 70.6e9),
}
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--model", choices=MODELS, default="8b")
ap.add_argument("--rank", type=int, default=16)
ap.add_argument("--mlp", action="store_true", help="also adapt the MLP (gate, up, down)")
ap.add_argument("--layers", type=int, help="adapt only the last N layers (default: all)")
a = ap.parse_args()
name, d, L, H, KV, hd, ff, total = MODELS[a.model]
layers = a.layers or L
mats = {"q": (d, H * hd), "k": (d, KV * hd), "v": (d, KV * hd), "o": (H * hd, d)}
if a.mlp:
mats.update({"gate": (d, ff), "up": (d, ff), "down": (ff, d)})
lora_per_layer = sum(a.rank * (i + o) for i, o in mats.values())
dora_per_layer = lora_per_layer + sum(o for _, o in mats.values())
lora = lora_per_layer * layers
dora = dora_per_layer * layers
print(f"{name}: {total / 1e9:.1f}B params · hidden {d} · {L} layers · GQA {H}q/{KV}kv heads")
print(f"adapting {', '.join(mats)} in {layers} layers at rank {a.rank}\n")
print(" per layer:")
for m, (i, o) in mats.items():
print(f" {m:5s} {i:>6} × {o:<6} full {i * o / 1e6:7.2f} M LoRA r·(in+out) = {a.rank * (i + o) / 1e3:7.1f} K")
GB = 1e9
rows = [
("full fine-tune", total, 16 * total),
("LoRA", lora, 2 * total + 16 * lora),
("DoRA", dora, 2 * total + 16 * dora),
("QLoRA (4-bit base)", lora, 0.5625 * total + 16 * lora),
]
print(f"\n {'method':20s} {'trainable':>14s} {'share':>8s} {'train memory*':>14s} {'checkpoint':>12s}")
for m, t, mem in rows:
ckpt = f"{2 * total / GB:.1f} GB" if m.startswith("full") else f"{2 * t / 1e6:.0f} MB"
print(f" {m:20s} {t:>14,.0f} {100 * t / total:7.3f}% {mem / GB:11.1f} GB {ckpt:>12s}")
print("\n * weights + gradients + optimizer only. Activations add more and grow with batch × sequence length;"
"\n gradient checkpointing shrinks them. Compare with the 'Peak mem' mlx_lm.lora prints.")
run_variants.sh Bash · 40 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 17 · step 2 — the same model, same data, three ways (videos 14, 15, 16).
#
# full every weight trains lr 1e-5 (full FT needs a smaller step)
# lora frozen base + low-rank B·A in every layer lr 1e-4
# dora LoRA on direction + a magnitude vector lr 1e-4
#
# Model: Qwen2.5-0.5B-Instruct in bf16. Small enough that FULL fine-tuning fits on a
# 16 GB Mac (≈0.5B × 16 bytes ≈ 8 GB + activations). QLoRA needs a quantized base; for
# full fine-tuning the base must NOT be quantized, so we use bf16 for all three.
#
# --mask-prompt loss on the assistant answer only (video 15)
# --num-layers -1 all layers, so the three runs are comparable
# --grad-checkpoint recompute activations: less memory, a bit slower (video 14)
#
# ~5–15 min per variant on M-series. Logs → logs/<variant>.log, weights → adapters/<variant>/
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
cd "$(dirname "$0")"
MODEL=${MODEL:-mlx-community/Qwen2.5-0.5B-Instruct-bf16}
ITERS=${ITERS:-300}
VARIANTS=(${VARIANTS:-full lora dora})
DATA=../../data/tickets
mkdir -p logs adapters
python sft_data_check.py || { echo "fix the data first"; exit 1; }
for v in "${VARIANTS[@]}"; do
lr=1e-4; [[ $v == full ]] && lr=1e-5
echo "▸ $v (lr $lr, $ITERS iters)"
start=$SECONDS
mlx_lm.lora --model "$MODEL" --train --data "$DATA" \
--fine-tune-type "$v" --num-layers -1 --mask-prompt --grad-checkpoint \
--batch-size 4 --iters "$ITERS" --learning-rate "$lr" \
--steps-per-report 20 --steps-per-eval 100 \
--adapter-path "adapters/$v" 2>&1 | tee "logs/$v.log"
echo "wall_seconds $(( SECONDS - start ))" >> "logs/$v.log"
done
echo "▸ next: python compare_variants.py"
sft_data_check.py Python · 95 lines
"""
Lesson 17 · step 0 — check an SFT dataset BEFORE you spend compute on it (video 15).
Catches the mistakes that quietly ruin training runs:
✗ lines that are not valid JSON, or have no "messages"
✗ conversations that do not END with an assistant turn (nothing to learn)
✗ empty contents, unknown roles
✗ exact duplicates (the model over-learns them)
✗ test-set leakage: training examples whose user text appears in the held-out test set
✗ rows longer than your max sequence length (silently truncated = broken answers)
~ label balance (for the triage task: category counts parsed from the answers)
python sft_data_check.py # data/tickets/train.jsonl vs test.jsonl
python sft_data_check.py my_train.jsonl --test my_test.jsonl --max-tokens 2048
Exit code 1 if any ✗ check fails, so it can gate a training script.
"""
import argparse
import collections
import json
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("train", nargs="?", default=str(ROOT / "data/tickets/train.jsonl"))
ap.add_argument("--test", default=str(ROOT / "data/tickets/test.jsonl"))
ap.add_argument("--max-tokens", type=int, default=1024, help="trainer's max sequence length")
a = ap.parse_args()
rows, problems = [], collections.defaultdict(list)
for n, line in enumerate(Path(a.train).read_text().splitlines(), 1):
if not line.strip():
continue
try:
obj = json.loads(line)
except json.JSONDecodeError:
problems["not valid JSON"].append(n)
continue
msgs = obj.get("messages")
if not isinstance(msgs, list) or not msgs:
problems['missing "messages"'].append(n)
continue
roles = [m.get("role") for m in msgs]
if any(r not in ("system", "user", "assistant", "tool") for r in roles):
problems["unknown role"].append(n)
if roles[-1] != "assistant":
problems["does not end with an assistant turn"].append(n)
if any(not str(m.get("content", "")).strip() for m in msgs):
problems["empty content"].append(n)
toks = sum(len(str(m.get("content", ""))) for m in msgs) // 4 + 8 * len(msgs) # ~4 chars/token + template
if toks > a.max_tokens:
problems[f"longer than --max-tokens {a.max_tokens}"].append(n)
rows.append((n, obj, toks))
# duplicates
seen = {}
for n, obj, _ in rows:
key = json.dumps(obj["messages"], sort_keys=True)
if key in seen:
problems["exact duplicate"].append(n)
seen.setdefault(key, n)
# test leakage: user text that also appears in the test set
test_texts = set()
if Path(a.test).exists():
for line in Path(a.test).read_text().splitlines():
if line.strip():
t = json.loads(line)
test_texts.add((t.get("ticket") or next((m["content"] for m in t.get("messages", []) if m["role"] == "user"), "")).strip())
for n, obj, _ in rows:
user = next((m["content"] for m in obj["messages"] if m["role"] == "user"), "").strip()
if user in test_texts:
problems["user text also in the TEST set (leakage)"].append(n)
# label balance (triage-specific, skipped if answers are not JSON)
cats = collections.Counter()
for _, obj, _ in rows:
try:
cats[json.loads(obj["messages"][-1]["content"]).get("category", "?")] += 1
except Exception:
pass
toks = sorted(t for *_, t in rows)
print(f"{a.train}\n rows {len(rows)} tokens p50 ≈{toks[len(toks) // 2] if toks else 0} max ≈{toks[-1] if toks else 0}")
if cats:
total = sum(cats.values())
print(" label balance " + " ".join(f"{c} {100 * k / total:.0f}%" for c, k in cats.most_common()))
blocking = {k: v for k, v in problems.items()}
if not blocking:
print(" ✓ no problems found")
for k, v in blocking.items():
print(f" ✗ {k}: {len(v)} rows (first: line {v[0]})")
if "exact duplicate" in blocking:
print("\nDuplicates make the model over-learn those rows; dedupe before training.")
sys.exit(1 if blocking else 0)