Full vs LoRA vs DoRA
course 20 of 22
lesson 17 · lessons/17-full-vs-peft-mlx
1 h 05 Free Mac
What you will do and why
You train one real model three ways on the same 800 support tickets: change every number, or train a small add-on (LoRA, and its variant DoRA). One table then shows what each cost, what it gained and what it broke.
Why it matters: Lesson example: the fully trained model saves a file of about 990 MB; the LoRA add-on saves about 9 MB, 110 times smaller, with accuracy close to full.
You are done when: compare_variants.py has printed one table with rows for the untrained model and for full, LoRA and DoRA, showing peak memory, file size, task accuracy and the general quiz.
Start this first
Run it from the 03-labs folder with the Python environment on (source .venv/bin/activate). It checks the data, downloads the model the first time (about 1 GB) and trains it three times: 15 to 45 minutes in all. Leave it running while you watch. For the other commands, open a second terminal window, go to 03-labs, run source .venv/bin/activate, then cd lessons/17-full-vs-peft-mlx.
cd lessons/17-full-vs-peft-mlx && bash run_variants.shSupervised Fine-Tuning · 4:02
Download mp4 (15.2 MB)Chapters
In this video What training-by-example data looks like, why only the answer is graded, how much data you need, and how to read the training curves.
3 key points
Supervised fine-tuning (SFT) teaches by example: a prompt in, the ideal answer out, and only the answer is graded.
Each ticket example holds a system message (the standing instructions), the customer’s ticket and the ideal JSON answer. The trainer skips the first two when scoring (masking the prompt);
run_variants.shturns this on with--mask-prompt.A few hundred clean examples beat many noisy ones.
Typical ranges for ticket triage with LoRA on a 7B model: 50 examples mostly fix the format, 500 add 8 to 12 points of accuracy, 2,000 add 12 to 16, then it flattens. 20,000 noisy ones often do worse than 2,000 clean. Today’s data: 800 training tickets.
Read the two loss curves together: training loss (how wrong it is on the examples it learns from) and validation loss (how wrong on examples held back).
Both falling: it is learning. Training falling while validation rises: it is memorising, so stop earlier or add data. Both flat and high: check the data, or raise the learning rate (the size of each nudge). The video’s usual starting values: about 0.0001 (written 1e-4) for LoRA, 10 times smaller for full.
In plain words
Section titled “In plain words”Today you train one small model three ways on the same 800 support tickets. Full fine-tuning changes every number in it. LoRA freezes it and trains a small add-on; DoRA is a LoRA variant. One table then shows what each cost (memory, time, file size), what it gained (accuracy) and what it broke (general knowledge).
Picture it
Full fine-tuning reprints the whole textbook with your changes, and while it works it needs about 8 times the book’s room for drafts and notes. LoRA sticks a pad of notes on the pages that matter: cheap, easy to swap, and the original text stays readable underneath. With LoRA almost all the weight you carry is the book itself.
With real numbersQwen2.5-0.5B-Instruct, lesson 17’s calculator and expected table
- The model: 490 million numbers, stored at 2 bytes each (bf16, a 16-bit format): 0.98 GB, about 1 GB.
- Full fine-tuning: 490 million x 16 bytes = 7.8 GB, 8 times the model. The working numbers for each batch of examples (activations) come on top: the lesson expects a peak of about 8 to 10 GB.
- LoRA at rank 16 (the add-on’s size setting, in the calculator’s example): 2,162,688 trained numbers, 0.44% of the model. Memory: the 0.98 GB frozen model plus 2.16 million x 16 bytes (0.03 GB), about 1.0 GB. With activations, about 2 to 3 GB at peak.
- Saved file, from the lesson’s expected table: about 990 MB for full (a whole new model) against about 9 MB for LoRA, 110 times smaller. Your own run’s sizes are the real figures.
- The same sums for an 8B model: 128 GB to train fully, 16.2 GB with LoRA, 4.7 GB with QLoRA (LoRA on a 4-bit copy of the model).
Words to know
- Full fine-tuning
- Training that changes every number in the model. Example: 7.8 GB of training memory for the 0.5B model.
- LoRA and DoRA
- Freeze the model and train a small add-on beside it. DoRA also trains one extra number per row of each grid, setting how strong that row is; it often scores 1 to 2 accuracy points higher and trains a bit slower. Example: under 1% of the numbers trained.
- Checkpoint
- The file a training run saves: the whole changed model for full fine-tuning, only the add-on for LoRA. Example: about 990 MB against about 9 MB in the lesson’s expected table.
- Forgetting
- Getting better at the trained task and worse at everything else. Example: a drop in the 12-question general quiz.
From the lesson
lessons/17-full-vs-peft-mlx/README.md
What and why
Section titled “What and why”Today’s three videos (Full Fine-Tuning, Supervised Fine-Tuning and Parameter-Efficient Fine-Tuning) make three claims. This lesson tests all three on a real model:
- Full fine-tuning costs about 16 bytes per parameter to train, against about 2 to serve.
- LoRA and DoRA train under 1% of the weights and land close to full fine-tuning on a narrow task.
- Full fine-tuning risks more forgetting, meaning getting better at your task and worse at everything else.
You train Qwen2.5-0.5B-Instruct three ways on the same ticket data. Then you read cost (memory, minutes, checkpoint size), gain (task accuracy) and damage (a general quiz) from one table. The model is small enough that full fine-tuning fits on a 16 GB Mac.
The files, in order
Section titled “The files, in order”| step | file | runs on | does |
|---|---|---|---|
| 0 | sft_data_check.py |
anywhere | checks the data: JSON, roles, ends-with-assistant, duplicates, test leakage, length, label balance |
| 1 | peft_params.py |
anywhere | trainable parameters and training memory for full, LoRA, DoRA and QLoRA, from real model shapes |
| 2 | run_variants.sh |
Mac | mlx_lm.lora --fine-tune-type full / lora / dora, same data and settings |
| 3 | compare_variants.py |
Mac | one table: trainable %, peak memory, val loss, minutes, checkpoint MB, task accuracy, general quiz |
This checker earned its place while the course was being built. On its first run it found that the synthetic dataset had 301 training tickets identical to test tickets, which would have inflated every fine-tune score.
data/make_tickets.pynow dedupes, keeps the splits disjoint and keeps one phrasing per category for the test set only.
cd lessons/17-full-vs-peft-mlxpython sft_data_check.pypython peft_params.py --model 0.5b # what you are about to trainpython peft_params.py --model 8b --rank 16 # the same maths at a size customers usepython peft_params.py --model 70b --rank 64 --mlp
bash run_variants.sh # ~15–45 min for all threepython compare_variants.pyWhat each command does
python sft_data_check.pyChecks the 800 training tickets before any training. It looks for unreadable lines, unknown speakers (each message must be marked as the instructions, the customer, the model or a tool), examples that do not end with the answer, duplicates and over-long examples. It also catches any ticket that also sits in the 100-ticket test set (a leak). Look for
rows 800andno problems found.python peft_params.py --model 0.5bWorks out, from the model’s real shape, how many numbers each method trains and how much memory training needs for today’s 0.5B model. Look for 7.8 GB for full against about 1.0 GB for LoRA, which trains 2,162,688 numbers (0.441%). Its LoRA file size says about 4 MB, not the lesson’s expected 9 MB: the calculator assumes rank 16 and 2 bytes per saved number, while
run_variants.shuses the trainer’s own defaults. Your own run’s files give the real size.python peft_params.py --model 8b --rank 16The same sums for Llama 3.1 8B, a size customers use. Look for 128.0 GB for full, 16.2 GB for LoRA and 4.7 GB for QLoRA (LoRA on a 4-bit copy of the model); the saved file is 16.0 GB for full against 27 MB for LoRA. The Full Fine-Tuning video’s rough figures are about 20 million numbers, 20 GB and 40 MB; the calculator counts rank 16 on the attention grids exactly.
python peft_params.py --model 70b --rank 64 --mlpA 70B model at a higher rank, also adapting the MLP grids (the other big block in every layer). Look for 1129.6 GB for full, about 1.1 TB, against 154.5 GB for LoRA and 53.0 GB for QLoRA.
bash run_variants.shChecks the data again, then trains the model three ways, full, LoRA and DoRA, with the same tickets and settings: 300 steps of 4 examples (1,200 examples, 1.5 passes over the 800). Full uses a learning rate (the size of each nudge) 10 times smaller: 0.00001 against 0.0001, printed 1e-5 and 1e-4. Expect 5 to 15 minutes per method. If you started it before the video, let that run finish instead of starting a second one. Look for
Val loss(how wrong it is on 100 held-back tickets) falling every 100 steps, and aPeak memfigure in each run.python compare_variants.pyReads the three training logs, then tests the untrained model and each trained one on two things: the category of 100 held-out tickets (30 in wordings never seen in training), and Day 4’s 12 quick general questions. Look for task accuracy jumping from about 40% to high for all three, LoRA and DoRA close to full, and any quiz drop in the
fullrow.
Tight on memory? Run VARIANTS="lora dora" bash run_variants.sh and skip full. Short on
time? Use ITERS=150.
What you should see (shape, not exact numbers)
Section titled “What you should see (shape, not exact numbers)”How to read it
Each row is one method. trainable_% is how much of the model it trained; peak_mem_GB the most memory it used; minutes how long it took; val_loss how wrong it still was on held-back tickets. ckpt_MB adds up everything in the run’s folder, including the trainer’s in-between copies, so it can read several times the saved file. task_acc_% is the share of test tickets sorted correctly, and general_quiz the 12-question check for forgetting. The lesson’s table is the expected shape, not exact figures, and your trainer version may count a little differently. Yours lists the untrained model first, then dora, full, lora (alphabetical). Read across a row: what it cost, what it bought, what it broke.
variant trainable_% peak_mem_GB val_loss minutes ckpt_MB task_acc_% general_quizbase (no training) 0.0 nan nan 0.0 0.0 ~40 ~9/12full 100.0 ~8–10 lowest slowest ~990 high may droplora ~0.4 ~2–3 close fast ~9 high ≈ basedora ~0.45 ~3 close a bit slower ~9 high ≈ base- Peak memory: full is several times LoRA’s. That is claim 1.
- Checkpoint: about 1 GB for full against single-digit MB. That is why multi-LoRA serving works.
- Task accuracy: all three jump well above base, and LoRA and DoRA land close to full. That is claim 2. Look at the held-out phrasings too, since they test generalisation rather than memory.
- General quiz: if any variant drops, it is usually full. That is claim 3. On 12 questions it is a smoke alarm, not proof.
Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
Day 11’s real training run starts from a model stored at 16 bits per number (bf16). On Day 8 you trained an add-on for a 7-billion-number model stored at 4 bits. Why must this comparison of full fine-tuning and LoRA use a 16-bit model, not a 4-bit one?Show answerHide
In plain words
Full fine-tuning changes every stored number, and 4-bit numbers are too coarse to take the tiny nudges training makes. Day 8’s QLoRA worked only because it never changed the 4-bit numbers: it trained a separate add-on instead.
Picture it
A 4-bit number is like a ruler marked only in whole centimetres. Training wants to move each mark by a hair’s width, and on that ruler the nudge rounds away to nothing. QLoRA leaves the coarse ruler alone and writes the fine corrections on a separate note.
With real numberslesson 17, and Day 8’s lesson 11
- 4 bits can hold only 2 x 2 x 2 x 2 = 16 different values; 16 bits can hold 65,536.
- Day 11’s model (Qwen2.5-0.5B-Instruct) at 16 bits: 490 million numbers x 2 bytes = 0.98 GB, small enough to train fully on a 16 GB Mac (about 8 to 10 GB at peak).
- Day 8’s QLoRA: a 7B model frozen at 4 bits plus a small 16-bit add-on, about 6 to 10 GB to train. A full fine-tune of the same model needs about 120 GB (it has 7.6 billion numbers: 7.6 x 16 = 122 GB).
Words to know
- bf16
- A 16-bit number format used for training, 2 bytes per number. Example: Day 11’s model is 0.98 GB in bf16.
- QLoRA
- LoRA on a model stored in 4 bits: the model stays frozen, only the add-on trains. Example: about 6 to 10 GB for a 7B model on Day 8.
- Frozen
- Left unchanged during training. Example: the 4-bit model in QLoRA.
- Gradient
- The correction worked out for each trainable number at each training step. Example: full fine-tuning keeps one for every number.
Go deeper: the engineer version
The kit's question
Why must the base be bf16 here, and not the 4-bit model from lesson 11?
The kit's answer
Full fine-tuning updates the weights, and you can’t train 4-bit weights directly. QLoRA only works because the base stays frozen.
More detail: Quantized weights sit on a coarse grid, and rounding to that grid gives no useful gradient, so a training step cannot move them reliably. QLoRA keeps the 4-bit base frozen (only read, and turned back into 16-bit numbers for the math) and sends gradients only into the 16-bit adapter. run_variants.sh uses the bf16 base for all three runs so that full, LoRA and DoRA compare fairly.
Day 11’s memory calculator (peft_params.py --model 8b) says LoRA needs about 16 GB to train, but full fine-tuning needs 128 GB. Where do LoRA’s 16 GB go?Show answerHide
In plain words
Almost all of it is the model itself, frozen and stored at 2 bytes per number. The add-on and its training state are tiny. Store the frozen model in 4 bits instead (QLoRA) and it drops to about 5 GB.
Picture it
You hire a moving truck to deliver one small box: nearly all the cost is the truck, not the box. In LoRA the truck is the frozen model and the box is the add-on. QLoRA books a smaller truck: the same model squeezed into 4 bits.
With real numberspeft_params.py --model 8b --rank 16, lesson 17
- Frozen model: 8 billion numbers x 2 bytes = 16 GB.
- LoRA add-on at rank 16: 13.6 million numbers (0.17% of the model) x 16 bytes = 0.22 GB.
- Total: 16 + 0.22 = 16.2 GB. Full: 8 billion x 16 bytes = 128 GB, about 8 times as much.
- QLoRA: the frozen model at 4 bits is half a byte per number, plus a little for shared scale numbers: 0.5625 bytes. 8 billion x 0.5625 = 4.5 GB, plus 0.22 = 4.7 GB, about 5 GB.
Words to know
- Frozen model
- The original model, kept unchanged and only read during LoRA training. Example: 16 GB for an 8B model at 2 bytes per number.
- Adapter (add-on)
- The small set of extra numbers LoRA trains and saves. Example: 13.6 million numbers for an 8B model at rank 16.
- Optimizer state
- The running averages the optimizer (Adam) keeps per trained number to size each nudge. Example: 8 of the 16 bytes per number.
- QLoRA
- LoRA on a model stored in 4 bits. Example: 4.7 GB for an 8B model instead of 16.2 GB.
Go deeper: the engineer version
The kit's question
peft_params.py --model 8b says LoRA needs about 16 GB, but full needs 128. Where does LoRA’s 16 GB go?
The kit's answer
Almost all of it is the frozen bf16 base, 2 bytes × 8B. The adapter’s optimizer state is tiny. Switch to QLoRA and it drops to about 5 GB.
More detail: peft_params.py counts LoRA as 2 B x base parameters (frozen bf16 weights, no gradients or optimizer state) + 16 B x adapter parameters (bf16 weight and gradient, fp32 Adam m and v, fp32 master copy). QLoRA’s 0.5625 B per base weight is 4 bits plus the shared scales of each small block. Activations come on top of every figure and grow with batch x sequence length; --grad-checkpoint trims them.
In today’s runs, full fine-tuning makes nudges 10 times smaller than LoRA’s (a learning rate of 0.00001 against 0.0001). Why?Show answerHide
In plain words
Full fine-tuning moves every number in the model, including the ones that hold its general skills. Big steps across all of them would damage what the model already knows.
Picture it
Restoring a painting: if you may touch every inch, you use a fine brush and light strokes, or you ruin the parts that were already right. Painting on a clear sheet laid over it (LoRA), bolder strokes are safe, because the original underneath is untouched.
With real numbersrun_variants.sh and the Supervised Fine-Tuning video
- LoRA and DoRA: learning rate 1e-4 (0.0001). Full: 1e-5 (0.00001), 10 times smaller.
- Full moves all 490 million numbers. LoRA at rank 16 (the calculator’s example) moves about 2.2 million (0.44%) and the rest stay frozen. Your training log’s
Trainable parametersline gives today’s exact share. - The Supervised Fine-Tuning video gives the same pair: about 1e-4 for LoRA, about 10 times smaller for full.
- The damage would show in the general quiz: the untrained model scores about 9 of 12, and a drop below that is usually in the full row.
Words to know
- Learning rate
- How big a nudge each training step gives each number. Example: 1e-4 (0.0001) for LoRA, 1e-5 (0.00001) for full.
- Forgetting
- Getting better at the trained task and worse at everything else. Example: a lower general quiz score.
- General quiz
- Day 4’s 12 quick questions (math, facts, format), reused here to spot forgetting. Example: about 9 of 12 for the untrained model.
- Training step
- One update: the model sees a few examples and adjusts. Example: 300 steps of 4 tickets.
Go deeper: the engineer version
The kit's question
Why is full fine-tuning’s learning rate 10× smaller?
The kit's answer
Every weight moves, so a big step damages what the model already knows.
More detail: An Adam update moves each weight by roughly the learning rate, so 1e-5 across every weight is still a large total change to the network. Adapter updates are confined to a low-rank correction that starts at zero (B = 0), so a larger rate is safe. Too large a rate in full fine-tuning shows up as a general-quiz drop, the forgetting that claim 3 of the lesson looks for.
Explain what you learned
Section titled “Explain what you learned”Question Should we fine-tune the whole model or use LoRA?
One clear answer
I ran full, LoRA and DoRA on the same data. LoRA got within a couple of points of full at a fraction of the memory, with a 9 MB checkpoint instead of a gigabyte, and without denting general ability. For a narrow task, I start with LoRA and only go full when the eval gap is real.
What this means
- “I ran full, LoRA and DoRA on the same data”: I trained the same model three ways on the same 800 tickets with the same settings, so the comparison is fair. Quote your own table’s numbers.
- “LoRA got within a couple of points of full”: LoRA’s task accuracy landed within about 2 percentage points of full training. The Full Fine-Tuning video says the same: usually within 1 to 2 points on a narrow task.
- “at a fraction of the memory”: About 2 to 3 GB at peak for LoRA against about 8 to 10 GB for full, on the 0.5B model. On an 8B model: 16.2 GB against 128 GB.
- “with a 9 MB checkpoint instead of a gigabyte”: The lesson expects the saved LoRA add-on at about 9 MB, and the fully trained model as a whole new copy of about 990 MB. Say your own
adapters.safetensorssizes from today’s run. Small files are why one server can hold many customers’ add-ons. - “and without denting general ability”: The 12-question general quiz stayed about where the untrained model was (about 9 of 12): no forgetting.
- “For a narrow task, I start with LoRA”: For one well-defined job, like sorting tickets into 4 categories, LoRA is my first choice.
- “and only go full when the eval gap is real”: I switch to full fine-tuning only if tests on the customer’s own examples show LoRA falling measurably short.
Your numbersSaved on this device and collected in the Day 11 wrap-up.
Hint: The peak_mem_GB column of compare_variants.py. The table lists dora first; write them as full / LoRA / DoRA, separated by slashes.
Hint: The size of each saved file: run ls -lh adapters/*/adapters.safetensors in the lesson folder (M means about a million bytes). The ckpt_MB column adds up the whole folder, including the trainer’s in-between copies, so it can read several times higher. Write them as full / LoRA / DoRA.
Hint: The task_acc_% column: tickets whose category the model got right, out of 100 held-out tickets. The table lists the untrained model, then dora, full, lora; write them in the label’s order.
Hint: The general_quiz column of compare_variants.py. The table lists the untrained model, then dora, full, lora; write them in the label’s order.
Done when
Section titled “Done when”compare_variants.py has printed one table with rows for the untrained model and for full, LoRA and DoRA, showing peak memory, file size, task accuracy and the general quiz.
Stuck?
Section titled “Stuck?”- The full run stops with an out-of-memory error, or the Mac slows to a crawl.
- Full training peaks at about 8 to 10 GB. Close big apps and try again, or skip it. To skip it, first remove what the failed run left behind in the lesson folder, or
compare_variants.pywill try to test it:rm -rf logs/full.log adapters/full. Then runVARIANTS="lora dora" bash run_variants.sh. Record that full training did not fit on your Mac: it is a useful measured limit for your comparison. - Training takes too long.
- Run
ITERS=150 bash run_variants.sh: 150 steps instead of 300, about half the time. Note it next to your numbers. run_variants.shstops withfix the data first.sft_data_check.pyfound a problem in the training file and printed its type with the first line number. If you editeddata/tickets/, rebuild it from the03-labsfolder withpython data/make_tickets.py.compare_variants.pysaysNo logs yet: run bash run_variants.sh first.- It reads the
logs/folder insidelessons/17-full-vs-peft-mlx. Runbash run_variants.shthere first, thenpython compare_variants.pyfrom the same folder. compare_variants.pysays the model evaluation was skipped.- The accuracy and quiz columns need mlx-lm on an Apple-silicon Mac. Turn the environment on (
source .venv/bin/activatefrom03-labs) and run it again. The memory, time and file-size columns come from the logs either way. mlx_lm.lora: command not found- mlx-lm is missing from the environment. From
03-labs, runsource .venv/bin/activate, thenpip install -r requirements.txt(Day 1’s setup), thenbash run_variants.shagain.
Go deeper: the lab book's training lab for this step All fixes
Code in this step
Section titled “Code in this step”compare_variants.py Python · 99 lines
"""Lesson 17 · step 3 — one table: what each method cost, what it bought, what it broke.
From the training logs (works anywhere): trainable %, peak memory, final validation loss, wall time, checkpoint size on disk
From the models themselves (Apple silicon, needs mlx-lm): task accuracy triage category on the 100 held-out tickets (30 in phrasings never trained on) general quiz the 12 general questions from lesson 04, which checks for FORGETTING (video 14)
python compare_variants.py # logs + evaluation python compare_variants.py --logs-only # skip the model evaluation"""import argparseimport importlib.utilimport jsonimport refrom pathlib import Path
from felab import record, tablefrom felab.tickets import SCHEMA, SYSTEM, grade, load_test
HERE = Path(__file__).parentMODEL = "mlx-community/Qwen2.5-0.5B-Instruct-bf16"
def parse_log(path: Path) -> dict: txt = path.read_text(errors="ignore") m = re.search(r"Trainable parameters:\s*([\d.]+)%\s*\(([\d.]+)M", txt) peaks = [float(x) for x in re.findall(r"Peak mem\s*([\d.]+)\s*GB", txt)] vals = re.findall(r"Val loss\s*([\d.]+)", txt) wall = re.search(r"wall_seconds\s+(\d+)", txt) return { "trainable_%": float(m.group(1)) if m else float("nan"), "trainable_M": float(m.group(2)) if m else float("nan"), "peak_mem_GB": max(peaks) if peaks else float("nan"), "val_loss": float(vals[-1]) if vals else float("nan"), "minutes": int(wall.group(1)) / 60 if wall else float("nan"), }
def dir_mb(p: Path) -> float: return sum(f.stat().st_size for f in p.rglob("*") if f.is_file()) / 1e6 if p.exists() else float("nan")
def evaluate(adapter: str | None) -> tuple[float, int]: """Greedy-decode the test tickets and the general quiz with mlx-lm.""" from mlx_lm import generate, load model, tok = load(MODEL, adapter_path=adapter)
def ask(messages, max_tokens): prompt = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False) return generate(model, tok, prompt=prompt, max_tokens=max_tokens, verbose=False)
rows = load_test() ok = sum(grade(ask([{"role": "system", "content": SYSTEM}, {"role": "user", "content": r["ticket"]}], 60), r)["category_ok"] for r in rows) spec = importlib.util.spec_from_file_location("q", HERE.parent / "04-quantization" / "quality_check.py") q = importlib.util.module_from_spec(spec) spec.loader.exec_module(q) # safe: its main() only runs under __main__ quiz = sum(want.replace(" ", "").lower() in ask([{"role": "user", "content": qq}], 40).replace(" ", "").lower() for qq, want in q.QA) return 100 * ok / len(rows), quiz
def main() -> None: ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--logs-only", action="store_true") a = ap.parse_args()
variants = [p.stem for p in sorted((HERE / "logs").glob("*.log"))] if not variants: raise SystemExit("No logs yet: run bash run_variants.sh first.") rows = [] can_eval = not a.logs_only and importlib.util.find_spec("mlx_lm") is not None if can_eval: print("evaluating base model …", flush=True) acc, quiz = evaluate(None) rows.append({"variant": "base (no training)", "trainable_%": 0.0, "peak_mem_GB": float("nan"), "val_loss": float("nan"), "minutes": 0.0, "ckpt_MB": 0.0, "task_acc_%": acc, "general_quiz": f"{quiz}/12"}) for v in variants: r = {"variant": v, **parse_log(HERE / "logs" / f"{v}.log"), "ckpt_MB": dir_mb(HERE / "adapters" / v)} r.pop("trainable_M") if can_eval: print(f"evaluating {v} …", flush=True) acc, quiz = evaluate(str(HERE / "adapters" / v)) r.update({"task_acc_%": acc, "general_quiz": f"{quiz}/12"}) rows.append(r) record("17-variants", r) print("\n" + table(rows)) print("\nRead across a row: what it cost (memory, minutes, MB) → what it bought (task accuracy) →" "\nwhat it broke (general quiz; a drop means forgetting). On a narrow task, LoRA and DoRA" "\nshould land near full fine-tuning at a fraction of the memory and checkpoint size.") if not can_eval: print("\n(model evaluation skipped: needs mlx-lm on Apple silicon)")
if __name__ == "__main__": main()peft_params.py Python · 63 lines
"""Lesson 17 · step 1 — how many parameters does each method train, and how much memory does it need?
Counts come from the model's real shapes (config.json), not rules of thumb.For a weight matrix of shape (in × out), a LoRA adapter of rank r adds r·(in + out) parameters.DoRA adds one magnitude value per output channel on top.
Memory to TRAIN (weights + gradients + Adam state; activations excluded): full 16 B × all params (bf16 weights & grads, fp32 Adam m, v, master copy) LoRA 2 B × base + 16 B × adapter params (bf16 frozen base) QLoRA ≈0.56 B × base + 16 B × adapter params (4-bit frozen base)
python peft_params.py --model 8b --rank 16 python peft_params.py --model 70b --rank 64 --mlp python peft_params.py --model 0.5b # what lesson 17 trains on your Mac"""import argparse
# hidden, layers, attention heads, kv heads, head_dim, mlp (intermediate) — from each config.jsonMODELS = { "0.5b": ("Qwen2.5-0.5B", 896, 24, 14, 2, 64, 4864, 0.49e9), "7b": ("Qwen2.5-7B", 3584, 28, 28, 4, 128, 18944, 7.6e9), "8b": ("Llama-3.1-8B", 4096, 32, 32, 8, 128, 14336, 8.0e9), "70b": ("Llama-3.1-70B", 8192, 80, 64, 8, 128, 28672, 70.6e9),}
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)ap.add_argument("--model", choices=MODELS, default="8b")ap.add_argument("--rank", type=int, default=16)ap.add_argument("--mlp", action="store_true", help="also adapt the MLP (gate, up, down)")ap.add_argument("--layers", type=int, help="adapt only the last N layers (default: all)")a = ap.parse_args()
name, d, L, H, KV, hd, ff, total = MODELS[a.model]layers = a.layers or Lmats = {"q": (d, H * hd), "k": (d, KV * hd), "v": (d, KV * hd), "o": (H * hd, d)}if a.mlp: mats.update({"gate": (d, ff), "up": (d, ff), "down": (ff, d)})
lora_per_layer = sum(a.rank * (i + o) for i, o in mats.values())dora_per_layer = lora_per_layer + sum(o for _, o in mats.values())lora = lora_per_layer * layersdora = dora_per_layer * layers
print(f"{name}: {total / 1e9:.1f}B params · hidden {d} · {L} layers · GQA {H}q/{KV}kv heads")print(f"adapting {', '.join(mats)} in {layers} layers at rank {a.rank}\n")print(" per layer:")for m, (i, o) in mats.items(): print(f" {m:5s} {i:>6} × {o:<6} full {i * o / 1e6:7.2f} M LoRA r·(in+out) = {a.rank * (i + o) / 1e3:7.1f} K")
GB = 1e9rows = [ ("full fine-tune", total, 16 * total), ("LoRA", lora, 2 * total + 16 * lora), ("DoRA", dora, 2 * total + 16 * dora), ("QLoRA (4-bit base)", lora, 0.5625 * total + 16 * lora),]print(f"\n {'method':20s} {'trainable':>14s} {'share':>8s} {'train memory*':>14s} {'checkpoint':>12s}")for m, t, mem in rows: ckpt = f"{2 * total / GB:.1f} GB" if m.startswith("full") else f"{2 * t / 1e6:.0f} MB" print(f" {m:20s} {t:>14,.0f} {100 * t / total:7.3f}% {mem / GB:11.1f} GB {ckpt:>12s}")print("\n * weights + gradients + optimizer only. Activations add more and grow with batch × sequence length;" "\n gradient checkpointing shrinks them. Compare with the 'Peak mem' mlx_lm.lora prints.")run_variants.sh Bash · 40 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 17 · step 2 — the same model, same data, three ways (videos 14, 15, 16).## full every weight trains lr 1e-5 (full FT needs a smaller step)# lora frozen base + low-rank B·A in every layer lr 1e-4# dora LoRA on direction + a magnitude vector lr 1e-4## Model: Qwen2.5-0.5B-Instruct in bf16. Small enough that FULL fine-tuning fits on a# 16 GB Mac (≈0.5B × 16 bytes ≈ 8 GB + activations). QLoRA needs a quantized base; for# full fine-tuning the base must NOT be quantized, so we use bf16 for all three.## --mask-prompt loss on the assistant answer only (video 15)# --num-layers -1 all layers, so the three runs are comparable# --grad-checkpoint recompute activations: less memory, a bit slower (video 14)## ~5–15 min per variant on M-series. Logs → logs/<variant>.log, weights → adapters/<variant>/# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"MODEL=${MODEL:-mlx-community/Qwen2.5-0.5B-Instruct-bf16}ITERS=${ITERS:-300}VARIANTS=(${VARIANTS:-full lora dora})DATA=../../data/ticketsmkdir -p logs adapters
python sft_data_check.py || { echo "fix the data first"; exit 1; }
for v in "${VARIANTS[@]}"; do lr=1e-4; [[ $v == full ]] && lr=1e-5 echo "▸ $v (lr $lr, $ITERS iters)" start=$SECONDS mlx_lm.lora --model "$MODEL" --train --data "$DATA" \ --fine-tune-type "$v" --num-layers -1 --mask-prompt --grad-checkpoint \ --batch-size 4 --iters "$ITERS" --learning-rate "$lr" \ --steps-per-report 20 --steps-per-eval 100 \ --adapter-path "adapters/$v" 2>&1 | tee "logs/$v.log" echo "wall_seconds $(( SECONDS - start ))" >> "logs/$v.log"doneecho "▸ next: python compare_variants.py"sft_data_check.py Python · 95 lines
"""Lesson 17 · step 0 — check an SFT dataset BEFORE you spend compute on it (video 15).
Catches the mistakes that quietly ruin training runs: ✗ lines that are not valid JSON, or have no "messages" ✗ conversations that do not END with an assistant turn (nothing to learn) ✗ empty contents, unknown roles ✗ exact duplicates (the model over-learns them) ✗ test-set leakage: training examples whose user text appears in the held-out test set ✗ rows longer than your max sequence length (silently truncated = broken answers) ~ label balance (for the triage task: category counts parsed from the answers)
python sft_data_check.py # data/tickets/train.jsonl vs test.jsonl python sft_data_check.py my_train.jsonl --test my_test.jsonl --max-tokens 2048Exit code 1 if any ✗ check fails, so it can gate a training script."""import argparseimport collectionsimport jsonimport sysfrom pathlib import Path
ROOT = Path(__file__).resolve().parents[2]ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)ap.add_argument("train", nargs="?", default=str(ROOT / "data/tickets/train.jsonl"))ap.add_argument("--test", default=str(ROOT / "data/tickets/test.jsonl"))ap.add_argument("--max-tokens", type=int, default=1024, help="trainer's max sequence length")a = ap.parse_args()
rows, problems = [], collections.defaultdict(list)for n, line in enumerate(Path(a.train).read_text().splitlines(), 1): if not line.strip(): continue try: obj = json.loads(line) except json.JSONDecodeError: problems["not valid JSON"].append(n) continue msgs = obj.get("messages") if not isinstance(msgs, list) or not msgs: problems['missing "messages"'].append(n) continue roles = [m.get("role") for m in msgs] if any(r not in ("system", "user", "assistant", "tool") for r in roles): problems["unknown role"].append(n) if roles[-1] != "assistant": problems["does not end with an assistant turn"].append(n) if any(not str(m.get("content", "")).strip() for m in msgs): problems["empty content"].append(n) toks = sum(len(str(m.get("content", ""))) for m in msgs) // 4 + 8 * len(msgs) # ~4 chars/token + template if toks > a.max_tokens: problems[f"longer than --max-tokens {a.max_tokens}"].append(n) rows.append((n, obj, toks))
# duplicatesseen = {}for n, obj, _ in rows: key = json.dumps(obj["messages"], sort_keys=True) if key in seen: problems["exact duplicate"].append(n) seen.setdefault(key, n)
# test leakage: user text that also appears in the test settest_texts = set()if Path(a.test).exists(): for line in Path(a.test).read_text().splitlines(): if line.strip(): t = json.loads(line) test_texts.add((t.get("ticket") or next((m["content"] for m in t.get("messages", []) if m["role"] == "user"), "")).strip())for n, obj, _ in rows: user = next((m["content"] for m in obj["messages"] if m["role"] == "user"), "").strip() if user in test_texts: problems["user text also in the TEST set (leakage)"].append(n)
# label balance (triage-specific, skipped if answers are not JSON)cats = collections.Counter()for _, obj, _ in rows: try: cats[json.loads(obj["messages"][-1]["content"]).get("category", "?")] += 1 except Exception: pass
toks = sorted(t for *_, t in rows)print(f"{a.train}\n rows {len(rows)} tokens p50 ≈{toks[len(toks) // 2] if toks else 0} max ≈{toks[-1] if toks else 0}")if cats: total = sum(cats.values()) print(" label balance " + " ".join(f"{c} {100 * k / total:.0f}%" for c, k in cats.most_common()))blocking = {k: v for k, v in problems.items()}if not blocking: print(" ✓ no problems found")for k, v in blocking.items(): print(f" ✗ {k}: {len(v)} rows (first: line {v[0]})")if "exact duplicate" in blocking: print("\nDuplicates make the model over-learn those rows; dedupe before training.")sys.exit(1 if blocking else 0)