About 10 minutes, free, and it works on your phone. Lock in today before you move on.
Done when
Your table puts full fine-tuning, LoRA and DoRA side by side on the same model and data: peak memory, saved file size, task accuracy and the general quiz. You can also say why the toy’s rank-1 add-on stayed stuck.
Fields you filled in on the step pages are already here. Nothing leaves this browser.
Step 1 · Training toys, part 1
How close each method gets on new examples (lower is better). Rank 2 matches full with 6% of the trained numbers; rank 1 cannot, because the task needs two kinds of change.
Example: 0.0002 / 0.1321 / 0.0000 (the lesson’s run)
error
Now the task needs 16 kinds of change. Even the biggest add-on the script tries (rank 8) stays above full: for a big shift, the rank must grow, or you train fully.
Example: 0.0002 / 0.0825 (a run of the script with its default seed; the last digits may differ)
error
Step 2 · Full vs LoRA vs DoRA
The most of your Mac’s memory each method needed. Full needs several times LoRA’s: the 16-bytes-per-number bill.
Example: about 8 to 10 / 2 to 3 / 3 (the lesson’s expected shape)
GB
What you store, ship and swap. A LoRA file this small is why one server can hold many customers’ add-ons (multi-LoRA).
Example: about 990 / 9 / 9 (the lesson’s expected files)
MB
What training bought. Expect all three well above the untrained model, with LoRA and DoRA close to full.
Example: untrained about 40; the three trained rows high and close together
%
What training broke, if anything: a drop means forgetting. On 12 questions it is a smoke alarm, not proof.
Example: untrained about 9; LoRA and DoRA about the same; full may drop
In Day 11’s toy (one small layer of 4,096 numbers), the rank-1 add-on never catches up with full fine-tuning, even with 200 examples. Why? (The rank is how many separate kinds of change the add-on can make at once.)
Show answerHide answer
In plain words
The change this task needs has two separate kinds (two directions), and a rank-1 add-on can make only one. More examples cannot supply the missing one; only a bigger rank can.
Picture it
Treasure lies to the north-east, and you may only move along a straight track running east. However many tries you get, you stop at the closest point on the track, never at the treasure. Add a second track running north and you reach it exactly.
With real numberslesson 16’s toy_finetune.py, the lesson’s sample run
The task needs a rank-2 change: two separate kinds of change (the script’s --true-rank is 2 by default).
Rank 1: 64 + 64 = 128 trained numbers, 3.1% of the layer. Error 0.3937, 0.1951 and 0.1321 at 16, 48 and 200 examples (0 is perfect).
Rank 2: 256 trained numbers, 6.2%. Error at 200 examples: 0.0000, as good as full fine-tuning (0.0002).
So rank 1 trains half as many numbers as rank 2 but misses one whole kind of change: at 200 examples its error is still 0.1321, while rank 2’s is 0.0000.
Words to know
Rank
How many separate kinds of change an add-on can make at once. Example: the toy’s task needs 2.
Held-out error
How wrong the trained layer is on new examples it never saw; lower is better. Example: 0.1321 for rank 1 at 200 examples.
Adapter (add-on)
The small set of extra numbers LoRA trains while the model stays frozen. Example: 128 numbers at rank 1.
Frozen
Left unchanged during training. Example: the toy’s pretrained layer W0.
Go deeper: the engineer version
The kit's question
Why does LoRA r=1 never catch up, even with 200 examples?
The kit's answer
The change the task needs is rank 2, and a rank-1 adapter cannot represent it.
More detail: LoRA’s update is B·A with B of shape 64 x r and A of shape r x 64, so its rank is at most r. The toy’s target shift dW_true is built from two random directions, so it has rank 2 (toy_finetune.py line 38). A rank-1 B·A can capture at best the strongest single direction of dW_true (its top singular direction); the remaining direction stays as error however many examples you add. With --true-rank 16 every rank the script tries (up to 8) falls short the same way.
Day 11’s real training run starts from a model stored at 16 bits per number (bf16). On Day 8 you trained an add-on for a 7-billion-number model stored at 4 bits. Why must this comparison of full fine-tuning and LoRA use a 16-bit model, not a 4-bit one?
Show answerHide answer
In plain words
Full fine-tuning changes every stored number, and 4-bit numbers are too coarse to take the tiny nudges training makes. Day 8’s QLoRA worked only because it never changed the 4-bit numbers: it trained a separate add-on instead.
Picture it
A 4-bit number is like a ruler marked only in whole centimetres. Training wants to move each mark by a hair’s width, and on that ruler the nudge rounds away to nothing. QLoRA leaves the coarse ruler alone and writes the fine corrections on a separate note.
With real numberslesson 17, and Day 8’s lesson 11
4 bits can hold only 2 x 2 x 2 x 2 = 16 different values; 16 bits can hold 65,536.
Day 11’s model (Qwen2.5-0.5B-Instruct) at 16 bits: 490 million numbers x 2 bytes = 0.98 GB, small enough to train fully on a 16 GB Mac (about 8 to 10 GB at peak).
Day 8’s QLoRA: a 7B model frozen at 4 bits plus a small 16-bit add-on, about 6 to 10 GB to train. A full fine-tune of the same model needs about 120 GB (it has 7.6 billion numbers: 7.6 x 16 = 122 GB).
Words to know
bf16
A 16-bit number format used for training, 2 bytes per number. Example: Day 11’s model is 0.98 GB in bf16.
QLoRA
LoRA on a model stored in 4 bits: the model stays frozen, only the add-on trains. Example: about 6 to 10 GB for a 7B model on Day 8.
Frozen
Left unchanged during training. Example: the 4-bit model in QLoRA.
Gradient
The correction worked out for each trainable number at each training step. Example: full fine-tuning keeps one for every number.
Go deeper: the engineer version
The kit's question
Why must the base be bf16 here, and not the 4-bit model from lesson 11?
The kit's answer
Full fine-tuning updates the weights, and you can’t train 4-bit weights directly. QLoRA only works because the base stays frozen.
More detail: Quantized weights sit on a coarse grid, and rounding to that grid gives no useful gradient, so a training step cannot move them reliably. QLoRA keeps the 4-bit base frozen (only read, and turned back into 16-bit numbers for the math) and sends gradients only into the 16-bit adapter. run_variants.sh uses the bf16 base for all three runs so that full, LoRA and DoRA compare fairly.
Day 11’s memory calculator (peft_params.py --model 8b) says LoRA needs about 16 GB to train, but full fine-tuning needs 128 GB. Where do LoRA’s 16 GB go?
Show answerHide answer
In plain words
Almost all of it is the model itself, frozen and stored at 2 bytes per number. The add-on and its training state are tiny. Store the frozen model in 4 bits instead (QLoRA) and it drops to about 5 GB.
Picture it
You hire a moving truck to deliver one small box: nearly all the cost is the truck, not the box. In LoRA the truck is the frozen model and the box is the add-on. QLoRA books a smaller truck: the same model squeezed into 4 bits.
With real numberspeft_params.py --model 8b --rank 16, lesson 17
LoRA add-on at rank 16: 13.6 million numbers (0.17% of the model) x 16 bytes = 0.22 GB.
Total: 16 + 0.22 = 16.2 GB. Full: 8 billion x 16 bytes = 128 GB, about 8 times as much.
QLoRA: the frozen model at 4 bits is half a byte per number, plus a little for shared scale numbers: 0.5625 bytes. 8 billion x 0.5625 = 4.5 GB, plus 0.22 = 4.7 GB, about 5 GB.
Words to know
Frozen model
The original model, kept unchanged and only read during LoRA training. Example: 16 GB for an 8B model at 2 bytes per number.
Adapter (add-on)
The small set of extra numbers LoRA trains and saves. Example: 13.6 million numbers for an 8B model at rank 16.
Optimizer state
The running averages the optimizer (Adam) keeps per trained number to size each nudge. Example: 8 of the 16 bytes per number.
QLoRA
LoRA on a model stored in 4 bits. Example: 4.7 GB for an 8B model instead of 16.2 GB.
Go deeper: the engineer version
The kit's question
peft_params.py --model 8b says LoRA needs about 16 GB, but full needs 128. Where does LoRA’s 16 GB go?
The kit's answer
Almost all of it is the frozen bf16 base, 2 bytes × 8B. The adapter’s optimizer state is tiny. Switch to QLoRA and it drops to about 5 GB.
More detail:peft_params.py counts LoRA as 2 B x base parameters (frozen bf16 weights, no gradients or optimizer state) + 16 B x adapter parameters (bf16 weight and gradient, fp32 Adam m and v, fp32 master copy). QLoRA’s 0.5625 B per base weight is 4 bits plus the shared scales of each small block. Activations come on top of every figure and grow with batch x sequence length; --grad-checkpoint trims them.
In today’s runs, full fine-tuning makes nudges 10 times smaller than LoRA’s (a learning rate of 0.00001 against 0.0001). Why?
Show answerHide answer
In plain words
Full fine-tuning moves every number in the model, including the ones that hold its general skills. Big steps across all of them would damage what the model already knows.
Picture it
Restoring a painting: if you may touch every inch, you use a fine brush and light strokes, or you ruin the parts that were already right. Painting on a clear sheet laid over it (LoRA), bolder strokes are safe, because the original underneath is untouched.
With real numbersrun_variants.sh and the Supervised Fine-Tuning video
LoRA and DoRA: learning rate 1e-4 (0.0001). Full: 1e-5 (0.00001), 10 times smaller.
Full moves all 490 million numbers. LoRA at rank 16 (the calculator’s example) moves about 2.2 million (0.44%) and the rest stay frozen. Your training log’s Trainable parameters line gives today’s exact share.
The Supervised Fine-Tuning video gives the same pair: about 1e-4 for LoRA, about 10 times smaller for full.
The damage would show in the general quiz: the untrained model scores about 9 of 12, and a drop below that is usually in the full row.
Words to know
Learning rate
How big a nudge each training step gives each number. Example: 1e-4 (0.0001) for LoRA, 1e-5 (0.00001) for full.
Forgetting
Getting better at the trained task and worse at everything else. Example: a lower general quiz score.
General quiz
Day 4’s 12 quick questions (math, facts, format), reused here to spot forgetting. Example: about 9 of 12 for the untrained model.
Training step
One update: the model sees a few examples and adjusts. Example: 300 steps of 4 tickets.
Go deeper: the engineer version
The kit's question
Why is full fine-tuning’s learning rate 10× smaller?
The kit's answer
Every weight moves, so a big step damages what the model already knows.
More detail: An Adam update moves each weight by roughly the learning rate, so 1e-5 across every weight is still a large total change to the network. Adapter updates are confined to a low-rank correction that starts at zero (B = 0), so a larger rate is safe. Too large a rate in full fine-tuning shows up as a general-quiz drop, the forgetting that claim 3 of the lesson looks for.
Under 20 seconds each. Record yourself once and listen back.
Step 1 · Training toys, part 1
QuestionWe want to tune our model so a scoring model rates its answers higher. Any risk?
One clear answer
A reward model is a proxy for what the customer wants. Optimise it too hard and it stops tracking the goal, so I keep a KL leash, and where correctness can be checked by a program I’d rather use a grader, which is what reinforcement fine-tuning on Fireworks does.
What this means
“A reward model is a proxy for what the customer wants”: This line previews Day 12, where you run the reward-model toy. A reward model is a second model trained to score answers the way people would. It is a stand-in (a proxy) for what the customer really wants, not the real thing.
“Optimise it too hard and it stops tracking the goal”: Push the main model to chase that score too hard and it finds answers the scorer likes that are not really better (reward hacking). In the Day 12 toy, as the model is allowed to drift further from where it started, the score rises about 16% (+1.69 to +1.96) while the answer’s true value falls about 20% (+4.30 to +3.45).
“so I keep a KL leash”: I charge the model for drifting too far from the model it started as, so it cannot wander off to exploit the scorer. The KL penalty measures that drift (KL, short for Kullback-Leibler divergence, is a standard measure of how far two models’ choices differ). A setting called beta (β) sets the leash’s length.
“where correctness can be checked by a program I’d rather use a grader”: When a program can check the answer, such as tests passing or the category matching, I score with that program (a grader) instead of a learned scorer. It checks the thing itself, so there is far less to exploit.
“which is what reinforcement fine-tuning on Fireworks does”: Fireworks’ reinforcement fine-tuning (RFT) trains a model with exactly that kind of program-based grader.
Step 2 · Full vs LoRA vs DoRA
QuestionShould we fine-tune the whole model or use LoRA?
One clear answer
I ran full, LoRA and DoRA on the same data. LoRA got within a couple of points of full at a fraction of the memory, with a 9 MB checkpoint instead of a gigabyte, and without denting general ability. For a narrow task, I start with LoRA and only go full when the eval gap is real.
What this means
“I ran full, LoRA and DoRA on the same data”: I trained the same model three ways on the same 800 tickets with the same settings, so the comparison is fair. Quote your own table’s numbers.
“LoRA got within a couple of points of full”: LoRA’s task accuracy landed within about 2 percentage points of full training. The Full Fine-Tuning video says the same: usually within 1 to 2 points on a narrow task.
“at a fraction of the memory”: About 2 to 3 GB at peak for LoRA against about 8 to 10 GB for full, on the 0.5B model. On an 8B model: 16.2 GB against 128 GB.
“with a 9 MB checkpoint instead of a gigabyte”: The lesson expects the saved LoRA add-on at about 9 MB, and the fully trained model as a whole new copy of about 990 MB. Say your own adapters.safetensors sizes from today’s run. Small files are why one server can hold many customers’ add-ons.
“and without denting general ability”: The 12-question general quiz stayed about where the untrained model was (about 9 of 12): no forgetting.
“For a narrow task, I start with LoRA”: For one well-defined job, like sorting tickets into 4 categories, LoRA is my first choice.
“and only go full when the eval gap is real”: I switch to full fine-tuning only if tests on the customer’s own examples show LoRA falling measurably short.
2 worked examples from today’s videos. Work each on paper, then reveal.
From the video:Full Fine-Tuning4:17
Whiteboard 1 of 2
How much memory does it take to train a model by changing every number in it (full fine-tuning)? Work it out for a 1-billion, an 8-billion and a 70-billion-number model. Then compare LoRA on the 70B.
Given, in plain words
Training keeps about 16 bytes for every number in the model: the number itself (2), its gradient, the direction to nudge it (2), the optimizer’s two running averages (4 + 4) and a precise master copy (4). The working numbers for each group of examples (activations) come on top. One H100 GPU holds 80 GB.
Reveal the answerHide the answer
Answer · in plain words
About 16 GB, 130 GB and 1.1 TB (terabytes, about 1,100 GB). So a 1B model trains on one big GPU, an 8B needs two, and a 70B needs sixteen plus software that splits the job. LoRA on the 70B needs about 160 GB: 2 to 4 GPUs.
Picture it
Moving house: the furniture (the model) is only part of the load. For training you also pack a spare copy of every item, a note on where each should move, and two running tallies of past moves. The truck must be about 8 times bigger than for the furniture alone.
Careful
Memory, not computing speed, decides whether you can train fully. The first fix is often a smaller model: the video notes that a full fine-tune of an 8B can beat a LoRA on a 70B.
Worked answer, step by step
Bytes per number: 2 (the number, 16-bit) + 2 (its gradient) + 4 + 4 (the optimizer’s two running averages, 32-bit) + 4 (the 32-bit master copy) = 16 bytes.
1B model: 1 billion x 16 bytes = 16 GB: one large GPU (an 80 GB H100 has room to spare).
8B model: 8 x 16 = 128 GB, about 130. One 80 GB H100 is not enough; two give 160 GB.
70B model: 70 x 16 = 1,120 GB, about 1.1 TB. 1,120 ÷ 80 = 14 GPUs for that alone. The video plans 16 H100s (1,280 GB), which leaves room for the activations on top, with ZeRO or FSDP splitting the weights, gradients and optimizer state across them.
LoRA on the same 70B: the frozen model at 2 bytes, 70 x 2 = 140 GB, plus a small add-on and its training state: about 160 GB, 2 to 4 H100s. Step 2’s calculator gives 154.5 GB at rank 64 with the MLP grids adapted too (--mlp), before activations.
Compare serving the 70B: 70 x 2 = 140 GB. Full training needs 8 times that (16 ÷ 2).
Go deeper: the engineer version
The kit's question
Worked example · what full fine-tuning needs · ≈16 bytes per parameter, plus activations
The kit's answer
1B model: ≈ 16 GB · one big GPU. 8B model: ≈ 130 GB · 2× H100 80 GB. 70B model: ≈ 1.1 TB · 16× H100 + sharding. LoRA on the same 70B: ≈ 160 GB · 2–4× H100. Memory, not compute, decides whether you can full fine-tune.
More detail: The 16 bytes are mixed-precision Adam: bf16 weights and gradients (2 + 2), Adam’s fp32 first and second moments, its two running averages (4 + 4), and an fp32 master copy (4). Lesson 16’s toy counts 12 bytes per trainable number (the gradient plus the two moments at 4 bytes each), leaving out the weights. Activations grow with batch x sequence length; gradient checkpointing recomputes them (about 30% slower), and 8-bit optimizers cut the bill to about 10 bytes per parameter. ZeRO and FSDP shard weights, gradients and optimizer state, so each GPU holds the total ÷ N.
Words to know
Gradient
The correction worked out for each trainable number at each training step. Example: 2 of the 16 bytes.
Optimizer (Adam)
The rule that nudges each number; Adam keeps two running averages per number. Example: 8 of the 16 bytes.
Activations
The working numbers computed for each example during training, kept for the correction step. Example: they come on top of the 16 bytes.
Sharding (ZeRO, FSDP)
Splitting the weights, gradients and optimizer state across several GPUs. Example: a 70B full fine-tune over 16 H100s.
From the video:Parameter-Efficient Fine-Tuning3:50
Whiteboard 2 of 2
One grid of numbers inside a model (a projection) is 4,096 by 4,096. You add a LoRA add-on of rank 16 beside it. How many numbers does the add-on train, and what share of the grid is that? And across a whole 8B model?
Given, in plain words
LoRA freezes the grid W and trains two thin grids beside it: A, 16 rows by 4,096, and B, 4,096 by 16 (16 is the rank). Multiplying B by A gives a full 4,096 by 4,096 grid of corrections, which is added to W.
Reveal the answerHide the answer
Answer · in plain words
The add-on trains 131,072 numbers, under 1% of the grid’s 16.8 million. Across a whole 8B model, adapting the attention grids, that is a fraction of 1% of all its numbers.
Picture it
The grid is a spreadsheet with 16.8 million cells. Instead of editing cells, you write one 16-column list and one 16-row list; multiplying them fills in a correction for every cell at once. Two thin lists, one big effect.
Why it pays
50 customers with full fine-tunes need 50 files of 16 GB and 50 deployments. 50 LoRA add-ons of about 40 MB each fit on one deployment (the video’s second worked example).
Worked answer, step by step
The full grid: 4,096 x 4,096 = 16,777,216 numbers, about 16.8 million.
A: 16 x 4,096 = 65,536 numbers. B: 4,096 x 16 = 65,536.
The add-on: 65,536 + 65,536 = 131,072 numbers.
Share of the grid: 131,072 ÷ 16,777,216 = 0.78%, about 0.8%.
Whole 8B model, attention only: the video estimates 0.2 to 0.5% of all numbers. Step 2’s calculator, at rank 16 on the four attention grids of all 32 layers, gives 13,631,488 numbers: 0.17%. It comes in lower because two of those grids (k and v) are 4,096 by 1,024, a quarter the size (GQA). Four full-size grids would give 4 x 131,072 x 32 = 16,777,216 numbers, 0.21%.
The rank is the dial: at rank 32, A and B double (262,144 numbers, 1.6% of the grid). More room to learn, more memory.
Go deeper: the engineer version
The kit's question
Worked example · the size of one adapter · one 4096 × 4096 projection · rank r = 16
The kit's answer
full matrix W: 16.8 M. LoRA A (r × 4096): 65,536. LoRA B (4096 × r): 65,536. trainable share: 0.8% of this matrix. whole 8B model, attention only: ≈ 0.2–0.5% of weights. Rank r is the dial: bigger r, more capacity, more memory.
More detail: LoRA adds r·(in + out) parameters per adapted matrix (peft_params.py line 5); the layer computes W·x + (α / r)·B·A·x with W frozen. B starts at zero, so training starts exactly at the pretrained model (lesson 16’s train_lora()). After training, B·A can be merged into W for zero extra latency, or kept as a separate file for multi-LoRA serving. At 2 bytes per number, the Full Fine-Tuning video’s roughly 20 million adapter numbers for an 8B model make its roughly 40 MB file.
Words to know
Rank (LoRA rank)
The add-on’s size setting, the width of its thin grids: how many separate kinds of change it can make at once. Example: 16 here.
Projection (weight grid)
One grid of learned numbers that a layer multiplies its input by. Example: 4,096 x 4,096 in an 8B model’s attention.
Adapter (add-on)
The small set of extra numbers LoRA trains and saves. Example: 131,072 for this one grid.
Multi-LoRA
One server holds one model and many add-ons, applying the right one to each request. Example: 50 customers on one deployment.
How you will use this
When a customer asks “full fine-tune or LoRA?”, you answer with your own table: the memory, file size, accuracy and forgetting of today’s three runs. Day 12 adds the last kind of training, preference tuning: teaching a model which of two answers people prefer.