Skip to content

Training toys, part 1

course 19 of 22

lesson 16 · lessons/16-training-toy

30 min Free Anywhere

What you will do and why

Two ways to train, shrunk to one small layer (a single grid of 4,096 numbers) so you see every number and finish in seconds. Change all of them, or train a small add-on (LoRA) beside them, and see when the add-on is too small.

Why it matters: Imagine teaching a tiny model a new ticket-label rule. Changing all its numbers works, but you want to see whether a smaller add-on can learn it too. In the lesson’s toy, an add-on that changes 256 numbers learns the rule while a smaller one does not. The point is to test whether the add-on has enough room to learn, not to assume the smallest setting always works.

You are done when: python toy_finetune.py has printed its table, and you can say why rank 1 stays stuck at 200 examples (0.1321 in the lesson’s run) while rank 2 matches full fine-tuning.

Full Fine-Tuning · 4:17

Download mp4 (15.6 MB)

Chapters

In this video Why changing every number in a model is the most powerful and the most expensive way to train it, and when it is worth it.

3 key points

  1. Training needs about 16 bytes of memory per number in the model; running it needs about 2.

    Each number brings its gradient (the direction to nudge it), two running averages and a precise master copy. A 7-billion-number model runs in about 14 GB (7 x 2) but needs about 112 GB (7 x 16) to train: 8 times as much.

  2. Memory, not computing speed, decides whether you can train fully at all.

    An 8-billion-number (8B) model needs about 130 GB (8 x 16 = 128), so two 80 GB H100s (NVIDIA’s data-center GPU). A 70B needs about 1.1 TB (terabytes: 70 x 16 = 1,120 GB): sixteen H100s plus software that splits the job across them.

  3. On a narrow task, LoRA usually lands within 1 to 2 accuracy points of full training, and forgets less of what the model already knew.

    Same 8B model: full training changes 8 billion numbers in about 130 GB and saves a 16 GB file. LoRA trains about 20 million in about 20 GB and saves about 40 MB. Go full for a big shift, such as a new language, with hundreds of thousands of examples.

Parameter-Efficient Fine-Tuning · 3:50

Download mp4 (14.3 MB)

Chapters

In this video How LoRA customises a huge model by training two thin grids of numbers beside each frozen one, and when to pick its relatives: QLoRA, DoRA and multi-LoRA.

3 key points

  1. LoRA freezes the model and trains two thin grids, A and B. Multiplied together, they make a full-size correction that is added to the frozen grid.

    For one 4,096 x 4,096 grid (16.8 million numbers) at rank 16 (the add-on’s size setting), A and B hold 65,536 numbers each: 131,072 in all, 0.8% of that grid.

  2. The rank is the dial: a higher rank gives the add-on more room to learn, and costs more memory.

    Rank 8 to 16 suits narrow tasks, 32 to 64 harder ones. Across a whole 8B model, adapting the attention grids, the video puts it at 0.2 to 0.5% of all numbers. Step 2’s calculator gives 0.17% at rank 16, because two of Llama 3.1 8B’s four attention grids are a quarter size (a design called GQA). Four full-size grids would give 0.21%.

  3. Pick QLoRA when memory is short, DoRA when quality lags, multi-LoRA when you have many customers.

    QLoRA is LoRA on a 4-bit copy of the model; DoRA is a LoRA variant that often scores 1 to 2 points higher. Fifty customers with full fine-tunes need 50 files of 16 GB (800 GB) and 50 deployments. Fifty LoRA add-ons of about 40 MB (2 GB in all) fit on one deployment, the right one applied per request.

Real training takes minutes to hours. This toy shrinks the model to one small layer of 4,096 numbers and trains it in seconds. You compare changing all of them (full fine-tuning) with training a small add-on beside them (LoRA). Part 2, on Day 12, uses a toy for preference tuning (teaching a model which of two answers people prefer).

Picture it

Full fine-tuning re-edits every page of an encyclopedia. LoRA leaves the pages alone and adds sticky notes. The rank is how many notes you may use: if the book needs two different kinds of fix, one note cannot hold both, however many examples you study.

With real numberslesson 16’s toy_finetune.py, the lesson’s sample run

  • The toy layer: a grid of 64 x 64 = 4,096 numbers. Full fine-tuning trains all 4,096.
  • LoRA at rank 2: two thin grids, 64 x 2 and 2 x 64, so 128 + 128 = 256 trained numbers, 6.2% of the layer.
  • The script counts 12 bytes per trained number: its gradient and the optimizer’s two running averages, at 4 bytes each in a real trainer. Full: 4,096 x 12 = 49,152 bytes (48 KB, far smaller than one phone photo). Rank 2: 256 x 12 = 3,072 bytes, 16 times less.
  • With 200 examples, rank 2 reaches an error of 0.0000 (0 is perfect), as good as full (0.0002). Rank 1 stays at 0.1321. With only 16 examples every method sits near 0.4 (0.36 to 0.40).
  • Scale the toy’s bill to an 8-billion-number model: 8 billion x 12 bytes = 96 GB. The Full Fine-Tuning video’s fuller count is 16 bytes per number: the number itself (2), its gradient (2, stored smaller than the toy counts it), the two averages (4 + 4) and a precise master copy (4). 8 billion x 16 = 128 GB, 8 times a 16 GB Mac.

Words to know

Full fine-tuning
Training that changes every number in the model. Example: all 4,096 numbers of the toy layer.
LoRA
Freeze the model and train two thin grids, A and B; multiplied together, they make a correction added to a layer. Example: 256 trained numbers at rank 2.
Rank
The add-on’s size setting, the width of its two thin grids: how many separate kinds of change it can make at once. Example: the toy’s task needs 2 kinds; a rank-1 add-on can make only 1.
Adam (the optimizer)
The usual rule for nudging each trained number; it keeps two running averages for every one. Example: adam() in toy_finetune.py.

From the lesson

lessons/16-training-toy/README.md

Real training runs take minutes to hours, which makes it hard to see what an algorithm does. These two scripts shrink the problem until every number fits on one screen, while keeping the real algorithms intact.

script shrinks lets you see
toy_finetune.py (part 1, today) a model → one 64×64 layer full fine-tuning vs LoRA ranks: trainable parameters, optimizer memory, held-out error as the data grows
toy_alignment.py (part 2, Day 12) all possible text → 6 answers to one ticket base → SFT → reward model → RLHF (β sweep and a sampled policy-gradient run) → DPO, including reward hacking
  • toy_finetune.py → train_lora(): W = W0 + B @ A with W0 frozen. The gradient reaches B and A through the chain rule, and those are the only trainables. B starts at zero, so training starts exactly at the pretrained model.
  • toy_finetune.py → adam(): look at what it stores for every trainable number. That storage is the memory bill from the Full Fine-Tuning video in this step.
Terminal window
cd lessons/16-training-toy
python toy_finetune.py # ~20 s
python toy_finetune.py --true-rank 16 # a bigger shift: small ranks can't keep up

What each command does

  1. python toy_finetune.py

    Trains one 64 x 64 layer five ways: full fine-tuning, then LoRA at rank 1, 2, 4 and 8, each on 16, 48 and 200 examples. It needs only numpy, no GPU, and takes about 20 seconds. Look for rank 1 stuck well above the others in the n=200 column (0.1321 in the lesson’s run), while rank 2 matches full (0.0000 against 0.0002).

  2. python toy_finetune.py --true-rank 16

    The same test, but the change the task needs now has 16 separate kinds instead of 2. Look for every LoRA row, even rank 8, staying above full fine-tuning in the n=200 column: rank 8 at about 0.08 against full’s 0.0002. An add-on smaller than the change the task needs cannot keep up, however much data it gets.

How to read it

Each row is one way to train the layer. trainable is how many numbers it changes; adam_bytes is its training memory at 12 bytes each. The n= columns are its error on 500 new examples after training on 16, 48 or 200 (lower is better, 0 is perfect). Your run also prints rank 4 and rank 8, which match full at 200 like rank 2. With only 16 examples every row does badly (0.36 to 0.40): fewer examples than the layer is wide (64) cannot pin the change down.

method trainable share adam_bytes n=16 n=48 n=200
full 4,096 100.0% 49,152 0.3600 0.1084 0.0002
lora r=1 128 3.1% 1,536 0.3937 0.1951 0.1321 ← rank too small: plateaus
lora r=2 256 6.2% 3,072 0.4024 0.0996 0.0000 ← = full, at 6% of the trainables

Part 2 of this lesson, toy_alignment.py, prints a table of its own. You run it on Day 12, step 1 (Training toys, part 2).

1 question. Say your answer out loud, then tap to check it.

In Day 11’s toy (one small layer of 4,096 numbers), the rank-1 add-on never catches up with full fine-tuning, even with 200 examples. Why? (The rank is how many separate kinds of change the add-on can make at once.)Show answerHide

In plain words

The change this task needs has two separate kinds (two directions), and a rank-1 add-on can make only one. More examples cannot supply the missing one; only a bigger rank can.

Picture it

Treasure lies to the north-east, and you may only move along a straight track running east. However many tries you get, you stop at the closest point on the track, never at the treasure. Add a second track running north and you reach it exactly.

With real numberslesson 16’s toy_finetune.py, the lesson’s sample run

  • The task needs a rank-2 change: two separate kinds of change (the script’s --true-rank is 2 by default).
  • Rank 1: 64 + 64 = 128 trained numbers, 3.1% of the layer. Error 0.3937, 0.1951 and 0.1321 at 16, 48 and 200 examples (0 is perfect).
  • Rank 2: 256 trained numbers, 6.2%. Error at 200 examples: 0.0000, as good as full fine-tuning (0.0002).
  • So rank 1 trains half as many numbers as rank 2 but misses one whole kind of change: at 200 examples its error is still 0.1321, while rank 2’s is 0.0000.

Words to know

Rank
How many separate kinds of change an add-on can make at once. Example: the toy’s task needs 2.
Held-out error
How wrong the trained layer is on new examples it never saw; lower is better. Example: 0.1321 for rank 1 at 200 examples.
Adapter (add-on)
The small set of extra numbers LoRA trains while the model stays frozen. Example: 128 numbers at rank 1.
Frozen
Left unchanged during training. Example: the toy’s pretrained layer W0.
Go deeper: the engineer version

The kit's question

Why does LoRA r=1 never catch up, even with 200 examples?

The kit's answer

The change the task needs is rank 2, and a rank-1 adapter cannot represent it.

More detail: LoRA’s update is B·A with B of shape 64 x r and A of shape r x 64, so its rank is at most r. The toy’s target shift dW_true is built from two random directions, so it has rank 2 (toy_finetune.py line 38). A rank-1 B·A can capture at best the strongest single direction of dW_true (its top singular direction); the remaining direction stays as error however many examples you add. With --true-rank 16 every rank the script tries (up to 8) falls short the same way.

Question We want to tune our model so a scoring model rates its answers higher. Any risk?

One clear answer

A reward model is a proxy for what the customer wants. Optimise it too hard and it stops tracking the goal, so I keep a KL leash, and where correctness can be checked by a program I’d rather use a grader, which is what reinforcement fine-tuning on Fireworks does.

What this means

  • “A reward model is a proxy for what the customer wants”: This line previews Day 12, where you run the reward-model toy. A reward model is a second model trained to score answers the way people would. It is a stand-in (a proxy) for what the customer really wants, not the real thing.
  • “Optimise it too hard and it stops tracking the goal”: Push the main model to chase that score too hard and it finds answers the scorer likes that are not really better (reward hacking). In the Day 12 toy, as the model is allowed to drift further from where it started, the score rises about 16% (+1.69 to +1.96) while the answer’s true value falls about 20% (+4.30 to +3.45).
  • “so I keep a KL leash”: I charge the model for drifting too far from the model it started as, so it cannot wander off to exploit the scorer. The KL penalty measures that drift (KL, short for Kullback-Leibler divergence, is a standard measure of how far two models’ choices differ). A setting called beta (β) sets the leash’s length.
  • “where correctness can be checked by a program I’d rather use a grader”: When a program can check the answer, such as tests passing or the category matching, I score with that program (a grader) instead of a learned scorer. It checks the thing itself, so there is far less to exploit.
  • “which is what reinforcement fine-tuning on Fireworks does”: Fireworks’ reinforcement fine-tuning (RFT) trains a model with exactly that kind of program-based grader.

Your numbersSaved on this device and collected in the Day 11 wrap-up.

Hint: The n=200 column of python toy_finetune.py, rows full, lora r=1 and lora r=2. Write all three, for example “0.0002 / 0.1321 / 0.0000”.

Hint: The n=200 column of python toy_finetune.py --true-rank 16, rows full and lora r=8.

python toy_finetune.py has printed its table, and you can say why rank 1 stays stuck at 200 examples (0.1321 in the lesson’s run) while rank 2 matches full fine-tuning.

ModuleNotFoundError: No module named 'numpy'
The Python environment is off in this terminal. From the 03-labs folder run source .venv/bin/activate and try again. The toys need nothing else, so on another machine pip install numpy is enough.
Your numbers differ a little from the lesson’s.
The toy draws its random numbers from a fixed starting point (--seed 0), so the shape should match: rank 1 stuck, rank 2 and above matching full at 200 examples. Small differences in the last digits do not matter. Run python toy_finetune.py --seed 1 to see the same shape with other random numbers.
toy_finetune.py Python · 115 lines
"""
Lesson 16 · part 1 — full fine-tuning vs LoRA on ONE layer, in numpy (seconds, no GPU).
Setup
A "pretrained" layer W0 (64 × 64) already does a general job well.
The new task needs W0 + ΔW, where ΔW is low-rank (rank 2) — like most real fine-tunes.
We try it with 16, 48 and 200 training examples for the new task.
We train it five ways with plain gradient descent:
full update all 4,096 weights of W (full fine-tuning)
lora r=1…8 freeze W0, train B (64×r) and A (r×64); W = W0 + B·A (LoRA)
and report, for each:
trainable how many numbers the optimizer must track
adam_bytes optimizer + gradient memory at 12 bytes per trainable param
task_err error on NEW held-out task examples (did it learn the task, or just memorise?)
python toy_finetune.py
python toy_finetune.py --true-rank 16 # a bigger shift: now small ranks can't keep up
"""
import argparse
import numpy as np
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--d", type=int, default=64, help="layer width")
ap.add_argument("--true-rank", type=int, default=2, help="rank of the change the task really needs")
ap.add_argument("--n-train", default="16,48,200", help="comma list of training-set sizes")
ap.add_argument("--steps", type=int, default=5000)
ap.add_argument("--seed", type=int, default=0)
a = ap.parse_args()
rng = np.random.default_rng(a.seed)
d = a.d
# ── the world ────────────────────────────────────────────────────────────────
W0 = rng.normal(0, 1 / np.sqrt(d), (d, d)) # "pretrained" weights
U, V = rng.normal(0, 1, (d, a.true_rank)), rng.normal(0, 1, (a.true_rank, d))
dW_true = 0.6 * U @ V / np.sqrt(d * a.true_rank) # the low-rank shift the task needs
W_task = W0 + dW_true
X_test = rng.normal(0, 1, (500, d)) # held-out task inputs
def mse(W, X, Wref):
return float(np.mean((X @ W.T - X @ Wref.T) ** 2))
def adam(params, grads, state, lr=1e-2, b1=0.9, b2=0.999, eps=1e-8):
"""Plain Adam. Note what it stores: m and v for EVERY trainable number — that is the memory bill."""
state["t"] = state.get("t", 0) + 1
for k in params:
m = state.setdefault("m_" + k, np.zeros_like(params[k]))
v = state.setdefault("v_" + k, np.zeros_like(params[k]))
m[:] = b1 * m + (1 - b1) * grads[k]
v[:] = b2 * v + (1 - b2) * grads[k] ** 2
mh, vh = m / (1 - b1 ** state["t"]), v / (1 - b2 ** state["t"])
params[k] -= lr * mh / (np.sqrt(vh) + eps)
def train_full(X_train, Y_train):
P = {"W": W0.copy()}
st = {}
for _ in range(a.steps):
err = X_train @ P["W"].T - Y_train # forward + residual
G = 2 * err.T @ X_train / len(X_train) # dLoss/dW, same shape as W
adam(P, {"W": G}, st)
return P["W"], P["W"].size
def train_lora(r, X_train, Y_train):
# Standard LoRA init: A small random, B zero → training starts exactly at W0.
P = {"A": rng.normal(0, 1 / np.sqrt(d), (r, d)), "B": np.zeros((d, r))}
st = {}
for _ in range(a.steps):
W = W0 + P["B"] @ P["A"] # W0 is frozen; only B·A changes
err = X_train @ W.T - Y_train
G = 2 * err.T @ X_train / len(X_train) # gradient w.r.t. the effective W …
adam(P, {"B": G @ P["A"].T, "A": P["B"].T @ G}, st) # … chained into B and A (the only trainables)
return W0 + P["B"] @ P["A"], P["A"].size + P["B"].size
sizes = [int(x) for x in a.n_train.split(",")]
methods = ["full"] + [f"lora r={r}" for r in (1, 2, 4, 8)]
err = {m: {} for m in methods}
count = {}
for n in sizes:
X_train = rng.normal(0, 1, (n, d)) # n task examples
Y_train = X_train @ W_task.T + rng.normal(0, 0.02, (n, d)) # labels, a little noise
W, count["full"] = train_full(X_train, Y_train)
err["full"][n] = mse(W, X_test, W_task)
for r in (1, 2, 4, 8):
W, count[f"lora r={r}"] = train_lora(r, X_train, Y_train)
err[f"lora r={r}"][n] = mse(W, X_test, W_task)
print(f"layer {d}×{d} · the task needs a rank-{a.true_rank} change · held-out task error (lower is better)\n")
head = f"{'method':10s} {'trainable':>10s} {'share':>7s} {'adam_bytes':>11s} " + " ".join(f"{'n=' + str(n):>9s}" for n in sizes)
print(head)
print("-" * len(head))
for m in methods:
n_t = count[m]
print(f"{m:10s} {n_t:10,d} {100 * n_t / W0.size:6.1f}% {n_t * 12:11,d} " + " ".join(f"{err[m][n]:9.4f}" for n in sizes))
print(f"""
How to read it
• Enough data (right-hand column): LoRA with r ≥ {a.true_rank} (the true rank of the change) matches full
fine-tuning, while training only a few percent of the numbers.
• r below the true rank plateaus however much data you add: rank is the capacity dial (video 16).
• Little data (left-hand column): nothing generalises well. With fewer examples than the layer
is wide, no method can pin down the whole change. More clean data beats a cleverer method
(video 15).
• adam_bytes is the gradient and optimizer bill at 12 bytes per trainable number. Scale 4,096
up to 8 billion and you have the memory wall of full fine-tuning (video 14). Scale it to
20 million and you have LoRA's.
• A linear toy can't show "forgetting" of general skills; lesson 17 checks that on a real model.
""")