Skip to content

Day 11 · Training methods in miniature

Part 2 · Day 11 of 12

0 of 12 days done

About 1 h 45 Free Mac

Today: Compare two ways to train a model on your own examples. Change every number in it (full fine-tuning), or freeze it and train a small add-on beside it (LoRA). First on a tiny practice model you can see in full. Then on a real small model on your Mac, measuring what each costs, gains and breaks.

Example

Imagine teaching a support bot to use your company’s ticket labels. One route rewrites every learned number in the model; another leaves the model alone and trains a small add-on. For a large model, the first route needs far more memory than a regular Mac has. Today you will try both routes on a smaller model and see how much memory they use and whether either forgets what it knew before.

By the end you’ll have

One table comparing full fine-tuning, LoRA and DoRA on the same model and the same data: memory, saved file size, task accuracy and forgetting.

Expected: The lesson expects full fine-tuning to peak at about 8 to 10 GB of memory and save a file of about 990 MB. LoRA peaks at about 2 to 3 GB and saves about 9 MB, with task accuracy close to full.

The 6 numbers you write down

  1. Toy error with 200 examples: full / rank 1 / rank 2

    How close each method gets on new examples (lower is better). Rank 2 matches full with 6% of the trained numbers; rank 1 cannot, because the task needs two kinds of change.

    Example: 0.0002 / 0.1321 / 0.0000 (the lesson’s run)

  2. With --true-rank 16, error with 200 examples: full / rank 8

    Now the task needs 16 kinds of change. Even the biggest add-on the script tries (rank 8) stays above full: for a big shift, the rank must grow, or you train fully.

    Example: 0.0002 / 0.0825 (a run of the script with its default seed; the last digits may differ)

  3. Peak memory while training: full / LoRA / DoRA (GB)

    The most of your Mac’s memory each method needed. Full needs several times LoRA’s: the 16-bytes-per-number bill.

    Example: about 8 to 10 / 2 to 3 / 3 (the lesson’s expected shape)

  4. Saved file size: full / LoRA / DoRA (MB)

    What you store, ship and swap. A LoRA file this small is why one server can hold many customers’ add-ons (multi-LoRA).

    Example: about 990 / 9 / 9 (the lesson’s expected files)

  5. Task accuracy: untrained / full / LoRA / DoRA (%)

    What training bought. Expect all three well above the untrained model, with LoRA and DoRA close to full.

    Example: untrained about 40; the three trained rows high and close together

  6. General quiz: untrained / full / LoRA / DoRA (out of 12)

    What training broke, if anything: a drop means forgetting. On 12 questions it is a smoke alarm, not proof.

    Example: untrained about 9; LoRA and DoRA about the same; full may drop

How you will use this: When a customer asks “full fine-tune or LoRA?”, you answer with your own table: the memory, file size, accuracy and forgetting of today’s three runs. Day 12 adds the last kind of training, preference tuning: teaching a model which of two answers people prefer.

Before you start

  1. An Apple-silicon Mac, with Day 1’s Python environment.

    Step 1’s toys run anywhere with Python and numpy (a Python math library), in seconds. Step 2 trains a real model with MLX (Apple’s own software for running and training models on Apple chips), so it needs an M-series Mac. Day 1’s setup installed both. Turn the environment on from the 03-labs folder:

    source .venv/bin/activate

    Every command today runs with it on. Open each new terminal window the same way.

  2. About 5 GB of free disk and 8 to 10 GB of free memory.

    Step 2 downloads a small model, Qwen2.5-0.5B-Instruct: 490 million numbers x 2 bytes = about 1 GB. After full training it saves a whole new copy (about 990 MB). The trainer also keeps in-between copies as it goes, so the full run’s folder can grow to several GB. Full training peaks at about 8 to 10 GB of memory, so close big apps first. On a tight Mac you can skip the full run; step 2’s Stuck? section shows how.

  3. About 45 minutes of steady laptop time.

    Step 2’s three training runs take 15 to 45 minutes together and keep the chip busy. You start them before that step’s video and read on while they run.

About 2 minutes. Say your answer out loud, then tap to check it. From Day 10 · Capstone and rehearsal.

Why must the small draft model share the big model’s tokenizer? (The tokenizer is the part that cuts text into tokens and gives each token a number.)Show answerHide

In plain words

The big model checks the guesses as token numbers, not as text. If the two models cut and number text differently, the same number means different words, and the check means nothing.

Picture it

Two warehouse clerks check an order by product code. That only works if they use the same catalogue: in one, code 4521 is a lamp; in the other, a sofa.

With real numbersthe lesson’s two models, lesson 13

  • Big model: Llama 3.1 8B at 4 bits. Draft: Llama 3.2 1B at 4 bits, from the same family, so it uses the same token numbers.
  • The draft sends up to 8 guesses a round, as token numbers (--draft-max 8).
  • The big model compares each guessed number with the number it would have picked, and keeps the matching run.
  • On Fireworks the lesson uses the same pair, Llama 3.1 8B checking a Llama 3.2 1B draft, with 4 guesses a round.

Words to know

Tokenizer
The part of a model that cuts text into tokens and numbers them. Example: Llama 3.1 8B and Llama 3.2 1B share one.
Token ID
The number a tokenizer gives each token; models compare these numbers, not the text.
Vocabulary
The full list of tokens a tokenizer knows, each with its own number.
Go deeper: the engineer version

The kit's question

Why must the draft share the tokenizer?

The kit's answer

The target verifies token ids, not text.

More detail: Verification compares the draft’s proposed token ids with the target’s own next-token choices position by position (with sampling, it accepts against the target’s probabilities), so both models must share one vocabulary and tokenization. A draft from another family proposes ids that mean different strings to the target. N-gram speculation and predicted outputs avoid the issue, because their drafts come from the target’s own tokens.

Why does speculative decoding help less when the server is busy with many people at once (high concurrency)?Show answerHide

In plain words

With one person, the chip spends most of each step waiting for memory, so it has spare math power to check guesses almost for free. With many people, that spare power already goes to the other people’s tokens, so checking guesses now costs real time.

Picture it

A delivery van on a single drop has lots of empty space, so carrying extra parcels on the off chance costs nothing. Once it is full of other customers’ parcels, every extra parcel pushes out a real one.

With real numbersthe lab book’s DGX Spark measurements and lesson 13’s formula

  • One person: writing reads the whole model for each token, and the math units mostly wait (Day 3: memory-bound).
  • Llama 3.1 8B at 8 bits on a DGX Spark (lab book): 20.5 tokens a second for 1 person, 368 in total for 32.
  • 368 ÷ 20.5 = 18 times the output from the same memory reads: the spare math is now in use.
  • Checking 4 guesses means the math for 5 tokens per person instead of 1, and every rejected guess is wasted math.
  • So speculation pays most for one or a few people who want fast replies.

Words to know

Latency-bound
A workload where what matters is how fast each reply arrives, typical of one or a few people at a time.
Memory-bound
Slowed by how fast memory delivers data, not by math. Example: writing for one person.
Batch size
How many requests share one pass over the model. Example: 1, then 32, on the Spark.
Compute
The chip’s math power. Example: what checking guesses uses.
Go deeper: the engineer version

The kit's question

Why does speculation help less at high concurrency?

The kit's answer

With a full batch the GPU is already busy, so spare compute for verification shrinks. It shines on latency-bound, low-batch workloads.

More detail: At batch 1, decode does about two FLOPs per parameter (1 to 2 per byte fetched at 16 or 8 bits), against an H100 ridge point of about 300 (989 teraflops ÷ 3.35 TB/s, the GPU Bandwidth video), so compute sits idle and verifying k drafted tokens in the same weight read is almost free. As the batch grows, arithmetic intensity rises toward the compute roof; verification adds k + 1 positions per sequence, and rejected positions become wasted compute that displaces other users’ tokens. Speculation shines on latency-bound, low-batch workloads; at high concurrency its gain shrinks and can turn negative.

For editing tasks, where most of the answer repeats the input, what can you use instead of a draft model?Show answerHide

In plain words

Use text you already have as the guess. N-gram speculation finds the last few words it wrote earlier in the prompt and guesses what came next there. Fireworks’ predicted outputs let you send the expected answer, such as the original document, as the draft.

Picture it

Proofreading: instead of an intern drafting from scratch, you hand the editor last week’s version. Most of it passes at a glance, and only the changed lines need real work.

With real numbersfireworks_spec.sh, the lab book and lesson 13’s formula

  • The script’s comment names the option: --ngram-speculation-length=3, no draft model; it reuses n-grams (short runs of tokens) from the prompt.
  • The lab book lists four ways on Fireworks: the default draft model, your own draft model, n-gram, or predicted outputs.
  • Suppose an edit keeps 9 of 10 guesses (α = 0.9, the highest rate in lesson 13’s table; measure it on real edits). With 4 guesses: 1 + 0.9 + 0.81 + 0.73 + 0.66 = 4.1 tokens per big pass.
  • With no draft model to run, guessing costs almost nothing, so the ideal speed-up approaches those 4.1 times. Real gains come in lower.

Words to know

N-gram
A run of n tokens in a row. Example: a 3-gram is 3 tokens.
N-gram speculation
Guessing the next tokens by copying what followed the same words earlier in the prompt; no draft model.
Predicted outputs
A Fireworks option: you send the expected answer as the draft for the model to check.
Go deeper: the engineer version

The kit's question

What is the zero-model alternative for editing tasks?

The kit's answer

N-gram speculation or Fireworks predicted outputs: the draft comes from the prompt itself.

More detail: N-gram (prompt-lookup) speculation matches the last few generated tokens against earlier text in the context and proposes the tokens that followed; predicted outputs let the client send the expected completion, for example the original file for a code edit. Both need no draft model and suit edits, rewrites and extraction, where the output copies long spans of the input. Measure acceptance on the real traffic, as with any drafter.

Start step 1: Training toys, part 1 (30 min)

Your progress is saved on this device.