Skip to content

Day 9 · Scale-out

Part 1 · Day 9 of 12

0 of 12 days done

About 1 h 40 14 GB model download · See Step 1 About half a GPU-hour Mac and Fireworks

Long day. Step 3 is optional and needs a DGX Spark (2 to 3 hours).

Today: Learn why some very large models still write fast, and time one on your Mac. Then let a small draft model guess ahead for a big one, and measure when that pays, on your Mac and on a paid Fireworks deployment. Step 3 is optional: it needs an NVIDIA DGX Spark (a desktop AI computer), and the rest of the course works without it.

Example

Imagine two support bots running on the same Mac. One model file is much bigger, so you might expect its reply to appear more slowly. In the lesson’s sample it actually writes faster: it loads only the specialist parts needed for each piece of the reply. Today you will compare the two writing speeds and find out why file size alone does not predict the customer’s wait.

By the end you’ll have

Your own estimate of how much of a mixture-of-experts model is read per token, plus a table of writing speeds with and without a draft model.

Expected: gpt-oss:20b (the mixture-of-experts model) faster than Llama 3.1 8B, with about a quarter of it read per token (25% in the sample). With a draft model, predictable prompts about 1.9 times faster (62 to 118 tokens a second), creative ones only a little (61.5 to 70).

The 10 numbers you write down

  1. Dense Llama 3.1 8B: writing speed (tok/s)

    How fast a model writes when it reads all of itself for every token.

    Example: 84.0 in the lesson’s sample (a Mac like an M4 Max).

  2. Mixture-of-experts gpt-oss:20b: writing speed (tok/s)

    How fast a model writes when it reads only the experts it picks.

    Example: 118.0 in the sample: faster than the dense model, from a file 2.8 times bigger.

  3. gpt-oss:20b: GB read per token / active %

    Your stopwatch estimate of how much of the model is read for each token.

    Example: 3.5 GB and 25.1% in the sample, against 17.1% published: within 2 times, a good result.

  4. Predictable prompts (structured): tok/s without / with speculation

    How much speculation speeds up easy-to-guess output on your Mac.

    Example: 62.0 / 118.0 in the lesson’s sample: 1.9 times.

  5. Creative prompts (prose): tok/s without / with speculation

    How much speculation helps when the next words are hard to guess.

    Example: 61.5 / 70.0 in the sample: 1.14 times.

  6. Acceptance rate α: structured / prose

    The number that decides whether speculation pays.

    Example: 0.780 / 0.410 in the sample.

  7. Fireworks with speculation: tok/s structured / prose

    The same test on a Fireworks deployment with a 1B draft model guessing 4 tokens a round.

    Example: The lesson has no sample for this run. Expect the structured row faster than the prose row.

  8. vLLM on the Spark: maximum concurrency at 8,192 tokens (x)

    How many full-length conversations fit in vLLM’s memory for notes at once.

    Example: The lesson prints no sample. The kit’s numbers give at most (0.8 × 128 − 8) ÷ 1.07 = about 88; expect fewer after vLLM’s own working memory. The Serving With vLLM video’s log, from another machine, shows about 16.

  9. Spark sweep: people at once before the first-token wait passes 1 second / total tok/s at that point

    How many people one Spark serves while 95 of 100 still get their first token within a second.

    Example: The lesson prints no sample. For scale, the lab book’s SGLang figure for this model: 368 tokens a second in total at 32 people.

  10. SGLang: cached_tokens / prompt_tokens

    How much of the prompt SGLang reused instead of reading again.

    Example: cached_tokens close to prompt_tokens on the same line: nearly the whole opening (the script estimates about 3,028 tokens) was reused.

How you will use this: Two of Day 10’s 20-second drill questions come from today: “Why is the MoE faster than the smaller dense model?” and “Will speculative decoding help us?”. The sizing memo also quotes your speculation speeds from results/13-spec.csv.

Before you start

  1. About 14 GB of free disk space. A Mac with 24 GB of memory or more (16 GB: use the sample numbers).

    Step 1 downloads gpt-oss:20b, a 13.8 GB model. Llama 3.1 8B (4.9 GB) is already there from Day 1, so only this download is new. Day 1’s rule keeps a model under about 70% of memory. On a 16 GB Mac that is 0.7 × 16 = 11.2 GB, too little, so use the lesson’s sample numbers there. Start the download now.

    From the 03-labs folder:

    ollama pull gpt-oss:20b

    Ollama must be running. If you restarted your Mac since Day 1, start it first with ollama serve &.

  2. llama.cpp from Day 2, with port 8080 free.

    Step 2 serves Day 2’s 4-bit Llama 3.1 8B again, now with a small draft model (Llama 3.2 1B) that it downloads on first start. Stop any llama.cpp server still running from Days 2, 4 or 5 first: they all use port 8080.

  3. Your Fireworks key, spend cap and firectl from Day 1.

    Step 2 ends on a paid Fireworks deployment for about 30 to 45 minutes. The script deletes it when it finishes, even if you press Ctrl+C. At the lab book’s H100 price of $8 an hour, 30 to 45 minutes costs about $4 to $6.

  4. Optional step 3 only: a DGX Spark you can reach from your Mac.

    Step 3 takes 2 to 3 hours and needs an NVIDIA DGX Spark (a desktop AI computer) or a similar NVIDIA GPU machine. Put its network address in 03-labs/.env as SPARK_IP, set up SSH keys, and add a Hugging Face token as HF_TOKEN for models that need one. No Spark? Read step 3, answer its checks, then tap Skip.

About 2 minutes. Say your answer out loud, then tap to check it. From Day 8 · Training day, local.

Day 8’s lesson keeps two sets of tickets out of training: 100 validation tickets, checked during training, and 100 test tickets. Why is the final score taken on the test tickets (test.jsonl) and not the validation ones (valid.jsonl)?Show answerHide

In plain words

You used the validation tickets to decide when to stop training, so they helped pick the model and can flatter it. The test tickets played no part in any decision, so their score is honest.

Picture it

A coach decides when the team has practised enough by watching its scores on one practice test. A good score on that same test afterwards proves little. The fair check is a fresh test nobody has looked at.

With real numbersthe kit’s ticket data (data/make_tickets.py) and Day 8’s training script

  • 1,000 tickets, split three ways: 800 to train on, 100 to check progress during training (validation), 100 kept back for the final grade (test).
  • Training checks the validation tickets every 50 steps, and you stop where their loss flattens: in the sample, 0.198 at step 100 and 0.121 at step 300.
  • That stopping point was chosen by looking at the validation tickets, so they are no longer neutral.
  • The test tickets never touch training or the stop decision. No test ticket repeats a training ticket, and 30 of the 100 use wordings training never saw.
  • The final grade uses the first 60 test tickets, so each one is worth 1.67 points (100 ÷ 60).

Words to know

Validation set
Examples kept out of training and checked during it, to decide when to stop. Example: valid.jsonl, 100 tickets.
Test set (held-out set)
Examples kept out of training and out of every decision, used only for the final score. Example: test.jsonl, 100 tickets; the eval uses 60.
Validation loss
A score of how wrong the model still is on the validation set; lower is better. Example: 0.412, 0.198 and 0.121 at steps 50, 100 and 300.
Leak (data leakage)
Data used to judge a model has also shaped it, so the score looks better than it is. Example: the validation tickets, once they set the stopping point.
Go deeper: the engineer version

The kit's question

Why do we evaluate on test.jsonl and not valid.jsonl?

The kit's answer

Validation loss guided when to stop, so it has leaked into the decision. Test is untouched.

More detail: Early stopping on validation loss is a form of model selection, so the validation score is optimistic; the test split stays untouched for the final number. make_tickets.py de-duplicates across all splits and keeps one phrasing per category for the test set only, so the test score measures generalisation, not memory. One caution: the lab book’s standalone snippet builds its validation file with head -50 train.jsonl after copying all of train.jsonl into training, so those 50 rows are also training rows and its validation loss would look better than it is. Day 8’s script uses the separate data/tickets/valid.jsonl.

When would you tell a customer to fine-tune on their own machines (a laptop, or AI chips called GPUs that they own) rather than on a managed platform such as Fireworks?Show answerHide

In plain words

When their data is not allowed to leave their own machines. Or when they want fast, free experiments first, before paying a managed platform to train the final version and serve it to real users.

Picture it

A chef tests new recipes in their own kitchen: it costs nothing, and the secret sauce stays secret. For a 500-guest wedding, they hire a caterer with the staff and ovens to serve everyone at once.

With real numbersDay 8’s lesson, Day 7’s Fireworks run (lesson sample figures and its summary line), and the lab book’s GPU price

  • On the Mac: $0 to train, and the 800 tickets never leave the machine. Training takes about 15 to 40 minutes.
  • But one Mac serves one person at a time with no uptime promise: COMPARISON.md answers “no” for 100 users at once.
  • On Fireworks (Day 7’s sample): training cost cents to about a dollar, but answering requests (serving) needs a GPU reserved by the hour. 38 minutes cost $5.07: $5.07 ÷ 38 x 60 = about $8 an hour, the lab book’s price.
  • Accuracy is close either way in the samples: category 78% to 96% on the Mac, 80% to 97% on Fireworks.
  • The usual pattern, from COMPARISON.md: try ideas on your own machine first (prototype locally), then run the real service on a managed platform (ship managed).

Words to know

Managed platform
A service that runs everything between the GPUs and the customer’s app: the engine, scaling, routing and the models. Example: Fireworks on Day 7.
Endpoint
A web address that answers requests for one model. Example: http://localhost:8081/v1 for Day 8’s fine-tune.
Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: about $8 an hour.
SLA (service level agreement)
A written promise to a customer about uptime or speed. Example: COMPARISON.md says the Mac has none.
Go deeper: the engineer version

The kit's question

When would you tell a customer to fine-tune locally or on their own GPUs?

The kit's answer

When data can’t leave their environment, or for fast iteration before a managed production run.

More detail: Local or self-hosted training fits data that must stay inside the customer’s environment, and cheap, fast iteration. Managed wins on time to a scalable endpoint, autoscaling and an SLA. On Fireworks a LoRA cannot run on serverless (shared, pay-per-token) models; it needs a dedicated deployment, so the serving hours, not the training, are the cost to plan (the lab book). Day 7’s lesson calls training “a few cents” in its sample and “about a dollar” in its summary line; the QLoRA On Your Own Box video says about a dollar. If data residency is the only blocker, the lab book notes that Fireworks can also run inside the customer’s own cloud account (BYOC), or pin a dedicated deployment to a region at 1.5 times the rate. The QLoRA On Your Own Box video puts it this way: “the local versus managed trade-off is a conversation, not a religion.”

The add-on you trained (the adapter) is about 7 MB. Why does that small size matter when the model runs as a service?Show answerHide

In plain words

One copy of the base model can keep hundreds of these small add-ons in memory and apply the right one to each request. So every customer can have their own tuned model for about the price of one shared server.

Picture it

One base recipe, a different spice packet per customer (the field guide’s comparison). The kitchen cooks one big pot and adds a small packet to each order, instead of running a separate kitchen for every customer.

With real numbersDay 7’s lesson (12 business units), Day 8’s sample run and memory script, and the lab book’s $8-an-hour GPU price

  • Adapter: 7.1M in Day 8’s sample, about 7.4 MB. Base model: about 4.3 GB, which is 4,300 MB. 4,300 ÷ 7.4 = about 580, so the base is hundreds of times bigger.
  • 100 adapters: 100 x 7.4 MB = 740 MB, about a sixth of one base model (740 ÷ 4,300 = 0.17).
  • Day 7’s example customer wants 12 tuned models, one per business unit (department). Each on its own dedicated GPU at $8 an hour: 12 x $8 = $96 an hour.
  • The same 12 as adapters on one shared base: about $8 an hour, if one deployment can carry all their traffic.
  • Running all month (730 hours): $8 x 730 = $5,840 for one deployment, against 12 x $5,840 = $70,080 for twelve.

Words to know

Adapter (LoRA adapter)
The small set of extra numbers LoRA trains and saves; the base model stays unchanged. Example: about 7 MB in Day 8’s sample.
Base model
The model a fine-tune starts from and leaves unchanged. Example: Qwen2.5 7B Instruct at 4 bits, about 4.3 GB.
Multi-LoRA
One server keeps one base model and many adapters, applying the right adapter to each request. Example: 12 business units on one deployment.
Tenant
One customer, or business unit, sharing a platform with others. Example: per-tenant fine-tunes, one adapter each.
Go deeper: the engineer version

The kit's question

The adapter is 7 MB. Why does that matter for serving?

The kit's answer

One base model can hold hundreds of adapters in memory, which makes per-tenant fine-tunes affordable.

More detail: Multi-LoRA servers keep one base in GPU memory and apply the right adapter per request, batching requests for different adapters together; the field guide says Fireworks serves hundreds of fine-tunes on one base model at the base model’s price. Fusing the adapter into the weights, as Day 8’s 2_fuse_and_serve.sh does for simplicity, gives that up: each fused model is a full copy. The size is the script’s estimate: 8 layers x 4 grids x 2 x 3,584 x rank 8 = 1,835,008 numbers, 3.7 MB at 2 bytes each. It treats all four grids as 3,584 x 3,584; Qwen2.5 7B’s key and value grids are 3,584 x 512 (lesson 17’s peft_params.py), which gives 1,441,792, and mlx-lm’s own defaults decide which grids get an add-on. Lesson 11’s README says about 4 MB, its sample prints 7.1M, and it does not say why they differ. du -sh adapters measures the whole folder, which also keeps copies saved during training, so quote your own ls -lh adapters/adapters.safetensors figure.

Start step 1: MoE vs dense (40 min)

Your progress is saved on this device.