Skip to content

Day 10 · Capstone and rehearsal

Part 1 · Day 10 of 12

0 of 12 days done

About 2 h 40 Free Anywhere

Today: Turn what you measured on Days 1 to 9 into a sizing memo: the written advice a field engineer sends a customer after the first call, on how to run their work and what each option costs. Build a 10-minute, five-slide talk from it and give it out loud once. Then write one specific criticism of Fireworks, backed by something you hit yourself.

Example

Imagine a customer asks whether to pay for each AI request or rent a chip that runs all month. For the traffic in the video’s example, paying per request is about $146 a month; keeping one chip reserved is about $5,840. That bill alone makes the first option easier to justify at this traffic level. Today you will use your own speed and capacity measurements to write a recommendation a customer can act on.

By the end you’ll have

results/SIZING_MEMO.md, five slides built from it, and one product criticism with a number you measured and a source, with the date you read it.

Expected: Paying per token costs the memo’s example customer $484 a month, or $282 with caching (reusing the repeated opening of each prompt), whatever you measured; the GPU bill depends on how many people one copy served in your Day 3 test.

The 10 numbers you write down

  1. People one copy holds within the speed promise, and copies needed at the busiest moment

    Your Day 3 measurement drives everything below it: fewer people per copy means more copies and a bigger dedicated bill.

    Example: 8 / 18 (the practice server, the kit’s sample: 120 ÷ 8 x 1.2 = 18)

  2. Monthly bill paying per token, without and with caching

    What the customer pays with no hardware of its own. The memo assumes 90% of each prompt is cached, at half price by default.

    Example: $484 / $282 (the example customer)

  3. Monthly bill on dedicated GPUs

    A fixed bill, paid whether the GPUs are busy or idle.

    Example: $105,120 (18 x $8 x 730)

  4. Crossover: tickets a month at which dedicated costs the same as serverless with caching

    Below it, serverless is cheaper. The figure keeps today’s copies fixed, so re-size the copies before you quote a volume near it. A fine-tune or a strict speed promise can need dedicated GPUs at any volume.

    Example: about 744 million (744,475,921), 372 times the example’s 2 million

  5. How long your talk took, the first time out loud

    One timed run shows where you ramble and where a number is missing.

    Example: 10

  6. What you hit, in one sentence

    The criticism itself, specific enough that someone at Fireworks could reproduce it.

    Example: Training my LoRA cost cents, but serving it needed a dedicated deployment billed by the hour (Day 7).

  7. Your evidence: numbers and sources

    What turns an opinion into a finding.

    Example: About $0.15 to train by the script’s estimate (0.3 million tokens x $0.50); the eval deployment ran 38 minutes, $5.07. LoRA docs: dedicated only. Pricing page: fine-tuned models at base-model prices. The lab book read both on 23 September 2026; read them again and write your own date.

  8. Who it hurts, and how much

    Why the product team should care, in the customer’s money or time.

    Example: A customer with one fine-tune and light traffic, 20 million tokens a day: about $122 a month at the base model’s per-token price, but up to $5,840 a month to keep its dedicated H100 ready around the clock, 48 times more.

  9. What you would change

    Turns the criticism into a product proposal: the lab book’s “turn recurring pain into product proposals”.

    Example: Per-token serving for LoRAs on the most-used base models. First, make the pricing page and the LoRA docs agree, and show the hourly cost when a deployment starts.

  10. Your short version, to say out loud

    What you say when they ask what you would change about Fireworks.

    Example: Fine-tuning on Fireworks was easy: my LoRA trained for about 15 cents. Serving it cost $5 for a 38-minute eval, because trained LoRAs run only on a dedicated GPU, $5,840 a month if kept ready. I’d propose per-token serving for LoRAs on popular base models. How do field engineers get that onto the roadmap?

How you will use this: A customer can use your sizing memo to choose how to run the system and what to budget. Its five slides help you explain the recommendation to a team. The memo draws on your own speed benchmark (Days 1 and 3) and model-improvement results (Days 7 and 8); the product criticism records a real obstacle you found.

Before you start

  1. Your results from Days 1 to 9.

    The memo reads six files in 03-labs/results/: the Day 1 and 2 timing table, Day 3’s concurrency sweep, Day 5’s caching test, Day 6’s valid-reply rate, the Day 7 and 8 eval table and Day 9’s speculation test. Day 1’s smoke test also saved practice-server rows in all six. Once the memo is built, check that it quotes your own runs, not rows marked mock. From the 03-labs folder:

    ls results/*.csv

    Look for 01-latency.csv, 03-sweep.csv, 06-prefix.csv, 07-structured.csv, 09-eval.csv and 13-spec.csv. A missing file shows up in the memo as TODO, never as a guess.

  2. Your numbers from the earlier wrap-ups.

    Each day’s wrap-up saved the numbers you typed, on this device. Open the wrap-up pages for Days 1, 3, 5 and 7 and press Copy all on each to gather what your slides need: the Day 1 timings, the Day 3 capacity, the Day 5 caching speed-up and the Day 7 eval table.

  3. Nothing still billing on Fireworks.

    Day 9 ran a paid deployment. Today is free, so first confirm nothing was left running. From the 03-labs folder:

    make fw-check

    It should list no deployments. If one is there, delete it with firectl deployment delete followed by its id.

About 2 minutes. Say your answer out loud, then tap to check it. From Day 9 · Scale-out.

A customer asks why their 120B mixture-of-experts model (120 billion parameters, split among many specialists) writes faster than their 70B dense model (70 billion, all used for every token). What do you tell them?Show answerHide

In plain words

Writing speed depends on how much the model reads for each token, not on its total size. The MoE reads about 5 billion parameters per token; the dense model reads all 70 billion.

Picture it

Looking something up in a 1,200-page encyclopedia is quick when the index sends you to 50 pages. Reading a 700-page manual cover to cover is slow, although it is the smaller book. You still need shelf space for all 1,200 pages.

With real numberslesson 12’s published sizes and the lab book’s DGX Spark measurements

  • gpt-oss 120B: 117 billion parameters in total, 5.1 billion used per token (the published figures in lesson 12’s script).
  • A 70B dense model uses all 70 billion for every token: 70 ÷ 5.1 = about 14 times as much to read.
  • Writing speed is capped at memory speed ÷ bytes read per token, so fewer bytes means faster writing.
  • Measured on a DGX Spark (lab book): gpt-oss 120B writes 55.4 tokens a second; Llama 3.1 70B stored at 8 bits writes 2.7.
  • 55.4 ÷ 2.7 = about 20 times faster, although the MoE holds more parameters (117 billion against 70).
  • The measured gap is bigger than 14 partly because the lab book’s two runs used different server software (llama.cpp and SGLang) and number formats. Read it as roughly 14 to 20 times.

Words to know

Parameters
The model’s learned numbers. Example: 70B means 70 billion.
Active parameters
The parameters read to write one token. Example: 5.1 billion for gpt-oss 120B.
Decode
The writing phase, one token at a time; its speed is set by the bytes read per token.
DGX Spark
NVIDIA’s desktop AI computer: 128 GB of memory moving 273 GB a second.
Go deeper: the engineer version

The kit's question

A customer asks why their 120B MoE is faster than their 70B dense. What do you say?

The kit's answer

About 5B active parameters against 70B read per token, so decode is bandwidth ÷ active bytes.

More detail: Decode is memory-bound: tok/s is about bandwidth × efficiency ÷ bytes read per token. For an MoE those bytes are the chosen experts plus the parts every token uses (attention, embeddings, router) and the KV cache, each at its own precision. The lab book inverts the Spark figure directly: 273 ÷ 55 ≈ 5 GB per token out of a 63 GB file, under 10% of the model. The two figures also differ in format: the 70B is FP8, 1 byte per parameter; gpt-oss 120B is MXFP4, a 4-bit format. Memory is still sized on the total: all 59 GiB (63.4 GB) must fit.

Why can a mixture-of-experts model be harder, not easier, to serve to many people at once on GPUs?Show answerHide

In plain words

Different people’s tokens need different experts, so they cannot all share one read of the model, as they do with a dense model. Spread over several GPUs, tokens must also travel between chips to reach their experts.

Picture it

A bus is efficient because everyone rides the same route. If every passenger needs a different part of town, you end up running many half-empty minibuses. And if the specialists work in different buildings, patients spend time walking between them.

With real numbersDay 3’s practice server and the field guide’s expert example

  • Dense model on Day 3’s practice server: 8 people got about 5 times one person’s total output, because they share each read of the model.
  • The field guide’s MoE example: 4 experts in a layer, 2 used per token.
  • One person: 2 of the 4 experts are read, half of that layer’s experts.
  • Two people whose tokens pick different pairs: all 4 experts are read, and each read serves only one person.
  • With the 4 experts on 4 GPUs, as in the field guide, each token is sent to the GPUs that hold its 2 experts, and the results are sent back.

Words to know

Expert
One specialist block in an MoE; only the few the router picks are read for a token.
Batching
Serving several people with one pass over the model.
Expert parallelism
Placing different experts on different GPUs, so tokens travel to the GPUs that hold their experts.
All-to-all
A network step where every GPU sends data to every other GPU. Example: moving tokens to their experts.
Go deeper: the engineer version

The kit's question

Why can MoE be harder to serve at high concurrency on GPUs?

The kit's answer

Different tokens hit different experts, so batches fragment. Expert parallelism across GPUs and all-to-all traffic also add complexity.

More detail: Batched decode shares one weight read across the batch. In an MoE layer each token’s router picks its own experts, so as the batch grows the set of experts read approaches all of them, while each expert’s share of the batch shrinks: more bytes per step and smaller, less efficient matrix multiplies. Expert parallelism puts experts on different GPUs and needs an all-to-all exchange of token activations in every MoE layer, out to the experts and back. That adds network traffic, and the busiest expert sets the pace.

A model needs 160 GB of memory, and each H100 GPU has 80 GB, so it must be split over 2 GPUs. Do you split every layer (one of the model’s stacked processing stages) across both (TP=2, tensor parallel) or give each GPU half of the layers (PP=2, pipeline parallel)?Show answerHide

In plain words

Split every layer (TP=2) when both GPUs sit in one server with a very fast link between them: each token then finishes sooner. Use the half-the-layers split (PP=2) only when the link between the GPUs is slow.

Picture it

Two cooks can split every dish: each chops half the vegetables, and they combine their halves before every next step. That constant passing only works side by side. Or one cook makes starters and the other mains: one hand-off per order, fine in separate kitchens. But each order waits for both, and a cook stands idle when orders are few.

With real numbersthe check’s numbers, the field guide and lesson 14’s two-box notes

  • 160 GB ÷ 80 GB per GPU = 2 GPUs at the very least, before any room for notes.
  • TP=2: each GPU holds half of every layer, and the two sync inside every layer, many times per token.
  • Inside one server, NVLink joins H100s at about 900 GB a second (field guide): fast enough for constant syncing.
  • Two DGX Sparks are joined by a 200-gigabit cable: 200 ÷ 8 = 25 GB a second. NVLink’s 900 adds both directions together; even halved to 450, it is 18 times the cable (450 ÷ 25).
  • Over a link like that, PP=2 often does better: one hand-off per token between the two halves, instead of syncs in every layer.

Words to know

Tensor parallelism (TP)
Splitting every layer across GPUs so they work on each token together; they sync inside every layer.
Pipeline parallelism (PP)
Giving each GPU a block of layers; each token passes from one GPU to the next.
NVLink
NVIDIA’s very fast direct link between GPUs inside one server. Example: about 900 GB/s on H100.
Pipeline bubble
Time a GPU sits idle in a pipeline, waiting for work from the stage before it.
Go deeper: the engineer version

The kit's question

The model needs 160 GB and each H100 has 80 GB. TP=2 or PP=2?

The kit's answer

TP=2 inside one NVLink node for latency. PP only if the link between GPUs is slow.

More detail: TP shards each weight matrix and all-reduces activations inside every layer, so it cuts per-token latency but is bound by interconnect latency and bandwidth: keep it inside one NVLink domain. PP places consecutive layer blocks on each GPU and passes activations once per stage boundary. It tolerates slower links but adds per-hop latency, and bubbles unless enough requests keep every stage busy. Across boxes the kit’s advice flips to trying PP first, because TP wants NVLink-class bandwidth (TWO_BOX.md). Note that 2 × 80 GB leaves no headroom for a 160 GB model: the Fitting Big Models video adds about 10% for activations and the runtime.

Start step 1: Sizing memo and 10-minute talk (2 h)

Your progress is saved on this device.