Skip to content

Day 10 wrap-up and drill

  1. Overview
  2. Step 1
  3. Step 2
  4. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

Your sizing memo shows a crossover volume (the monthly volume at which serverless and dedicated cost the same). You have five slides, one product criticism with a number and a source with the date you read it, and you have given the 10-minute talk out loud once.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 1 · Sizing memo and 10-minute talk

Your Day 3 measurement drives everything below it: fewer people per copy means more copies and a bigger dedicated bill.

Example: 8 / 18 (the practice server, the kit’s sample: 120 ÷ 8 x 1.2 = 18)

What the customer pays with no hardware of its own. The memo assumes 90% of each prompt is cached, at half price by default.

Example: $484 / $282 (the example customer)

A fixed bill, paid whether the GPUs are busy or idle.

Example: $105,120 (18 x $8 x 730)

Below it, serverless is cheaper. The figure keeps today’s copies fixed, so re-size the copies before you quote a volume near it. A fine-tune or a strict speed promise can need dedicated GPUs at any volume.

Example: about 744 million (744,475,921), 372 times the example’s 2 million

One timed run shows where you ramble and where a number is missing.

Example: 10

Step 2 · One product criticism

The criticism itself, specific enough that someone at Fireworks could reproduce it.

Example: Training my LoRA cost cents, but serving it needed a dedicated deployment billed by the hour (Day 7).

What turns an opinion into a finding.

Example: About $0.15 to train by the script’s estimate (0.3 million tokens x $0.50); the eval deployment ran 38 minutes, $5.07. LoRA docs: dedicated only. Pricing page: fine-tuned models at base-model prices. The lab book read both on 23 September 2026; read them again and write your own date.

Why the product team should care, in the customer’s money or time.

Example: A customer with one fine-tune and light traffic, 20 million tokens a day: about $122 a month at the base model’s per-token price, but up to $5,840 a month to keep its dedicated H100 ready around the clock, 48 times more.

Turns the criticism into a product proposal: the lab book’s “turn recurring pain into product proposals”.

Example: Per-token serving for LoRAs on the most-used base models. First, make the pricing page and the LoRA docs agree, and show the hourly cost when a deployment starts.

What you say when they ask what you would change about Fireworks.

Example: Fine-tuning on Fireworks was easy: my LoRA trained for about 15 cents. Serving it cost $5 for a 38-minute eval, because trained LoRAs run only on a dedicated GPU, $5,840 a month if kept ready. I’d propose per-token serving for LoRAs on popular base models. How do field engineers get that onto the roadmap?

Under 20 seconds each. Record yourself once and listen back.

Step 1 · Sizing memo and 10-minute talk

Question A new customer asks: should we run this on serverless or on dedicated GPUs? How would you approach it?

One clear answer

I’d start with their traffic shape and SLO, measure a concurrency curve on the candidate model, apply the cheap levers first (caching and structured output), and only then decide between serverless and dedicated, with the crossover point written down.

What this means
  • “I’d start with their traffic shape and SLO”: Before any test, I write down how much they send and when (the traffic shape) and their speed promise (the SLO, service level objective). The memo’s example: 120 people at the busiest moment, 3,200 tokens in and 60 out per ticket, 2 million tickets a month, and 95 of 100 first words within 1,000 ms.
  • “measure a concurrency curve on the candidate model”: I run Day 3’s test on the model I would propose: raise the number of people using it at the same time until the first word comes too late. On the practice server, with 8 people, 95 of 100 got their first word within 41.4 ms (thousandths of a second). With 16, 95 of 100 got it within 3,357 ms, over 3 seconds: about 80 times longer, because requests were waiting in line.
  • “apply the cheap levers first (caching and structured output)”: I use the fixes that need no new hardware before buying any: reuse the repeated opening of each prompt (Day 5) and force well-formed replies so nothing is retried (Day 6). If 90% of each prompt is cached at half price, as the memo assumes, the bill falls from $484 to $282 a month.
  • “only then decide between serverless and dedicated”: Then I choose between paying per token on shared models (serverless) and renting GPUs by the hour (dedicated), using the numbers above.
  • “with the crossover point written down”: The memo states the monthly volume at which the two cost the same, so the customer knows when to look again. With the practice server’s capacity: about 744 million tickets a month, 372 times the example’s 2 million. That keeps today’s 18 copies fixed, and they could not carry 372 times the traffic, so read it as: at this capacity, dedicated does not pay. Serverless for now.

2 worked examples from today’s videos. Work each on paper, then reveal.

From the video: The Platform Map 2:51

Whiteboard 1 of 2

Build slide 4 of your talk, the cost options, for the Platform Map video’s customer: 20 million tokens (word pieces) a day, 16 million in and 4 million out, on gpt-oss 120B (a large open model, about 120 billion learned numbers). Price three options, serverless, serverless with caching and one dedicated GPU, and say where they cross over.

Given, in plain words

Serverless prices (lab book): $0.15 per million tokens in, $0.015 per million for input the server has cached (reused from an earlier request), $0.60 per million out. A dedicated H100 GPU: $8 an hour, busy or idle; a month is 730 hours. For the caching option, assume 90% of the input is a repeated opening, as the capstone memo does.

Reveal the answerHide the answer

Answer · in plain words

Serverless costs about $146 a month, or about $87 with caching. One dedicated H100 costs $5,840. Recommend serverless with caching, and put the crossover on the slide: the GPU only pays off at about 67 times today’s traffic.

Picture it

Taxis at full fare, taxis with a regular-rider discount, or a car of your own. The discount costs nothing extra, so it always beats full fare. The car costs the same every month, so it only pays off once you ride so much that even your discounted fares add up to more.

The 90% cached share is an assumption

Measure the customer’s real reuse (Day 5) before it goes on a slide. And check with a sweep (Day 3) that one GPU could carry the traffic at the crossover.

Worked answer, step by step

  1. Serverless: input 16 million x $0.15 per million = $2.40 a day; output 4 million x $0.60 = $2.40. Total $4.80 a day.
  2. Per month: $4.80 x 30.4 days (730 hours ÷ 24) = about $146.
  3. Serverless with caching: 90% of the input, 14.4 million tokens, at $0.015 per million = $0.216; the other 1.6 million at $0.15 = $0.24. Input $0.456 a day instead of $2.40.
  4. Output is unchanged: $0.456 + $2.40 = $2.856 a day, x 30.4 = about $87 a month.
  5. Dedicated: $8 x 730 hours = $5,840 a month, however much traffic it carries.
  6. Crossover: serverless bills grow in step with traffic; the GPU’s does not. So they meet at $5,840 ÷ $146 = 40 times today’s traffic, or $5,840 ÷ $87 = about 67 times against serverless with caching.
  7. Slide 4 says: serverless with caching now; dedicated only past about 67 times today’s traffic, or when a fine-tune or a strict speed promise needs its own GPU.
Go deeper: the engineer version

The kit's question

Worked example · 20M tokens a day · 16M input · 4M output · gpt-oss-120b pricing

The kit's answer

input: 16M × $0.15 = $2.40. output: 4M × $0.60 = $2.40. per day: $4.80. per month: ≈ $146. one dedicated H100: $8/hr × 730 h = $5,840. Recommendation: Serverless until utilisation justifies the GPU — and say so plainly.

More detail: The video’s $146 is $4.80 x 30.42 = $146.0; the narration rounds it to “about a hundred and fifty a month”. The cached price is the lab book’s for gpt-oss 120B ($0.015 per million, 90% off). build_memo.py assumes 90% of input is cached but at only 50% off, so pass --cached-discount 0.9 for this model. The lab book’s memo template asks for these rows: Option A, serverless with its cache hit rate and monthly total; Option B, dedicated with its break-even volume. A region-pinned deployment costs 1.5 times as much ($8,760 a month), which moves the crossovers to 60 and about 100 times today’s traffic.

Words to know

Cached input
Input tokens the server reused from an earlier request, billed at a discount. Example: $0.015 instead of $0.15 per million for gpt-oss 120B.
Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: $8 an hour, $5,840 a month.
Crossover (break-even)
The traffic at which two ways of serving cost the same; below it, paying per token is cheaper. Example: about 67 times today’s traffic here.
Utilisation
The share of time a GPU does useful work. Example: the video’s rule, serverless until utilisation justifies the GPU.

From the video: The Platform Map 2:51

Whiteboard 2 of 2

The table right after today’s rewatch (you watched it on Day 6): a customer’s routine requests go to a large frontier model at $6.00 per million output tokens. A fine-tuned 8B model (8 billion learned numbers) would cost $0.20. How much cheaper is that, and what must be true before you promise the saving on Fireworks?

Given, in plain words

Per million output tokens (the tokens the model writes): frontier model $6.00, fine-tuned 8B $0.20. Task accuracy 87% before, 91% after. p95 latency (95 of 100 requests finish within it) 1.9 s before, 0.6 s after. The video says these numbers are examples, not measurements. On Fireworks a trained LoRA fine-tune is served only on a dedicated deployment, $8 an hour (lab book).

Reveal the answerHide the answer

Answer · in plain words

30 times cheaper per token, and better at the task. But on Fireworks the fine-tune needs its own GPU, about $5,840 a month kept ready around the clock. The saving is real only once the traffic pays for that GPU, or several fine-tunes share it.

Picture it

A new supplier sells each box at a thirtieth of your current price, but only if you rent their warehouse by the month. For a big shop that is a bargain; for a small one the rent eats the saving.

Careful

The break-even counts only output tokens at the frontier price and assumes one GPU can carry the traffic: measure both before quoting. The gap between the video’s per-token price and Fireworks’ dedicated-only rule is also a ready-made product criticism for today’s step 2.

Worked answer, step by step

  1. Unit cost: $6.00 ÷ $0.20 = 30. Per output token the fine-tuned model is 30 times cheaper, as the video says.
  2. Quality: task accuracy rises from 87% to 91%, and p95 latency falls from 1.9 s to 0.6 s.
  3. The catch on Fireworks: a trained LoRA serves only on a dedicated deployment. One H100 kept ready all month: $8 x 730 hours = $5,840.
  4. Break-even on output tokens alone: $5,840 ÷ $6.00 per million = about 973 million output tokens a month, about 32 million a day.
  5. Below that, the fixed GPU costs more than the frontier bill it replaces. Above it, the 30-times saving shows up.
  6. Shared by 12 fine-tunes (Day 7’s customer who wanted one per business unit, 12 in all), each carries $5,840 ÷ 12 = about $487 a month and breaks even at about 81 million output tokens a month ($487 ÷ $6.00).
  7. Answer: specialise the routine traffic once volume passes the break-even, or several fine-tunes share one deployment, and show the maths in the memo.
Go deeper: the engineer version

The kit's question

Specialise the routine traffic · frontier model · fine-tuned 8B. Specialisation is the other lever, and it compounds with the first.

The kit's answer

$ per 1M output: $6.00 · $0.20. task accuracy: 87% · 91%. p95 latency: 1.9 s · 0.6 s. what it cost: — · one afternoon + ~$2. Illustrative numbers, but the shape matches the published case studies. Move the routine ninety percent of traffic to a fine tuned eight billion parameter model and the unit cost falls by roughly thirty times, while quality on that narrow task usually goes up, not down.

More detail: The narration’s “roughly thirty times” is per output token: $6.00 ÷ $0.20 = 30. The video’s $0.20 equals the flat serverless price the lab book lists for 4B to 16B models, so it assumes the fine-tune is billed like its base model. The lab book warns that the pricing page says exactly that, “at the same price as base models”, while the LoRA docs say trained LoRAs deploy only to on-demand deployments: reconcile the two before promising serverless economics for a fine-tune. Multi-LoRA needs a BF16 deployment shape; FP8 and FP4 shapes cannot host adapters.

Words to know

Frontier model
One of the largest, most capable models, usually priced highest. Example: $6.00 per million output tokens here.
Unit cost
The price of one unit of work, such as a million tokens written. Example: $6.00 against $0.20.
LoRA
A small add-on trained on your examples while the big model stays unchanged. Example: Day 7’s ticket-sorting fine-tune.
Multi-LoRA
One deployment serving many LoRA add-ons on the same base model, so they share its cost. Example: 12 fine-tunes, about $487 a month each.

How you will use this

A customer can use your sizing memo to choose how to run the system and what to budget. Its five slides help you explain the recommendation to a team. The memo draws on your own speed benchmark (Days 1 and 3) and model-improvement results (Days 7 and 8); the product criticism records a real obstacle you found.