Skip to content

Sizing memo and 10-minute talk

course 17 of 22

lesson 15 · lessons/15-capstone

2 h Free Anywhere

What you will do and why

Turn what you measured on Days 1 to 9 into a sizing memo, the written advice a customer gets after the first call. Then build a 10-minute talk from it.

Why it matters: The memo’s example customer, a support-ticket assistant, handles 2 million tickets a month. Paying per token costs $484 a month, or $282 if the repeated opening of each prompt is reused (cached), as the memo assumes.

You are done when: results/SIZING_MEMO.md has your own number in every section you ran and a crossover volume written down. You have five slides and have given the talk out loud once, in about 10 minutes.

Rewatch from Day 6: The Platform Map, 1:18 to 1:59ShowHide

A field engineer is paid for the recommendation, not the benchmark. One script turns what you measured into a sizing memo: what the customer needs, how many server copies, what each option costs a month and which to pick.

Picture it

A builder’s quote for a new kitchen. It starts from your measurements, prices two or three options, recommends one and lists what could go wrong. You decide from the quote, not from the tape measure.

With real numbersthe memo’s example customer in build_memo.py, with the practice server’s capacity from Day 3

  • The customer: 2,000,000 tickets a month, each about 3,200 tokens sent to the model (about 2,400 words) and 60 written back (about 45 words), with 120 people at the busiest moment.
  • Paying per token (serverless), at the script’s built-in example prices ($0.07 per million tokens in, $0.30 per million out). In: 2,000,000 x 3,200 = 6,400 million tokens; 6,400 x $0.07 = $448. Out: 2,000,000 x 60 = 120 million; 120 x $0.30 = $36. Total $484 a month.
  • With caching: the script assumes 90% of each prompt is the same opening and bills that part at half price. Full price on 10%: $448 x 0.1 = $44.80. Half price on 90%: $448 x 0.9 ÷ 2 = $201.60. Input $246.40, plus $36 out = $282 a month.
  • Dedicated GPUs: the practice server holds 8 people per copy, so 120 ÷ 8 x 1.2 (20% spare) = 18 copies. 18 x $8 an hour x 730 hours = $105,120 a month. The 8 is a stand-in: a data-center GPU holds far more people, so a real dedicated bill would be lower.
  • Crossover: the GPUs cost $105,120 ÷ $282.40 = 372 times the cached serverless bill. That bill grows with traffic and the GPU bill does not, so they meet at 372 x 2 million = about 744 million tickets a month. This keeps today’s 18 copies fixed, and they could not carry 372 times the traffic: read it as “at this capacity, dedicated does not pay”. So: serverless now.
  • Your own memo uses your real-server sweep rows (Day 3’s llama.cpp run on your Mac) when they exist, so your copies, GPU bill and crossover will differ from these.

Words to know

Sizing memo
A short document recommending how to run a customer’s workload: how many copies, which option, what it costs, with the maths shown. Example: results/SIZING_MEMO.md.
Serverless
Fireworks’ shared, always-on models, billed per token you send and receive. Example: $282 a month for the example customer, with caching.
Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: $8 an hour, $5,840 for a month.
Crossover (break-even)
The traffic at which two ways of serving cost the same; below it, paying per token is cheaper. Example: about 744 million tickets a month here, if the current copies could carry it.

From the lesson

lessons/15-capstone/README.md

A Field Engineer is paid for the recommendation, not the benchmark. This lesson turns six of the CSV files in results/ into the document you would send after a discovery call, and a ten-minute talk you can give to the customer’s team.

build_memo.py reads what you measured on Days 1 to 9 and writes results/SIZING_MEMO.md with: needs → latency profile → capacity and replicas → levers (caching, schema, batch, LoRA, speculation) → three cost options and the serverless-versus-dedicated crossover → recommendation → risks. Anything you haven’t measured shows up as TODO rather than an invented number.

Terminal window
cd lessons/15-capstone
python build_memo.py
python build_memo.py --peak-concurrency 200 --tickets-per-month 5e6 --slo-ms 800 \
--serverless-in 0.15 --serverless-out 0.60 --gpu-hour 8 # your scenario, current prices
open ../../results/SIZING_MEMO.md

What each command does

  1. python build_memo.py

    Builds the memo for the script’s example customer: Acme Cloud’s support-ticket assistant, 120 people at the busiest moment, 2 million tickets a month, and a promise that 95 of 100 first words arrive within 1,000 ms (one second). It prints the memo and saves it as results/SIZING_MEMO.md. Look for $484 and $282 in section 5. A section with no result file says TODO (still to do), never a guess; Day 1’s smoke test saved practice-server numbers in all six files, so check each number comes from your own run. Each section names the lesson its numbers come from: lesson 01 is Day 1, 03 is Day 3, 05 is Day 4, 06 is Day 5, 07 and 08 are Day 6, 09 to 11 are Days 7 and 8, 13 is Day 9.

  2. python build_memo.py --peak-concurrency 200 --tickets-per-month 5e6 --slo-ms 800

    The same memo for a bigger customer: 200 people at once, 5 million tickets a month (5e6 means 5 x 1,000,000), 95 of 100 first words within 800 ms, and the lab book’s gpt-oss 120B prices ($0.15 per million tokens in, $0.60 out) with a GPU at $8 an hour. It overwrites the first memo. For a real customer, change these numbers to theirs. Look for $2,580 a month on serverless and $1,500 with caching. The script assumes cached input costs half price, but the lab book lists gpt-oss 120B’s at $0.015, 90% off: add --cached-discount 0.9 and the cached bill drops to $636.

  3. open ../../results/SIZING_MEMO.md

    Opens the memo in your Mac’s default app for .md files (Markdown: plain text with light formatting, such as # for a heading). If it prints an error instead, run open -e ../../results/SIZING_MEMO.md to open it in TextEdit. Look for seven sections, from “1. What they need” to “7. Risks and what I’d validate next”, and in section 3 the name of the server your capacity was measured on.

  1. The customer’s problem, in their numbers: peak concurrency, SLO, volume, format needs.
  2. What I measured: the Day 3 concurrency curve with the SLO line, plus the Day 1 TTFT/ITL table.
  3. Three levers: prefix caching (Day 5), schema-constrained output (Day 6) and a fine-tune (Days 7 and 8), each with your before/after.
  4. Cost options: serverless, serverless with caching, dedicated, and where they cross over.
  5. Recommendation, risks and the next two weeks of a proof of concept.
they ask your answer uses
“Why is our TTFT high but typing speed fine?” prefill vs decode, prompt length, queueing (Days 1 and 3)
“Can we run a 70B on one GPU?” weights + KV maths, quantization, TP (Days 4 and 9)
“Serverless or dedicated?” the crossover table, SLO, LoRA needs (today’s memo)
“Is fine-tuning worth it?” the eval table: quality, p95 and $/1k (Days 7 and 8)
“Why is the MoE faster than the smaller dense model?” active vs total params (Day 9)
“Will speculative decoding help us?” acceptance on their traffic (Day 9)
“Our agent bill is too high.” stable-first prompts, cached tokens, batch for evals (Days 5 and 6)

Question A new customer asks: should we run this on serverless or on dedicated GPUs? How would you approach it?

One clear answer

I’d start with their traffic shape and SLO, measure a concurrency curve on the candidate model, apply the cheap levers first (caching and structured output), and only then decide between serverless and dedicated, with the crossover point written down.

What this means

  • “I’d start with their traffic shape and SLO”: Before any test, I write down how much they send and when (the traffic shape) and their speed promise (the SLO, service level objective). The memo’s example: 120 people at the busiest moment, 3,200 tokens in and 60 out per ticket, 2 million tickets a month, and 95 of 100 first words within 1,000 ms.
  • “measure a concurrency curve on the candidate model”: I run Day 3’s test on the model I would propose: raise the number of people using it at the same time until the first word comes too late. On the practice server, with 8 people, 95 of 100 got their first word within 41.4 ms (thousandths of a second). With 16, 95 of 100 got it within 3,357 ms, over 3 seconds: about 80 times longer, because requests were waiting in line.
  • “apply the cheap levers first (caching and structured output)”: I use the fixes that need no new hardware before buying any: reuse the repeated opening of each prompt (Day 5) and force well-formed replies so nothing is retried (Day 6). If 90% of each prompt is cached at half price, as the memo assumes, the bill falls from $484 to $282 a month.
  • “only then decide between serverless and dedicated”: Then I choose between paying per token on shared models (serverless) and renting GPUs by the hour (dedicated), using the numbers above.
  • “with the crossover point written down”: The memo states the monthly volume at which the two cost the same, so the customer knows when to look again. With the practice server’s capacity: about 744 million tickets a month, 372 times the example’s 2 million. That keeps today’s 18 copies fixed, and they could not carry 372 times the traffic, so read it as: at this capacity, dedicated does not pay. Serverless for now.

Your numbersSaved on this device and collected in the Day 10 wrap-up.

Hint: Memo section 3, Capacity: the two numbers in bold (between ** marks when you open the plain file). Write both, for example 8 / 18. The first line ends with the server it was measured on.

Hint: Memo section 5, the first two rows. Write both, for example $484 / $282.

Hint: Memo section 5, the third row: copies x $8 an hour x 730 hours.

Hint: The Crossover line under the section 5 table.

Hint: Time yourself with the five slides. The target is 10 minutes, about 2 per slide.

results/SIZING_MEMO.md has your own number in every section you ran and a crossover volume written down. You have five slides and have given the talk out loud once, in about 10 minutes.

ModuleNotFoundError: No module named 'felab'
The virtual environment is off in this window. From the 03-labs folder run source .venv/bin/activate, then cd lessons/15-capstone and run the command again.
A section says TODO although you did that day.
The memo reads fixed file names in results/. From lessons/15-capstone, run ls ../../results/ and look for that day’s file (the list is under Before you start on the Day 10 overview). If it is missing, run that day’s main command again; the practice-server version is free.
Replicas, the dedicated cost and the crossover say TODO, but 03-sweep.csv is there.
No crowd size in your sweep kept the memo’s speed promise, by default 95 of 100 first words within 1,000 ms. Once your sweep has any real-server rows, the memo ignores the practice-server rows. From lessons/15-capstone, open ../../results/03-sweep.csv and read the ttft_p95 column (in ms). Pass the customer’s real target with --slo-ms, or run Day 3’s sweep again with smaller crowds, for example --levels 1,2,4.
The memo warns that capacity came from a laptop or mock server.
Expected: your Day 3 sweep ran on your Mac or the practice server. Data-center GPUs running a production engine such as vLLM hold far more people per copy, so treat the copies and the dedicated cost as placeholders and say so on your risks slide. Day 9’s optional DGX Spark step, or a Fireworks deployment, gives real figures.
The fine-tune row names mock mock-8b-lora instead of your own model.
The memo compares every row in results/09-eval.csv, including the practice-server rows Day 1’s smoke test saved. Open that file in TextEdit, delete the lines containing mock mock-8b, keep the first line (the column names), save, and run python build_memo.py again.
The crossover is hundreds of millions of tickets a month.
Not a bug in your data: the memo assumes today’s copies carry any volume. With a handful of people per copy, the dedicated option needs many GPUs, so serverless wins by a wide margin: with the practice server’s 8, the example’s crossover is about 744 million tickets, 372 times its volume. Replace the capacity with a real GPU measurement before you quote it.
open says no application knows how to open the memo.
Your Mac has no app set for .md files. Run open -e ../../results/SIZING_MEMO.md to open it in TextEdit instead.
The talk runs well over 10 minutes.
Five slides in 10 minutes is about 2 minutes each. Keep one number per point, and move tables to a backup slide you show only if asked.
build_memo.py Python · 144 lines
"""
Lesson 15 — turn everything you measured into a customer sizing memo.
Reads results/*.csv from lessons 01–14, picks the latest numbers, applies them to a
customer scenario, and writes results/SIZING_MEMO.md: the document you'd send after
a discovery call, and the backbone of a 10-minute project walkthrough.
python build_memo.py
python build_memo.py --peak-concurrency 200 --tickets-per-month 5e6 --slo-ms 800
Every number is traceable to a CSV row, and anything missing is marked TODO rather
than invented. That habit matters more than the numbers.
"""
import argparse
import csv
import math
from pathlib import Path
from felab.results import RESULTS
def latest(name: str) -> list[dict]:
p = RESULTS / f"{name}.csv"
return list(csv.DictReader(p.open())) if p.exists() else []
def f(x, fmt="{:,.0f}", todo="TODO (run the lesson)"):
try:
return fmt.format(float(x))
except (TypeError, ValueError):
return todo
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--customer", default="Acme Cloud (support triage agent)")
ap.add_argument("--peak-concurrency", type=int, default=120)
ap.add_argument("--tickets-per-month", type=float, default=2e6)
ap.add_argument("--slo-ms", type=float, default=1000, help="p95 TTFT target")
ap.add_argument("--in-tokens", type=int, default=3200, help="avg prompt tokens per ticket (agent preamble + ticket)")
ap.add_argument("--out-tokens", type=int, default=60)
ap.add_argument("--serverless-in", type=float, default=0.07, help="$ per 1M input tokens — check pricing page")
ap.add_argument("--serverless-out", type=float, default=0.30)
ap.add_argument("--cached-discount", type=float, default=0.5, help="fraction off cached input — check pricing")
ap.add_argument("--gpu-hour", type=float, default=8.0, help="$ per dedicated GPU-hour — check pricing")
a = ap.parse_args()
lat = latest("01-latency")
sweep = latest("03-sweep")
prefix = latest("06-prefix")
struct = latest("07-structured")
evals = latest("09-eval")
spec = latest("13-spec")
# --- capacity from the sweep: best concurrency under SLO for the most recent non-mock target
real = [r for r in sweep if r["target"] != "mock"] or sweep
under = [r for r in real if float(r["ttft_p95"]) <= a.slo_ms]
per_replica = max((int(r["conc"]) for r in under), default=None)
sweep_src = real[-1]["target"] if real else None
replicas = math.ceil(a.peak_concurrency / per_replica * 1.2) if per_replica else None # +20% headroom
# --- prefix caching share
stable = [r for r in prefix if r["order"] == "stable-first"]
speedup = stable[-1]["speedup_x"] if stable else None
# --- serverless monthly cost, with and without caching (assume 90% of input is the shared preamble)
n = a.tickets_per_month
cost_in = n * a.in_tokens / 1e6 * a.serverless_in
cost_out = n * a.out_tokens / 1e6 * a.serverless_out
cached_share = 0.9
cost_in_cached = cost_in * (1 - cached_share * a.cached_discount)
dedicated = replicas * a.gpu_hour * 730 if replicas else None
# --- quality
best_eval = max(evals, key=lambda r: float(r["category_%"])) if evals else None
worst_eval = min(evals, key=lambda r: float(r["category_%"])) if evals else None
schema = next((r for r in reversed(struct) if r["mode"] == "json_schema"), None)
prompt_only = next((r for r in reversed(struct) if r["mode"] == "prompt"), None)
per_ticket = (cost_in_cached + cost_out) / n
crossover = f"{dedicated / per_ticket:,.0f}" if dedicated else "TODO"
cap_warning = ("- ⚠ Capacity came from a laptop or mock server. Production GPUs with vLLM/SGLang hold far "
"more users per replica; re-run `sweep.py` against the Spark or a Fireworks deployment "
"before quoting replica counts.\n") if sweep_src in ("mock", "ollama", "llamacpp", "mlx") else ""
# (built outside the f-string so this runs on Python 3.10/3.11 too)
lat_lines = "".join(
f"- {r['target']} · {r['model']}: TTFT p50 {f(r['ttft_p50_ms'])} ms / p95 {f(r['ttft_p95_ms'])} ms"
f" · {f(r['tok_per_s'], '{:.0f}')} tok/s per user (prompt ≈{r['prompt_tok']} tok)\n"
for r in lat[-4:]) or "- TODO (run lesson 01)\n"
spec_text = ", ".join(f"{r['label']}/{r['traffic']} {f(r['tok_s'])} tok/s" for r in spec[-4:]) or "TODO"
md = f"""# Sizing memo: {a.customer}
*Generated by `lessons/15-capstone/build_memo.py` from measured results. Prices are inputs; verify them on the pricing page.*
## 1. What they need
- Peak **{a.peak_concurrency} concurrent** agent sessions, **p95 TTFT ≤ {a.slo_ms:.0f} ms**
- **{n:,.0f} tickets/month**, about {a.in_tokens:,} input and {a.out_tokens} output tokens each (the agent preamble dominates)
- Output must be **schema-valid JSON** for the ticketing integration
## 2. Latency profile (lesson 01)
{lat_lines}
## 3. Capacity (lesson 03)
- One replica holds **{per_replica or 'TODO'}** concurrent users within SLO (measured on `{sweep_src or '—'}`)
- Replicas for peak with 20% headroom: **{replicas or 'TODO'}**
{cap_warning}
## 4. Levers that change the bill
| lever | evidence | effect |
|---|---|---|
| Prefix caching: stable preamble first, `user` for affinity | lesson 06: **{f(speedup, '{:.1f}')}×** lower TTFT on warm calls | about 90% of input tokens billed at the cached rate |
| Constrained decoding (`json_schema`) | lesson 07: {f(prompt_only and prompt_only['schema_valid_%'], '{:.0f}')}% → **{f(schema and schema['schema_valid_%'], '{:.0f}')}%** valid | no retry loop, no parser failures |
| Batch API for evals and back-fills | lesson 08 | about 50% off for non-interactive work |
| LoRA fine-tune | lesson 09–11: category accuracy {f(worst_eval and worst_eval['category_%'], '{:.0f}')}% → **{f(best_eval and best_eval['category_%'], '{:.0f}')}%** ({best_eval['candidate'] if best_eval else '—'}) | a smaller model at higher accuracy; needs a dedicated deployment (multi-LoRA shares it) |
| Speculative decoding | lesson 13: {spec_text} | pays on structured output; measure α first |
## 5. Cost options (monthly, rough)
| option | how | $/month |
|---|---|---|
| Serverless, no caching | {n:,.0f} × tokens × list price | **${cost_in + cost_out:,.0f}** |
| Serverless with prefix caching | 90% of input cached at {a.cached_discount:.0%} off | **${cost_in_cached + cost_out:,.0f}** |
| Dedicated, {replicas or '?'} replicas × ${a.gpu_hour}/GPU-h × 730 h | predictable latency; required for the LoRA | **{('$' + format(dedicated, ',.0f')) if dedicated else 'TODO'}** |
**Crossover:** at about **{crossover}** tickets/month, dedicated costs the same as serverless with caching.
Below that, serverless wins. Above it, or when a LoRA or a strict SLO is needed, dedicated wins.
## 6. Recommendation
1. **Now:** serverless with the preamble restructured for caching, plus `json_schema`. It is the cheapest path to production.
2. **In 4–6 weeks,** if accuracy is the blocker: LoRA on a small base, served on dedicated capacity as a multi-LoRA deployment, justified by the eval table.
3. **Always:** nightly evals through the Batch API. Watch p95 TTFT, not averages.
## 7. Risks and what I'd validate next
- Traffic shape: burstiness at peak. Replay a real hour of traffic through `sweep.py`.
- Quality on real tickets, not synthetic ones: 200 hand-labelled examples, with the LLM judge calibrated against them.
- KV headroom if prompts grow (lesson 05): cap the context and consider an FP8 KV cache.
"""
out = RESULTS / "SIZING_MEMO.md"
out.write_text(md)
print(md)
print(f"\nsaved → {out}")
if __name__ == "__main__":
main()