Sizing memo and 10-minute talk
course 17 of 22
lesson 15 · lessons/15-capstone
2 h Free Anywhere
What you will do and why
Turn what you measured on Days 1 to 9 into a sizing memo, the written advice a customer gets after the first call. Then build a 10-minute talk from it.
Why it matters: The memo’s example customer, a support-ticket assistant, handles 2 million tickets a month. Paying per token costs $484 a month, or $282 if the repeated opening of each prompt is reused (cached), as the memo assumes.
You are done when: results/SIZING_MEMO.md has your own number in every section you ran and a crossover volume written down. You have five slides and have given the talk out loud once, in about 10 minutes.
Rewatch from Day 6: The Platform Map, 1:18 to 1:59ShowHide
In plain words
Section titled “In plain words”A field engineer is paid for the recommendation, not the benchmark. One script turns what you measured into a sizing memo: what the customer needs, how many server copies, what each option costs a month and which to pick.
Picture it
A builder’s quote for a new kitchen. It starts from your measurements, prices two or three options, recommends one and lists what could go wrong. You decide from the quote, not from the tape measure.
With real numbersthe memo’s example customer in build_memo.py, with the practice server’s capacity from Day 3
- The customer: 2,000,000 tickets a month, each about 3,200 tokens sent to the model (about 2,400 words) and 60 written back (about 45 words), with 120 people at the busiest moment.
- Paying per token (serverless), at the script’s built-in example prices ($0.07 per million tokens in, $0.30 per million out). In: 2,000,000 x 3,200 = 6,400 million tokens; 6,400 x $0.07 = $448. Out: 2,000,000 x 60 = 120 million; 120 x $0.30 = $36. Total $484 a month.
- With caching: the script assumes 90% of each prompt is the same opening and bills that part at half price. Full price on 10%: $448 x 0.1 = $44.80. Half price on 90%: $448 x 0.9 ÷ 2 = $201.60. Input $246.40, plus $36 out = $282 a month.
- Dedicated GPUs: the practice server holds 8 people per copy, so 120 ÷ 8 x 1.2 (20% spare) = 18 copies. 18 x $8 an hour x 730 hours = $105,120 a month. The 8 is a stand-in: a data-center GPU holds far more people, so a real dedicated bill would be lower.
- Crossover: the GPUs cost $105,120 ÷ $282.40 = 372 times the cached serverless bill. That bill grows with traffic and the GPU bill does not, so they meet at 372 x 2 million = about 744 million tickets a month. This keeps today’s 18 copies fixed, and they could not carry 372 times the traffic: read it as “at this capacity, dedicated does not pay”. So: serverless now.
- Your own memo uses your real-server sweep rows (Day 3’s llama.cpp run on your Mac) when they exist, so your copies, GPU bill and crossover will differ from these.
Words to know
- Sizing memo
- A short document recommending how to run a customer’s workload: how many copies, which option, what it costs, with the maths shown. Example:
results/SIZING_MEMO.md. - Serverless
- Fireworks’ shared, always-on models, billed per token you send and receive. Example: $282 a month for the example customer, with caching.
- Dedicated deployment
- A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: $8 an hour, $5,840 for a month.
- Crossover (break-even)
- The traffic at which two ways of serving cost the same; below it, paying per token is cheaper. Example: about 744 million tickets a month here, if the current copies could carry it.
From the lesson
lessons/15-capstone/README.md
What and why
Section titled “What and why”A Field Engineer is paid for the recommendation, not the benchmark. This lesson
turns six of the CSV files in results/ into the document you would send after a discovery call,
and a ten-minute talk you can give to the customer’s team.
build_memo.py reads what you measured on Days 1 to 9 and writes results/SIZING_MEMO.md with:
needs → latency profile → capacity and replicas → levers (caching, schema, batch,
LoRA, speculation) → three cost options and the serverless-versus-dedicated crossover
→ recommendation → risks. Anything you haven’t measured shows up as TODO rather
than an invented number.
cd lessons/15-capstonepython build_memo.pypython build_memo.py --peak-concurrency 200 --tickets-per-month 5e6 --slo-ms 800 \ --serverless-in 0.15 --serverless-out 0.60 --gpu-hour 8 # your scenario, current pricesopen ../../results/SIZING_MEMO.mdWhat each command does
python build_memo.pyBuilds the memo for the script’s example customer: Acme Cloud’s support-ticket assistant, 120 people at the busiest moment, 2 million tickets a month, and a promise that 95 of 100 first words arrive within 1,000 ms (one second). It prints the memo and saves it as
results/SIZING_MEMO.md. Look for $484 and $282 in section 5. A section with no result file says TODO (still to do), never a guess; Day 1’s smoke test saved practice-server numbers in all six files, so check each number comes from your own run. Each section names the lesson its numbers come from: lesson 01 is Day 1, 03 is Day 3, 05 is Day 4, 06 is Day 5, 07 and 08 are Day 6, 09 to 11 are Days 7 and 8, 13 is Day 9.python build_memo.py --peak-concurrency 200 --tickets-per-month 5e6 --slo-ms 800The same memo for a bigger customer: 200 people at once, 5 million tickets a month (
5e6means 5 x 1,000,000), 95 of 100 first words within 800 ms, and the lab book’s gpt-oss 120B prices ($0.15 per million tokens in, $0.60 out) with a GPU at $8 an hour. It overwrites the first memo. For a real customer, change these numbers to theirs. Look for $2,580 a month on serverless and $1,500 with caching. The script assumes cached input costs half price, but the lab book lists gpt-oss 120B’s at $0.015, 90% off: add--cached-discount 0.9and the cached bill drops to $636.open ../../results/SIZING_MEMO.mdOpens the memo in your Mac’s default app for
.mdfiles (Markdown: plain text with light formatting, such as # for a heading). If it prints an error instead, runopen -e ../../results/SIZING_MEMO.mdto open it in TextEdit. Look for seven sections, from “1. What they need” to “7. Risks and what I’d validate next”, and in section 3 the name of the server your capacity was measured on.
The 10-minute presentation (5 slides)
Section titled “The 10-minute presentation (5 slides)”- The customer’s problem, in their numbers: peak concurrency, SLO, volume, format needs.
- What I measured: the Day 3 concurrency curve with the SLO line, plus the Day 1 TTFT/ITL table.
- Three levers: prefix caching (Day 5), schema-constrained output (Day 6) and a fine-tune (Days 7 and 8), each with your before/after.
- Cost options: serverless, serverless with caching, dedicated, and where they cross over.
- Recommendation, risks and the next two weeks of a proof of concept.
Drill: answer each in about 20 seconds
Section titled “Drill: answer each in about 20 seconds”| they ask | your answer uses |
|---|---|
| “Why is our TTFT high but typing speed fine?” | prefill vs decode, prompt length, queueing (Days 1 and 3) |
| “Can we run a 70B on one GPU?” | weights + KV maths, quantization, TP (Days 4 and 9) |
| “Serverless or dedicated?” | the crossover table, SLO, LoRA needs (today’s memo) |
| “Is fine-tuning worth it?” | the eval table: quality, p95 and $/1k (Days 7 and 8) |
| “Why is the MoE faster than the smaller dense model?” | active vs total params (Day 9) |
| “Will speculative decoding help us?” | acceptance on their traffic (Day 9) |
| “Our agent bill is too high.” | stable-first prompts, cached tokens, batch for evals (Days 5 and 6) |
Explain what you learned
Section titled “Explain what you learned”Question A new customer asks: should we run this on serverless or on dedicated GPUs? How would you approach it?
One clear answer
I’d start with their traffic shape and SLO, measure a concurrency curve on the candidate model, apply the cheap levers first (caching and structured output), and only then decide between serverless and dedicated, with the crossover point written down.
What this means
- “I’d start with their traffic shape and SLO”: Before any test, I write down how much they send and when (the traffic shape) and their speed promise (the SLO, service level objective). The memo’s example: 120 people at the busiest moment, 3,200 tokens in and 60 out per ticket, 2 million tickets a month, and 95 of 100 first words within 1,000 ms.
- “measure a concurrency curve on the candidate model”: I run Day 3’s test on the model I would propose: raise the number of people using it at the same time until the first word comes too late. On the practice server, with 8 people, 95 of 100 got their first word within 41.4 ms (thousandths of a second). With 16, 95 of 100 got it within 3,357 ms, over 3 seconds: about 80 times longer, because requests were waiting in line.
- “apply the cheap levers first (caching and structured output)”: I use the fixes that need no new hardware before buying any: reuse the repeated opening of each prompt (Day 5) and force well-formed replies so nothing is retried (Day 6). If 90% of each prompt is cached at half price, as the memo assumes, the bill falls from $484 to $282 a month.
- “only then decide between serverless and dedicated”: Then I choose between paying per token on shared models (serverless) and renting GPUs by the hour (dedicated), using the numbers above.
- “with the crossover point written down”: The memo states the monthly volume at which the two cost the same, so the customer knows when to look again. With the practice server’s capacity: about 744 million tickets a month, 372 times the example’s 2 million. That keeps today’s 18 copies fixed, and they could not carry 372 times the traffic, so read it as: at this capacity, dedicated does not pay. Serverless for now.
Your numbersSaved on this device and collected in the Day 10 wrap-up.
Hint: Memo section 3, Capacity: the two numbers in bold (between ** marks when you open the plain file). Write both, for example 8 / 18. The first line ends with the server it was measured on.
Hint: Memo section 5, the first two rows. Write both, for example $484 / $282.
Hint: Memo section 5, the third row: copies x $8 an hour x 730 hours.
Hint: The Crossover line under the section 5 table.
Hint: Time yourself with the five slides. The target is 10 minutes, about 2 per slide.
Done when
Section titled “Done when”results/SIZING_MEMO.md has your own number in every section you ran and a crossover volume written down. You have five slides and have given the talk out loud once, in about 10 minutes.
Stuck?
Section titled “Stuck?”ModuleNotFoundError: No module named 'felab'- The virtual environment is off in this window. From the
03-labsfolder runsource .venv/bin/activate, thencd lessons/15-capstoneand run the command again. - A section says TODO although you did that day.
- The memo reads fixed file names in
results/. Fromlessons/15-capstone, runls ../../results/and look for that day’s file (the list is under Before you start on the Day 10 overview). If it is missing, run that day’s main command again; the practice-server version is free. - Replicas, the dedicated cost and the crossover say TODO, but
03-sweep.csvis there. - No crowd size in your sweep kept the memo’s speed promise, by default 95 of 100 first words within 1,000 ms. Once your sweep has any real-server rows, the memo ignores the practice-server rows. From
lessons/15-capstone, open../../results/03-sweep.csvand read thettft_p95column (in ms). Pass the customer’s real target with--slo-ms, or run Day 3’s sweep again with smaller crowds, for example--levels 1,2,4. - The memo warns that capacity came from a laptop or mock server.
- Expected: your Day 3 sweep ran on your Mac or the practice server. Data-center GPUs running a production engine such as vLLM hold far more people per copy, so treat the copies and the dedicated cost as placeholders and say so on your risks slide. Day 9’s optional DGX Spark step, or a Fireworks deployment, gives real figures.
- The fine-tune row names
mock mock-8b-lorainstead of your own model. - The memo compares every row in
results/09-eval.csv, including the practice-server rows Day 1’s smoke test saved. Open that file in TextEdit, delete the lines containingmock mock-8b, keep the first line (the column names), save, and runpython build_memo.pyagain. - The crossover is hundreds of millions of tickets a month.
- Not a bug in your data: the memo assumes today’s copies carry any volume. With a handful of people per copy, the dedicated option needs many GPUs, so serverless wins by a wide margin: with the practice server’s 8, the example’s crossover is about 744 million tickets, 372 times its volume. Replace the capacity with a real GPU measurement before you quote it.
opensays no application knows how to open the memo.- Your Mac has no app set for
.mdfiles. Runopen -e ../../results/SIZING_MEMO.mdto open it in TextEdit instead. - The talk runs well over 10 minutes.
- Five slides in 10 minutes is about 2 minutes each. Keep one number per point, and move tables to a backup slide you show only if asked.
Code in this step
Section titled “Code in this step”build_memo.py Python · 144 lines
"""Lesson 15 — turn everything you measured into a customer sizing memo.
Reads results/*.csv from lessons 01–14, picks the latest numbers, applies them to acustomer scenario, and writes results/SIZING_MEMO.md: the document you'd send aftera discovery call, and the backbone of a 10-minute project walkthrough.
python build_memo.py python build_memo.py --peak-concurrency 200 --tickets-per-month 5e6 --slo-ms 800
Every number is traceable to a CSV row, and anything missing is marked TODO ratherthan invented. That habit matters more than the numbers."""import argparseimport csvimport mathfrom pathlib import Path
from felab.results import RESULTS
def latest(name: str) -> list[dict]: p = RESULTS / f"{name}.csv" return list(csv.DictReader(p.open())) if p.exists() else []
def f(x, fmt="{:,.0f}", todo="TODO (run the lesson)"): try: return fmt.format(float(x)) except (TypeError, ValueError): return todo
def main() -> None: ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--customer", default="Acme Cloud (support triage agent)") ap.add_argument("--peak-concurrency", type=int, default=120) ap.add_argument("--tickets-per-month", type=float, default=2e6) ap.add_argument("--slo-ms", type=float, default=1000, help="p95 TTFT target") ap.add_argument("--in-tokens", type=int, default=3200, help="avg prompt tokens per ticket (agent preamble + ticket)") ap.add_argument("--out-tokens", type=int, default=60) ap.add_argument("--serverless-in", type=float, default=0.07, help="$ per 1M input tokens — check pricing page") ap.add_argument("--serverless-out", type=float, default=0.30) ap.add_argument("--cached-discount", type=float, default=0.5, help="fraction off cached input — check pricing") ap.add_argument("--gpu-hour", type=float, default=8.0, help="$ per dedicated GPU-hour — check pricing") a = ap.parse_args()
lat = latest("01-latency") sweep = latest("03-sweep") prefix = latest("06-prefix") struct = latest("07-structured") evals = latest("09-eval") spec = latest("13-spec")
# --- capacity from the sweep: best concurrency under SLO for the most recent non-mock target real = [r for r in sweep if r["target"] != "mock"] or sweep under = [r for r in real if float(r["ttft_p95"]) <= a.slo_ms] per_replica = max((int(r["conc"]) for r in under), default=None) sweep_src = real[-1]["target"] if real else None replicas = math.ceil(a.peak_concurrency / per_replica * 1.2) if per_replica else None # +20% headroom
# --- prefix caching share stable = [r for r in prefix if r["order"] == "stable-first"] speedup = stable[-1]["speedup_x"] if stable else None
# --- serverless monthly cost, with and without caching (assume 90% of input is the shared preamble) n = a.tickets_per_month cost_in = n * a.in_tokens / 1e6 * a.serverless_in cost_out = n * a.out_tokens / 1e6 * a.serverless_out cached_share = 0.9 cost_in_cached = cost_in * (1 - cached_share * a.cached_discount) dedicated = replicas * a.gpu_hour * 730 if replicas else None
# --- quality best_eval = max(evals, key=lambda r: float(r["category_%"])) if evals else None worst_eval = min(evals, key=lambda r: float(r["category_%"])) if evals else None schema = next((r for r in reversed(struct) if r["mode"] == "json_schema"), None) prompt_only = next((r for r in reversed(struct) if r["mode"] == "prompt"), None)
per_ticket = (cost_in_cached + cost_out) / n crossover = f"{dedicated / per_ticket:,.0f}" if dedicated else "TODO" cap_warning = ("- ⚠ Capacity came from a laptop or mock server. Production GPUs with vLLM/SGLang hold far " "more users per replica; re-run `sweep.py` against the Spark or a Fireworks deployment " "before quoting replica counts.\n") if sweep_src in ("mock", "ollama", "llamacpp", "mlx") else ""
# (built outside the f-string so this runs on Python 3.10/3.11 too) lat_lines = "".join( f"- {r['target']} · {r['model']}: TTFT p50 {f(r['ttft_p50_ms'])} ms / p95 {f(r['ttft_p95_ms'])} ms" f" · {f(r['tok_per_s'], '{:.0f}')} tok/s per user (prompt ≈{r['prompt_tok']} tok)\n" for r in lat[-4:]) or "- TODO (run lesson 01)\n" spec_text = ", ".join(f"{r['label']}/{r['traffic']} {f(r['tok_s'])} tok/s" for r in spec[-4:]) or "TODO"
md = f"""# Sizing memo: {a.customer}
*Generated by `lessons/15-capstone/build_memo.py` from measured results. Prices are inputs; verify them on the pricing page.*
## 1. What they need- Peak **{a.peak_concurrency} concurrent** agent sessions, **p95 TTFT ≤ {a.slo_ms:.0f} ms**- **{n:,.0f} tickets/month**, about {a.in_tokens:,} input and {a.out_tokens} output tokens each (the agent preamble dominates)- Output must be **schema-valid JSON** for the ticketing integration
## 2. Latency profile (lesson 01){lat_lines}## 3. Capacity (lesson 03)- One replica holds **{per_replica or 'TODO'}** concurrent users within SLO (measured on `{sweep_src or '—'}`)- Replicas for peak with 20% headroom: **{replicas or 'TODO'}**{cap_warning}## 4. Levers that change the bill| lever | evidence | effect ||---|---|---|| Prefix caching: stable preamble first, `user` for affinity | lesson 06: **{f(speedup, '{:.1f}')}×** lower TTFT on warm calls | about 90% of input tokens billed at the cached rate || Constrained decoding (`json_schema`) | lesson 07: {f(prompt_only and prompt_only['schema_valid_%'], '{:.0f}')}% → **{f(schema and schema['schema_valid_%'], '{:.0f}')}%** valid | no retry loop, no parser failures || Batch API for evals and back-fills | lesson 08 | about 50% off for non-interactive work || LoRA fine-tune | lesson 09–11: category accuracy {f(worst_eval and worst_eval['category_%'], '{:.0f}')}% → **{f(best_eval and best_eval['category_%'], '{:.0f}')}%** ({best_eval['candidate'] if best_eval else '—'}) | a smaller model at higher accuracy; needs a dedicated deployment (multi-LoRA shares it) || Speculative decoding | lesson 13: {spec_text} | pays on structured output; measure α first |
## 5. Cost options (monthly, rough)| option | how | $/month ||---|---|---|| Serverless, no caching | {n:,.0f} × tokens × list price | **${cost_in + cost_out:,.0f}** || Serverless with prefix caching | 90% of input cached at {a.cached_discount:.0%} off | **${cost_in_cached + cost_out:,.0f}** || Dedicated, {replicas or '?'} replicas × ${a.gpu_hour}/GPU-h × 730 h | predictable latency; required for the LoRA | **{('$' + format(dedicated, ',.0f')) if dedicated else 'TODO'}** |
**Crossover:** at about **{crossover}** tickets/month, dedicated costs the same as serverless with caching.Below that, serverless wins. Above it, or when a LoRA or a strict SLO is needed, dedicated wins.
## 6. Recommendation1. **Now:** serverless with the preamble restructured for caching, plus `json_schema`. It is the cheapest path to production.2. **In 4–6 weeks,** if accuracy is the blocker: LoRA on a small base, served on dedicated capacity as a multi-LoRA deployment, justified by the eval table.3. **Always:** nightly evals through the Batch API. Watch p95 TTFT, not averages.
## 7. Risks and what I'd validate next- Traffic shape: burstiness at peak. Replay a real hour of traffic through `sweep.py`.- Quality on real tickets, not synthetic ones: 200 hand-labelled examples, with the LLM judge calibrated against them.- KV headroom if prompts grow (lesson 05): cap the context and consider an FP8 KV cache.""" out = RESULTS / "SIZING_MEMO.md" out.write_text(md) print(md) print(f"\nsaved → {out}")
if __name__ == "__main__": main()