Step 1 · Sizing memo and 10-minute talk
Question A new customer asks: should we run this on serverless or on dedicated GPUs? How would you approach it?
One clear answer
I’d start with their traffic shape and SLO, measure a concurrency curve on the candidate model, apply the cheap levers first (caching and structured output), and only then decide between serverless and dedicated, with the crossover point written down.
What this means
- “I’d start with their traffic shape and SLO”: Before any test, I write down how much they send and when (the traffic shape) and their speed promise (the SLO, service level objective). The memo’s example: 120 people at the busiest moment, 3,200 tokens in and 60 out per ticket, 2 million tickets a month, and 95 of 100 first words within 1,000 ms.
- “measure a concurrency curve on the candidate model”: I run Day 3’s test on the model I would propose: raise the number of people using it at the same time until the first word comes too late. On the practice server, with 8 people, 95 of 100 got their first word within 41.4 ms (thousandths of a second). With 16, 95 of 100 got it within 3,357 ms, over 3 seconds: about 80 times longer, because requests were waiting in line.
- “apply the cheap levers first (caching and structured output)”: I use the fixes that need no new hardware before buying any: reuse the repeated opening of each prompt (Day 5) and force well-formed replies so nothing is retried (Day 6). If 90% of each prompt is cached at half price, as the memo assumes, the bill falls from $484 to $282 a month.
- “only then decide between serverless and dedicated”: Then I choose between paying per token on shared models (serverless) and renting GPUs by the hour (dedicated), using the numbers above.
- “with the crossover point written down”: The memo states the monthly volume at which the two cost the same, so the customer knows when to look again. With the practice server’s capacity: about 744 million tickets a month, 372 times the example’s 2 million. That keeps today’s 18 copies fixed, and they could not carry 372 times the traffic, so read it as: at this capacity, dedicated does not pay. Serverless for now.