Skip to content

Concurrency sweep

course 5 of 22

lesson 03 · lessons/03-concurrency-sweep

50 min Free Mac

What you will do and why

You measure how many people one server can serve at once before the wait for a reply to start passes the promised limit. That number, not a single speed figure, decides how many servers a customer needs.

Why it matters: Example: on the kit’s practice server (a stand-in with made-up numbers), 95 of 100 requests from 8 people at once get their first token (word piece) within 41.4 ms, thousandths of a second. At 16 people that wait jumps to 3,357 ms, over 3 seconds.

You are done when: The practice run prints holds up to C=8, and results/03-sweep.png shows both runs, practice and your Mac. You have written its test conditions in Your numbers: tokens in, tokens out, people at once and the promise.

Benchmarking Properly · 2:22

Download mp4 (8.9 MB)

Chapters

In this video Why one speed number misleads, and the table a field engineer hands a customer instead.

3 key points

  1. Write down the traffic before you test anything.

    Four lines: tokens in, tokens out, requests per second, and the speed promise with its percentile (for example p95: the wait 95 out of 100 requests stay under). The video’s example: 1,024 tokens in and 256 out per request, tested with 200 requests at each crowd size.

  2. Test at 1, 2, 4, 8, 16 and 32 people at once, and report the curve, not one number.

    In the video’s table, total output climbs from 24 to 264 tokens per second (11 times). Meanwhile the p95 wait for the first token (95 of 100 wait less) grows from 210 ms to 1,850 ms.

  3. The answer is the biggest crowd whose p95 wait stays under the promise. The average hides the problem.

    With a 1-second promise the video sizes at 16 people (720 ms). From 8 to 32 people the average wait only went from 290 to 410 ms, while the p95 went from 380 to 1,850 ms, almost 5 times. An average-only report would have called 32 people fine.

Batching & Scheduling · 2:58

Download mp4 (10.4 MB)

Chapters

In this video Three ways a server keeps its expensive chip busy: continuous batching, chunked prefill and speculative decoding.

The narrator’s “next video” is the file order. Your next step is below.

3 key points

  1. Continuous batching gives a free seat to the next request the moment a reply finishes, instead of waiting for the whole group.

    The video’s 12 requests on 4 seats: the chip is busy 87% of the time instead of 61%. All the work is done at step 17 instead of 24 (each step writes one token for every busy seat).

  2. Chunked prefill reads one huge prompt in pieces, so it cannot freeze everyone else’s replies.

    The video’s example is a 100,000-token prompt (about 75,000 words), read in chunks between other people’s writing steps. That prompt’s first token comes a little later; everyone else’s reply keeps appearing word by word.

  3. Speculative decoding cuts waiting: a small model drafts a few tokens and the big model checks them all in one pass, so the output does not change.

    Day 9 comes back to it. The video’s real case stacks three fixes (reusing a repeated prompt, chunked prefill, speculation) on the same 2 GPUs (the chips that run the model). The p95 wait for the first token fell from 2.8 s to 0.40 s, 7 times faster. The p95 gap between tokens fell from 48 to 26 ms, the part speculation speeds up.

GPU Bandwidth · 3:10

Download mp4 (12.2 MB)

Chapters

In this video Why one person’s writing speed is set by how fast memory delivers the model, and how to predict it on any chip.

3 key points

  1. Top writing speed for one person = memory speed ÷ model size.

    Each token needs the whole model fetched once. A 70-billion-number model at 8 bits (1 byte per number) is 70 GB. An H100 (NVIDIA’s data-center GPU) fetches 3,350 GB a second: 3,350 ÷ 70 = about 48 tokens per second.

  2. Real dense models (ones that use every number for each token) reach about 60 to 85% of that limit.

    The video’s DGX Spark (NVIDIA’s desktop AI computer) moves 273 GB a second. For an 8-billion-number model at 4 bits, which the video counts as 4.5 GB, the limit is 273 ÷ 4.5 = 61 tokens per second. It measured 38.7: 64% of the limit. Today’s 4-bit file is 4.9 GB, which gives 273 ÷ 4.9 = 56.

  3. Batching is how you put the chip’s idle math to work.

    An H100 can do about 989 trillion calculations a second but fetch only 3,350 GB a second: about 300 calculations per byte (989 ÷ 3.35 = 295). Serving one person needs only about 2 per number fetched, so the math sits idle. Serving many people from each fetch puts it to work.

To write each word piece, the server reads the whole model from memory once, and it shares that read among everyone it is serving. So extra people cost little until its seats are full. After that, newcomers wait in line and the wait for the first token jumps. That jump is the knee. The biggest crowd before it is one server’s capacity.

Picture it

A bus takes about as long to drive its route with 1 rider as with 8, so filling seats is nearly free. Once every seat is taken, the next riders wait at the stop, however fast the bus drives. Count the most riders one bus takes with nobody waiting too long. Divide the rush-hour crowd by it to know how many buses to run.

With real numberslesson 03’s practice server: 8 seats, a 500 ms promise, realistic behaviour with made-up numbers

  • 1 person: 54.1 tokens (word pieces) per second in total. 95 of 100 get their first token within 27.4 ms (thousandths of a second).
  • 8 people: 275.1 in total, 5 times as much (275.1 ÷ 54.1 = 5.1). Each person still gets 38.9 per second, and 95 of 100 get their first token within 41.4 ms.
  • 16 people: 294.3 in total, only 7% more than at 8. Now 8 people wait for a seat. 95 of 100 get their first token within 3,357 ms (over 3 seconds), 81 times longer (3,357 ÷ 41.4 = 81).
  • So one copy of the server (a replica) holds 8 people within the 500 ms promise.
  • If the busiest moment brings 120 people at once (the lesson’s example): 120 ÷ 8 = 15 copies, plus spare ones.

Words to know

Batching
Serving several people with one pass over the model. Example: 8 people share each fetch of the model.
TTFT (time to first token)
How long until the first token of the reply appears. Example: 41.4 ms (p95) at 8 people.
p95
The value 95 out of 100 requests stay under; the slowest 5 are the tail. Example: 3,357 ms at 16 people.
Goodput
Requests per second that kept the promise; late ones do not count. Example: 2.5 per second at 8 people, 0.4 at 16.

From the lesson

lessons/03-concurrency-sweep/README.md

A GPU decode step reads all the weights once, whether it serves 1 user or 16. Batching lets many users share that one read, so total throughput climbs almost for free. Then the batch slots fill up, new requests wait in a queue, and TTFT p95 shoots up. The last concurrency level before that knee is the capacity of one replica. Divide the customer’s peak traffic by it and you have a replica count.

tok/s ▲ ________ aggregate (flattens: slots full)
│ /
│ / ╱ TTFT p95 (the knee = queueing)
│ / ─────────────────── ╱ ─ ─ SLO
└──────────────────────────────> concurrent users
1 2 4 8 16 32

Goodput counts only the requests that met the SLO. That is the number to sell, not raw tokens per second.

  • sweep.py → run_level() shows that an asyncio.Semaphore(conc) is the concurrency level.
  • Prompts are unique, so the prefix cache can’t inflate the numbers.
  • plot_sweep.py puts both curves on one chart and draws the SLO line.
Terminal window
cd lessons/03-concurrency-sweep
python sweep.py --target mock --slo-ms 500 # mock has 8 slots → knee after C=8
python plot_sweep.py && open ../../results/03-sweep.png
# real: restart llama.cpp with 8 slots (context split across them)
bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 8
python sweep.py --target llamacpp --levels 1,2,4,8,16
python plot_sweep.py

What each command does

  1. python sweep.py --target mock --slo-ms 500

    Tests the practice server at 1, 2, 4, 8, 16 and 32 people at once. The promise: 95 of 100 first tokens within 500 ms (half a second). Each crowd size sends 3 requests per person, and at least 12 in all. Look for the line starting Capacity line: near the end. It says holds up to C=8 (C is people at once): the practice server’s 8 seats.

  2. python plot_sweep.py && open ../../results/03-sweep.png

    Draws the chart and opens it. Total output is the solid line (left scale). The p95 wait for the first token is the dashed line (right scale). The promise is the red dotted line. Look for the dashed line leaping after 8 while the solid line flattens.

  3. bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 8

    Restarts llama.cpp, Day 2’s real server (the program that runs the model on your Mac), with the same 4-bit model file (Q4_K_M) and 8 seats (slots). Its room for conversations, 16,384 tokens, is split across the seats: 2,048 each, the same as Day 2’s 8,192 over 4. It keeps its terminal busy. Open a second terminal, cd into the kit’s 03-labs/lessons/03-concurrency-sweep folder (the ../02-three-local-servers path only works from there), then run it and leave it running. Look for ctx=16384 and slots=8 on its first line.

  4. python sweep.py --target llamacpp --levels 1,2,4,8,16

    The same test against the real model on your Mac, from 1 to 16 people. With no --slo-ms, it judges against the script’s default promise, 1,000 ms (1 second), which is also the Day 10 sizing memo’s default. Look for ttft_p95 jumping between the 8 and 16 rows, and the Capacity line: near the end.

  5. python plot_sweep.py

    Redraws the chart with the practice run and your Mac’s run on the same axes, and saves it again to results/03-sweep.png (open it as before). The red promise line now shows your Mac run’s promise. Look for the same shape on your Mac: output flattens and the wait jumps once all 8 seats are busy.

How to read it

Your practice run prints six rows, 1 to 32 people at once (conc); the lesson shows three. Read across: total output (agg_tok_s), one person’s speed (user_tok_s), the first-token wait in ms, typical (ttft_p50) and 95 of 100 (ttft_p95), and requests per second that kept the promise (goodput_rps). From 1 to 8 people, output grows 5 times while the p95 wait only moves from 27.4 to 41.4 ms. At 16, output barely grows, the wait jumps to 3,357 ms and goodput falls from 2.5 to 0.4. That jump is the knee.

conc agg_tok_s user_tok_s ttft_p50 ttft_p95 goodput_rps
1 54.1 54.9 25.4 27.4 0.5
8 275.1 38.9 26.3 41.4 2.5 ← 5× throughput, still fast
16 294.3 38.9 2,776.3 3,357.0 0.4 ← slots full: queueing
Capacity line: TTFT p95 ≤ 500 ms holds up to C=8 concurrent users on this replica.

3 questions. Say your answer out loud, then tap to check it.

When 8 people used the practice server at once instead of 1, its total output rose about 5 times. Yet each person’s reply only slowed from 55 to 39 tokens (word pieces) per second. Why does sharing cost so little?Show answerHide

In plain words

To write each token, the chip must fetch the whole model from memory, and that fetch is the slow part. One fetch serves all 8 people at once, so each extra person adds only a little.

Picture it

A bus takes about the same time to drive its route with 1 rider or with 8. The drive is the slow part; letting a few more people on adds a little time at each stop. This stops working once the seats run out: then extra riders wait for the next bus, and that point is the knee.

With real numberslesson 03’s practice server: it behaves like a real one, but its numbers are made up for teaching

  • 1 person alone: 54.1 tokens per second in total; that person sees about 55 (54.9) once the reply is flowing.
  • 8 people: 275.1 in total, about 5 times the 54.1 (275.1 ÷ 54.1 = 5.1).
  • Each person: 54.9 falls to 38.9 tokens per second, 29% slower.
  • So each person keeps 71% of their solo speed while the server does 5 times the work.

Words to know

Decode
The writing phase; the model writes one token at a time and fetches the whole model for each. Example: 54.9 tokens per second for one person on the practice server.
Memory-bound
Slowed by how fast memory delivers data, not by how fast the chip calculates. Example: decode.
Batching
Serving several people with one pass over the model. Example: 8 people share each fetch.
Tokens per second (tok/s)
How many tokens are written each second. Example: 38.9 per person with 8 people.
Go deeper: the engineer version

The kit's question

Throughput went up 5×, but per-user speed only dropped from 55 to 39 tok/s. Why so cheap?

The kit's answer

Decode is memory-bound, so sharing one weight read across 8 users costs little extra.

More detail: Each decode step streams all the weights once, whatever the batch size. Each extra user still adds its own KV-cache reads and a little compute, which is why per-user speed falls somewhat (54.9 to 38.9). The practice server builds this in on purpose: one step takes 18 ms alone and grows 6% per extra user, so 18 x 1.42 = 25.6 ms with 8. That is 1,000 ÷ 18 = 55.6 tokens per second alone and 1,000 ÷ 25.56 = 39.1 with 8 (felab/mock_server.py, _itl). In roofline terms, one user’s decode does about 1 to 2 calculations per byte fetched, while an H100 needs about 300 to keep its math busy; batching raises that ratio. The 1-user total (54.1) sits a little under the per-user figure (54.9) because the per-user figure is the median speed once text starts flowing, while the total divides all tokens by wall-clock time, which includes the wait before the first token.

At the busiest moment, 120 people use the service at once. One copy of the server (a replica) handles 8 people while still keeping its speed promise (the SLO). How many copies do you need?Show answerHide

In plain words

120 ÷ 8 = 15 copies covers the busiest moment exactly. Add spare copies for surprises, so plan for about 18.

Picture it

A ride seats 8 per car, and 120 people want to ride at the same moment. You need 15 cars, plus a few spare in case one breaks down or a bigger crowd shows up.

With real numberslesson 03, and the rule in the Day 10 sizing memo

  • Busiest moment: 120 people at once.
  • One copy keeps its promise up to 8 people (the practice server’s limit).
  • 120 ÷ 8 = 15 copies.
  • The lesson’s answer, 18, adds 3 spare copies: 20% extra (15 x 1.2 = 18). The Day 10 sizing memo adds the same 20%.

Words to know

Replica
One complete running copy of the server and model. Example: 18 replicas for 120 people.
SLO (service level objective)
The speed promise. Example: 95 of 100 requests get their first token within 500 ms (half a second).
Peak traffic
The most people using the service at the same moment. Example: 120.
Headroom
Spare capacity kept for surprises. Example: plan 18 where 15 are needed.
Go deeper: the engineer version

The kit's question

Peak traffic is 120 concurrent users and one replica handles 8 within SLO. How many replicas?

The kit's answer

15, plus headroom. Say 18.

More detail: Replicas = peak concurrency ÷ per-replica capacity at the SLO, rounded up, plus headroom for traffic spikes, a replica restarting and forecast error. Lesson 15’s build_memo.py computes exactly this, math.ceil(a.peak_concurrency / per_replica * 1.2): 120 ÷ 8 x 1.2 = 18. Measure the per-replica capacity with the customer’s real prompt and reply lengths, because longer contexts move the knee.

On the sweep chart, the knee is the point where all the server’s seats are taken and new people start waiting in line. What moves that point?Show answerHide

In plain words

The knee moves to more people when more fit at once: more seats, more memory set aside for the notes (the KV cache), a smaller model, or shorter conversations. It moves to fewer people when one long prompt hogs the chip while it is read.

Picture it

A restaurant is full at a certain number of guests. More tables, more floor space or shorter meals seat more people; one huge order hogging the kitchen slows every other table.

With real numberslesson 03’s practice server and llama.cpp settings, and lessons 04 and 05 (Day 4)

  • Practice server: 8 seats, so the knee comes right after 8 people.
  • At 8 people, 95 of 100 got their first token within 41.4 ms (thousandths of a second); at 16, within 3,357 ms (over 3 seconds): about 80 times longer.
  • Seats share memory: llama.cpp splits 16,384 tokens of conversation room across 8 seats, 2,048 tokens (about 1,500 words) each. More seats means less room for each conversation.
  • Memory for notes sets the seats. Day 4’s calculator: 20 GB of spare memory holds 4 conversations of 32,768 tokens (about 25,000 words) with notes at 16 bits per number, and 8 at 8 bits.
  • A smaller model frees memory: Llama 3.1 8B at 4 bits is 4.9 GB against 8.5 GB at 8 bits, 3.6 GB more for notes (Day 4).

Words to know

Knee
The point where seats run out and waiting time shoots up. Example: after 8 people on the practice server.
Slot
One seat on the server, a conversation it can serve at the same time. Example: the practice server has 8.
Notes (KV cache)
The model’s memory of each conversation so far, kept in the same memory as the model. Example: 4 conversations of 32,768 tokens fit in 20 GB.
Prefill
The reading phase; the model reads the whole prompt at once before writing. Example: a 100,000-token prompt takes a long prefill.
Go deeper: the engineer version

The kit's question

What changes the knee?

The kit's answer

More slots or memory for KV cache, a smaller or quantized model, shorter contexts, and prefill/decode interference.

More detail: In llama.cpp the -c context is split across the --parallel slots (the comment in serve_llamacpp.sh says so), so slots and context length trade off directly. A quantized model frees memory for KV cache and also makes each decode step faster. Long prefills stall the decode steps of everyone batched with them; chunked prefill (Batching & Scheduling video) is the usual fix. The 41.4 and 3,357 ms figures are TTFT p95.

Question The vendor’s benchmark says this GPU is fast. How many of them do we need for launch?

One clear answer

Single-stream numbers sell hardware; concurrency curves size deployments. I sweep, find the highest concurrency that holds p95 TTFT under the SLO, and size replicas from peak traffic divided by that.

What this means

  • “Single-stream numbers sell hardware”: a speed measured with one person at a time makes a chip look good in a brochure, but says nothing about a crowd. Example: the practice server’s 54.9 tokens per second for one person.
  • “concurrency curves size deployments”: how many servers to buy comes from a chart of output and waiting time at 1, 2, 4, 8 and 16 people at once.
  • “I sweep”: I run the same test at each crowd size in turn, with the customer’s own prompt and reply lengths.
  • “find the highest concurrency that holds p95 TTFT under the SLO”: I find the biggest crowd at which 95 of 100 people still get their first token within the promised time. Practice server: 8 people, 41.4 ms against a 500 ms promise; at 16 it is 3,357 ms.
  • “and size replicas from peak traffic divided by that”: copies of the server needed = the most people at once ÷ people per copy. 120 ÷ 8 = 15, plus 20% spare = 18.

Your numbersSaved on this device and collected in the Day 3 wrap-up.

Hint: From the sweep’s settings: a 12-word prompt plus a number tag (about 17 tokens of text) in, up to 128 tokens out, 1 to 16 people at once, 3 requests per person, and the promise (1,000 ms unless you set --slo-ms).

Hint: The Capacity line after the llama.cpp sweep, for example holds up to C=8. Write the number of people; the promise is 1,000 ms unless you set --slo-ms.

Hint: The agg_tok_s column of the llama.cpp sweep: the 1 row and your capacity row. Write both, for example 54.1 / 275.1.

Hint: The user_tok_s column, the same two rows. Write both, for example 54.9 / 38.9.

Hint: Divide user_tok_s at 1 person (llama.cpp run) by the limit make check prints on its bandwidth line (about 56 tok/s on an M4 Pro), then multiply by 100. Expect 60 to 85%. M3 Max and M4 Max chips come in two memory speeds and make check assumes the faster one, so a low share may mean you have the slower one.

The practice run prints holds up to C=8, and results/03-sweep.png shows both runs, practice and your Mac. You have written its test conditions in Your numbers: tokens in, tokens out, people at once and the promise.

The first sweep stops with Connection refused (or Connection error).
The practice server is not running. From the 03-labs folder, with the Python environment on, run make mock &, then run the sweep again. make check shows which servers are up.
The llama.cpp sweep stops with Connection refused (or Connection error).
llama.cpp is not running yet, or is still loading the model. Start it in its own terminal with the serve_llamacpp.sh line above and wait until it says it is listening. make check shows llamacpp as UP when it is ready.
serve_llamacpp.sh says No such file or directory.
The new terminal is in the wrong folder. cd into the kit’s 03-labs/lessons/03-concurrency-sweep folder and run the line again.
Your Mac’s chart jumps after 4 people, not 8.
llama.cpp started with Day 2’s 4 seats. Its first line must say slots=8: stop it with Ctrl+C and start it again with Q4_K_M 16384 8 at the end.
plot_sweep.py prints a text chart, then open says 03-sweep.png does not exist.
The chart library (matplotlib) is missing, so the script falls back to text. In the 03-labs folder run source .venv/bin/activate, then pip install -r requirements.txt, then plot again.
plot_sweep.py says No results yet.
Run a sweep first. Every sweep adds its rows to 03-labs/results/03-sweep.csv, wherever you run it from, and the chart reads them from there.
plot_sweep.py Python · 61 lines
"""
Lesson 03 — draw the sweep: throughput up, TTFT knee, capacity line.
Reads results/03-sweep.csv (the latest run per target) and writes
results/03-sweep.png. Without matplotlib it prints a text chart instead.
python plot_sweep.py
"""
import csv
from pathlib import Path
from felab.results import RESULTS
path = RESULTS / "03-sweep.csv"
if not path.exists():
raise SystemExit("No results yet — run sweep.py first.")
rows = list(csv.DictReader(path.open()))
# Keep only the most recent sweep for each target (a sweep restarts at conc=1).
latest: dict[str, list[dict]] = {}
for r in rows:
if r["conc"] == "1":
latest[r["target"]] = []
latest.setdefault(r["target"], []).append(r)
try:
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt
except ImportError:
for tgt, rs in latest.items():
print(f"\n{tgt} (█ = aggregate tok/s, · = TTFT p95 ms)")
top = max(float(r["agg_tok_s"]) for r in rs)
for r in rs:
bar = "█" * int(40 * float(r["agg_tok_s"]) / top)
print(f" C={r['conc']:>3} {bar:<40} {float(r['agg_tok_s']):7.0f} tok/s · {float(r['ttft_p95']):6.0f} ms")
raise SystemExit
fig, ax1 = plt.subplots(figsize=(9, 5))
ax2 = ax1.twinx()
for tgt, rs in latest.items():
c = [int(r["conc"]) for r in rs]
ax1.plot(c, [float(r["agg_tok_s"]) for r in rs], "-o", label=f"{tgt} aggregate tok/s")
ax2.plot(c, [float(r["ttft_p95"]) for r in rs], "--s", label=f"{tgt} TTFT p95")
slo = float(rs[0]["slo_ms"])
ax2.axhline(slo, color="red", lw=1, ls=":")
ax2.text(c[0], slo * 1.03, f"SLO {slo:.0f} ms", color="red", fontsize=9)
ax1.set_xscale("log", base=2)
ax1.set_xticks(c)
ax1.set_xticklabels([str(x) for x in c])
ax1.set_xlabel("concurrent requests")
ax1.set_ylabel("aggregate tokens / s (solid)")
ax2.set_ylabel("TTFT p95, ms (dashed)")
ax1.set_title("Throughput rises, then the queue arrives")
h1, l1 = ax1.get_legend_handles_labels()
h2, l2 = ax2.get_legend_handles_labels()
ax1.legend(h1 + h2, l1 + l2, loc="upper left", fontsize=8)
out = Path(RESULTS / "03-sweep.png")
fig.tight_layout()
fig.savefig(out, dpi=130)
print(f"saved → {out}")
sweep.py Python · 92 lines
"""
Lesson 03 — concurrency sweep: the curve that sizes a deployment.
Single-user speed sells hardware. Concurrency curves size deployments.
For each concurrency level C (1, 2, 4, 8, 16, 32) we keep C requests in flight and
measure:
agg tok/s total tokens per second across everyone → goes UP with C (batching)
per-user tok/s what one user sees → goes DOWN slowly
TTFT p95 tail wait for the first token → flat, then a KNEE (queueing)
goodput requests/s that met the SLO → the number to sell
The capacity of one replica = the highest C where TTFT p95 is still under your SLO.
python sweep.py --target mock
python sweep.py --target mock --levels 1,2,4,8,16,32 --slo-ms 1000
python sweep.py --target llamacpp # start it with --parallel 8 for a fair fight
python plot_sweep.py # draws the chart from results/
"""
import argparse
import asyncio
import time
from felab import add_target_args, astream_once, banner, client, percentile, record, resolve, table, user
TOPICS = ["caching", "batching", "quantization", "attention", "tokenizers", "GPUs", "speculation", "MoE"]
async def run_level(cli, model: str, conc: int, total: int, max_tokens: int, slo_ms: float) -> dict:
"""Fire `total` requests with at most `conc` in flight at any moment."""
sem = asyncio.Semaphore(conc)
async def worker(i: int):
async with sem: # the semaphore IS the concurrency level
# Unique prompts so the prefix cache doesn't flatter us (see lesson 06).
prompt = f"#{i} Summarise the idea of {TOPICS[i % len(TOPICS)]} for a new engineer in 80 words."
return await astream_once(cli, model, user(prompt), max_tokens=max_tokens)
t0 = time.perf_counter()
samples = await asyncio.gather(*(worker(i) for i in range(total)))
wall = time.perf_counter() - t0
ttft = [s.ttft_ms for s in samples]
tokens = sum(s.tokens for s in samples)
per_user = percentile([s.decode_tps for s in samples if s.decode_tps], 50)
ok = sum(1 for s in samples if s.ttft_ms <= slo_ms)
return {
"conc": conc,
"agg_tok_s": tokens / wall,
"user_tok_s": per_user,
"ttft_p50": percentile(ttft, 50),
"ttft_p95": percentile(ttft, 95),
"goodput_rps": ok / wall, # only requests that met the SLO count
}
async def main_async(args) -> None:
t = resolve(args)
banner(t)
cli = client(t, asynchronous=True)
await astream_once(cli, t.model, user("warm up"), max_tokens=4)
rows = []
for conc in [int(x) for x in args.levels.split(",")]:
total = max(args.min_requests, conc * 3) # enough requests to reach steady state
print(f" running C={conc:>3} ({total} requests)…", flush=True)
r = await run_level(cli, t.model, conc, total, args.max_tokens, args.slo_ms)
rows.append(r)
record("03-sweep", {"target": t.name, "model": t.model.split("/")[-1], "slo_ms": args.slo_ms, **r})
print("\n" + table(rows))
fits = [r["conc"] for r in rows if r["ttft_p95"] <= args.slo_ms]
if fits:
best = max(fits)
print(f"\nCapacity line: TTFT p95 ≤ {args.slo_ms:.0f} ms holds up to C={best} concurrent users on this replica.")
else:
print(f"\nEven C=1 misses the {args.slo_ms:.0f} ms SLO — the model or hardware is too slow for this target.")
print("Plot it: python plot_sweep.py")
def main() -> None:
ap = add_target_args(argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter))
ap.add_argument("--levels", default="1,2,4,8,16,32")
ap.add_argument("--slo-ms", type=float, default=1000, help="TTFT p95 target (default 1000)")
ap.add_argument("--max-tokens", type=int, default=128)
ap.add_argument("--min-requests", type=int, default=12)
asyncio.run(main_async(ap.parse_args()))
if __name__ == "__main__":
main()