Concurrency sweep
course 5 of 22
lesson 03 · lessons/03-concurrency-sweep
50 min Free Mac
What you will do and why
You measure how many people one server can serve at once before the wait for a reply to start passes the promised limit. That number, not a single speed figure, decides how many servers a customer needs.
Why it matters: Example: on the kit’s practice server (a stand-in with made-up numbers), 95 of 100 requests from 8 people at once get their first token (word piece) within 41.4 ms, thousandths of a second. At 16 people that wait jumps to 3,357 ms, over 3 seconds.
You are done when: The practice run prints holds up to C=8, and results/03-sweep.png shows both runs, practice and your Mac. You have written its test conditions in Your numbers: tokens in, tokens out, people at once and the promise.
Benchmarking Properly · 2:22
Download mp4 (8.9 MB)Chapters
In this video Why one speed number misleads, and the table a field engineer hands a customer instead.
3 key points
Write down the traffic before you test anything.
Four lines: tokens in, tokens out, requests per second, and the speed promise with its percentile (for example p95: the wait 95 out of 100 requests stay under). The video’s example: 1,024 tokens in and 256 out per request, tested with 200 requests at each crowd size.
Test at 1, 2, 4, 8, 16 and 32 people at once, and report the curve, not one number.
In the video’s table, total output climbs from 24 to 264 tokens per second (11 times). Meanwhile the p95 wait for the first token (95 of 100 wait less) grows from 210 ms to 1,850 ms.
The answer is the biggest crowd whose p95 wait stays under the promise. The average hides the problem.
With a 1-second promise the video sizes at 16 people (720 ms). From 8 to 32 people the average wait only went from 290 to 410 ms, while the p95 went from 380 to 1,850 ms, almost 5 times. An average-only report would have called 32 people fine.
Batching & Scheduling · 2:58
Download mp4 (10.4 MB)Chapters
In this video Three ways a server keeps its expensive chip busy: continuous batching, chunked prefill and speculative decoding.
The narrator’s “next video” is the file order. Your next step is below.
3 key points
Continuous batching gives a free seat to the next request the moment a reply finishes, instead of waiting for the whole group.
The video’s 12 requests on 4 seats: the chip is busy 87% of the time instead of 61%. All the work is done at step 17 instead of 24 (each step writes one token for every busy seat).
Chunked prefill reads one huge prompt in pieces, so it cannot freeze everyone else’s replies.
The video’s example is a 100,000-token prompt (about 75,000 words), read in chunks between other people’s writing steps. That prompt’s first token comes a little later; everyone else’s reply keeps appearing word by word.
Speculative decoding cuts waiting: a small model drafts a few tokens and the big model checks them all in one pass, so the output does not change.
Day 9 comes back to it. The video’s real case stacks three fixes (reusing a repeated prompt, chunked prefill, speculation) on the same 2 GPUs (the chips that run the model). The p95 wait for the first token fell from 2.8 s to 0.40 s, 7 times faster. The p95 gap between tokens fell from 48 to 26 ms, the part speculation speeds up.
GPU Bandwidth · 3:10
Download mp4 (12.2 MB)Chapters
In this video Why one person’s writing speed is set by how fast memory delivers the model, and how to predict it on any chip.
3 key points
Top writing speed for one person = memory speed ÷ model size.
Each token needs the whole model fetched once. A 70-billion-number model at 8 bits (1 byte per number) is 70 GB. An H100 (NVIDIA’s data-center GPU) fetches 3,350 GB a second: 3,350 ÷ 70 = about 48 tokens per second.
Real dense models (ones that use every number for each token) reach about 60 to 85% of that limit.
The video’s DGX Spark (NVIDIA’s desktop AI computer) moves 273 GB a second. For an 8-billion-number model at 4 bits, which the video counts as 4.5 GB, the limit is 273 ÷ 4.5 = 61 tokens per second. It measured 38.7: 64% of the limit. Today’s 4-bit file is 4.9 GB, which gives 273 ÷ 4.9 = 56.
Batching is how you put the chip’s idle math to work.
An H100 can do about 989 trillion calculations a second but fetch only 3,350 GB a second: about 300 calculations per byte (989 ÷ 3.35 = 295). Serving one person needs only about 2 per number fetched, so the math sits idle. Serving many people from each fetch puts it to work.
In plain words
Section titled “In plain words”To write each word piece, the server reads the whole model from memory once, and it shares that read among everyone it is serving. So extra people cost little until its seats are full. After that, newcomers wait in line and the wait for the first token jumps. That jump is the knee. The biggest crowd before it is one server’s capacity.
Picture it
A bus takes about as long to drive its route with 1 rider as with 8, so filling seats is nearly free. Once every seat is taken, the next riders wait at the stop, however fast the bus drives. Count the most riders one bus takes with nobody waiting too long. Divide the rush-hour crowd by it to know how many buses to run.
With real numberslesson 03’s practice server: 8 seats, a 500 ms promise, realistic behaviour with made-up numbers
- 1 person: 54.1 tokens (word pieces) per second in total. 95 of 100 get their first token within 27.4 ms (thousandths of a second).
- 8 people: 275.1 in total, 5 times as much (275.1 ÷ 54.1 = 5.1). Each person still gets 38.9 per second, and 95 of 100 get their first token within 41.4 ms.
- 16 people: 294.3 in total, only 7% more than at 8. Now 8 people wait for a seat. 95 of 100 get their first token within 3,357 ms (over 3 seconds), 81 times longer (3,357 ÷ 41.4 = 81).
- So one copy of the server (a replica) holds 8 people within the 500 ms promise.
- If the busiest moment brings 120 people at once (the lesson’s example): 120 ÷ 8 = 15 copies, plus spare ones.
Words to know
- Batching
- Serving several people with one pass over the model. Example: 8 people share each fetch of the model.
- TTFT (time to first token)
- How long until the first token of the reply appears. Example: 41.4 ms (p95) at 8 people.
- p95
- The value 95 out of 100 requests stay under; the slowest 5 are the tail. Example: 3,357 ms at 16 people.
- Goodput
- Requests per second that kept the promise; late ones do not count. Example: 2.5 per second at 8 people, 0.4 at 16.
From the lesson
lessons/03-concurrency-sweep/README.md
What and why
Section titled “What and why”A GPU decode step reads all the weights once, whether it serves 1 user or 16. Batching lets many users share that one read, so total throughput climbs almost for free. Then the batch slots fill up, new requests wait in a queue, and TTFT p95 shoots up. The last concurrency level before that knee is the capacity of one replica. Divide the customer’s peak traffic by it and you have a replica count.
tok/s ▲ ________ aggregate (flattens: slots full) │ / │ / ╱ TTFT p95 (the knee = queueing) │ / ─────────────────── ╱ ─ ─ SLO └──────────────────────────────> concurrent users 1 2 4 8 16 32Goodput counts only the requests that met the SLO. That is the number to sell, not raw tokens per second.
Read the code first
Section titled “Read the code first”sweep.py → run_level()shows that anasyncio.Semaphore(conc)is the concurrency level.- Prompts are unique, so the prefix cache can’t inflate the numbers.
plot_sweep.pyputs both curves on one chart and draws the SLO line.
cd lessons/03-concurrency-sweeppython sweep.py --target mock --slo-ms 500 # mock has 8 slots → knee after C=8python plot_sweep.py && open ../../results/03-sweep.png
# real: restart llama.cpp with 8 slots (context split across them)bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 8python sweep.py --target llamacpp --levels 1,2,4,8,16python plot_sweep.pyWhat each command does
python sweep.py --target mock --slo-ms 500Tests the practice server at 1, 2, 4, 8, 16 and 32 people at once. The promise: 95 of 100 first tokens within 500 ms (half a second). Each crowd size sends 3 requests per person, and at least 12 in all. Look for the line starting
Capacity line:near the end. It saysholds up to C=8(C is people at once): the practice server’s 8 seats.python plot_sweep.py && open ../../results/03-sweep.pngDraws the chart and opens it. Total output is the solid line (left scale). The p95 wait for the first token is the dashed line (right scale). The promise is the red dotted line. Look for the dashed line leaping after 8 while the solid line flattens.
bash ../02-three-local-servers/serve_llamacpp.sh Q4_K_M 16384 8Restarts llama.cpp, Day 2’s real server (the program that runs the model on your Mac), with the same 4-bit model file (Q4_K_M) and 8 seats (slots). Its room for conversations, 16,384 tokens, is split across the seats: 2,048 each, the same as Day 2’s 8,192 over 4. It keeps its terminal busy. Open a second terminal,
cdinto the kit’s03-labs/lessons/03-concurrency-sweepfolder (the../02-three-local-serverspath only works from there), then run it and leave it running. Look forctx=16384andslots=8on its first line.python sweep.py --target llamacpp --levels 1,2,4,8,16The same test against the real model on your Mac, from 1 to 16 people. With no
--slo-ms, it judges against the script’s default promise, 1,000 ms (1 second), which is also the Day 10 sizing memo’s default. Look forttft_p95jumping between the 8 and 16 rows, and theCapacity line:near the end.python plot_sweep.pyRedraws the chart with the practice run and your Mac’s run on the same axes, and saves it again to
results/03-sweep.png(open it as before). The red promise line now shows your Mac run’s promise. Look for the same shape on your Mac: output flattens and the wait jumps once all 8 seats are busy.
What you should see (mock)
Section titled “What you should see (mock)”How to read it
Your practice run prints six rows, 1 to 32 people at once (conc); the lesson shows three. Read across: total output (agg_tok_s), one person’s speed (user_tok_s), the first-token wait in ms, typical (ttft_p50) and 95 of 100 (ttft_p95), and requests per second that kept the promise (goodput_rps). From 1 to 8 people, output grows 5 times while the p95 wait only moves from 27.4 to 41.4 ms. At 16, output barely grows, the wait jumps to 3,357 ms and goodput falls from 2.5 to 0.4. That jump is the knee.
conc agg_tok_s user_tok_s ttft_p50 ttft_p95 goodput_rps 1 54.1 54.9 25.4 27.4 0.5 8 275.1 38.9 26.3 41.4 2.5 ← 5× throughput, still fast 16 294.3 38.9 2,776.3 3,357.0 0.4 ← slots full: queueingCapacity line: TTFT p95 ≤ 500 ms holds up to C=8 concurrent users on this replica.Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
When 8 people used the practice server at once instead of 1, its total output rose about 5 times. Yet each person’s reply only slowed from 55 to 39 tokens (word pieces) per second. Why does sharing cost so little?Show answerHide
In plain words
To write each token, the chip must fetch the whole model from memory, and that fetch is the slow part. One fetch serves all 8 people at once, so each extra person adds only a little.
Picture it
A bus takes about the same time to drive its route with 1 rider or with 8. The drive is the slow part; letting a few more people on adds a little time at each stop. This stops working once the seats run out: then extra riders wait for the next bus, and that point is the knee.
With real numberslesson 03’s practice server: it behaves like a real one, but its numbers are made up for teaching
- 1 person alone: 54.1 tokens per second in total; that person sees about 55 (54.9) once the reply is flowing.
- 8 people: 275.1 in total, about 5 times the 54.1 (275.1 ÷ 54.1 = 5.1).
- Each person: 54.9 falls to 38.9 tokens per second, 29% slower.
- So each person keeps 71% of their solo speed while the server does 5 times the work.
Words to know
- Decode
- The writing phase; the model writes one token at a time and fetches the whole model for each. Example: 54.9 tokens per second for one person on the practice server.
- Memory-bound
- Slowed by how fast memory delivers data, not by how fast the chip calculates. Example: decode.
- Batching
- Serving several people with one pass over the model. Example: 8 people share each fetch.
- Tokens per second (tok/s)
- How many tokens are written each second. Example: 38.9 per person with 8 people.
Go deeper: the engineer version
The kit's question
Throughput went up 5×, but per-user speed only dropped from 55 to 39 tok/s. Why so cheap?
The kit's answer
Decode is memory-bound, so sharing one weight read across 8 users costs little extra.
More detail: Each decode step streams all the weights once, whatever the batch size. Each extra user still adds its own KV-cache reads and a little compute, which is why per-user speed falls somewhat (54.9 to 38.9). The practice server builds this in on purpose: one step takes 18 ms alone and grows 6% per extra user, so 18 x 1.42 = 25.6 ms with 8. That is 1,000 ÷ 18 = 55.6 tokens per second alone and 1,000 ÷ 25.56 = 39.1 with 8 (felab/mock_server.py, _itl). In roofline terms, one user’s decode does about 1 to 2 calculations per byte fetched, while an H100 needs about 300 to keep its math busy; batching raises that ratio. The 1-user total (54.1) sits a little under the per-user figure (54.9) because the per-user figure is the median speed once text starts flowing, while the total divides all tokens by wall-clock time, which includes the wait before the first token.
At the busiest moment, 120 people use the service at once. One copy of the server (a replica) handles 8 people while still keeping its speed promise (the SLO). How many copies do you need?Show answerHide
In plain words
120 ÷ 8 = 15 copies covers the busiest moment exactly. Add spare copies for surprises, so plan for about 18.
Picture it
A ride seats 8 per car, and 120 people want to ride at the same moment. You need 15 cars, plus a few spare in case one breaks down or a bigger crowd shows up.
With real numberslesson 03, and the rule in the Day 10 sizing memo
- Busiest moment: 120 people at once.
- One copy keeps its promise up to 8 people (the practice server’s limit).
- 120 ÷ 8 = 15 copies.
- The lesson’s answer, 18, adds 3 spare copies: 20% extra (15 x 1.2 = 18). The Day 10 sizing memo adds the same 20%.
Words to know
- Replica
- One complete running copy of the server and model. Example: 18 replicas for 120 people.
- SLO (service level objective)
- The speed promise. Example: 95 of 100 requests get their first token within 500 ms (half a second).
- Peak traffic
- The most people using the service at the same moment. Example: 120.
- Headroom
- Spare capacity kept for surprises. Example: plan 18 where 15 are needed.
Go deeper: the engineer version
The kit's question
Peak traffic is 120 concurrent users and one replica handles 8 within SLO. How many replicas?
The kit's answer
15, plus headroom. Say 18.
More detail: Replicas = peak concurrency ÷ per-replica capacity at the SLO, rounded up, plus headroom for traffic spikes, a replica restarting and forecast error. Lesson 15’s build_memo.py computes exactly this, math.ceil(a.peak_concurrency / per_replica * 1.2): 120 ÷ 8 x 1.2 = 18. Measure the per-replica capacity with the customer’s real prompt and reply lengths, because longer contexts move the knee.
On the sweep chart, the knee is the point where all the server’s seats are taken and new people start waiting in line. What moves that point?Show answerHide
In plain words
The knee moves to more people when more fit at once: more seats, more memory set aside for the notes (the KV cache), a smaller model, or shorter conversations. It moves to fewer people when one long prompt hogs the chip while it is read.
Picture it
A restaurant is full at a certain number of guests. More tables, more floor space or shorter meals seat more people; one huge order hogging the kitchen slows every other table.
With real numberslesson 03’s practice server and llama.cpp settings, and lessons 04 and 05 (Day 4)
- Practice server: 8 seats, so the knee comes right after 8 people.
- At 8 people, 95 of 100 got their first token within 41.4 ms (thousandths of a second); at 16, within 3,357 ms (over 3 seconds): about 80 times longer.
- Seats share memory: llama.cpp splits 16,384 tokens of conversation room across 8 seats, 2,048 tokens (about 1,500 words) each. More seats means less room for each conversation.
- Memory for notes sets the seats. Day 4’s calculator: 20 GB of spare memory holds 4 conversations of 32,768 tokens (about 25,000 words) with notes at 16 bits per number, and 8 at 8 bits.
- A smaller model frees memory: Llama 3.1 8B at 4 bits is 4.9 GB against 8.5 GB at 8 bits, 3.6 GB more for notes (Day 4).
Words to know
- Knee
- The point where seats run out and waiting time shoots up. Example: after 8 people on the practice server.
- Slot
- One seat on the server, a conversation it can serve at the same time. Example: the practice server has 8.
- Notes (KV cache)
- The model’s memory of each conversation so far, kept in the same memory as the model. Example: 4 conversations of 32,768 tokens fit in 20 GB.
- Prefill
- The reading phase; the model reads the whole prompt at once before writing. Example: a 100,000-token prompt takes a long prefill.
Go deeper: the engineer version
The kit's question
What changes the knee?
The kit's answer
More slots or memory for KV cache, a smaller or quantized model, shorter contexts, and prefill/decode interference.
More detail: In llama.cpp the -c context is split across the --parallel slots (the comment in serve_llamacpp.sh says so), so slots and context length trade off directly. A quantized model frees memory for KV cache and also makes each decode step faster. Long prefills stall the decode steps of everyone batched with them; chunked prefill (Batching & Scheduling video) is the usual fix. The 41.4 and 3,357 ms figures are TTFT p95.
Explain what you learned
Section titled “Explain what you learned”Question The vendor’s benchmark says this GPU is fast. How many of them do we need for launch?
One clear answer
Single-stream numbers sell hardware; concurrency curves size deployments. I sweep, find the highest concurrency that holds p95 TTFT under the SLO, and size replicas from peak traffic divided by that.
What this means
- “Single-stream numbers sell hardware”: a speed measured with one person at a time makes a chip look good in a brochure, but says nothing about a crowd. Example: the practice server’s 54.9 tokens per second for one person.
- “concurrency curves size deployments”: how many servers to buy comes from a chart of output and waiting time at 1, 2, 4, 8 and 16 people at once.
- “I sweep”: I run the same test at each crowd size in turn, with the customer’s own prompt and reply lengths.
- “find the highest concurrency that holds p95 TTFT under the SLO”: I find the biggest crowd at which 95 of 100 people still get their first token within the promised time. Practice server: 8 people, 41.4 ms against a 500 ms promise; at 16 it is 3,357 ms.
- “and size replicas from peak traffic divided by that”: copies of the server needed = the most people at once ÷ people per copy. 120 ÷ 8 = 15, plus 20% spare = 18.
Your numbersSaved on this device and collected in the Day 3 wrap-up.
Hint: From the sweep’s settings: a 12-word prompt plus a number tag (about 17 tokens of text) in, up to 128 tokens out, 1 to 16 people at once, 3 requests per person, and the promise (1,000 ms unless you set --slo-ms).
Hint: The Capacity line after the llama.cpp sweep, for example holds up to C=8. Write the number of people; the promise is 1,000 ms unless you set --slo-ms.
Hint: The agg_tok_s column of the llama.cpp sweep: the 1 row and your capacity row. Write both, for example 54.1 / 275.1.
Hint: The user_tok_s column, the same two rows. Write both, for example 54.9 / 38.9.
Done when
Section titled “Done when”The practice run prints holds up to C=8, and results/03-sweep.png shows both runs, practice and your Mac. You have written its test conditions in Your numbers: tokens in, tokens out, people at once and the promise.
Stuck?
Section titled “Stuck?”- The first sweep stops with
Connection refused(orConnection error). - The practice server is not running. From the
03-labsfolder, with the Python environment on, runmake mock &, then run the sweep again.make checkshows which servers are up. - The llama.cpp sweep stops with
Connection refused(orConnection error). - llama.cpp is not running yet, or is still loading the model. Start it in its own terminal with the
serve_llamacpp.shline above and wait until it says it is listening.make checkshowsllamacppas UP when it is ready. serve_llamacpp.shsaysNo such file or directory.- The new terminal is in the wrong folder.
cdinto the kit’s03-labs/lessons/03-concurrency-sweepfolder and run the line again. - Your Mac’s chart jumps after 4 people, not 8.
- llama.cpp started with Day 2’s 4 seats. Its first line must say
slots=8: stop it with Ctrl+C and start it again withQ4_K_M 16384 8at the end. plot_sweep.pyprints a text chart, thenopensays03-sweep.pngdoes not exist.- The chart library (matplotlib) is missing, so the script falls back to text. In the
03-labsfolder runsource .venv/bin/activate, thenpip install -r requirements.txt, then plot again. plot_sweep.pysaysNo results yet.- Run a sweep first. Every sweep adds its rows to
03-labs/results/03-sweep.csv, wherever you run it from, and the chart reads them from there.
Code in this step
Section titled “Code in this step”plot_sweep.py Python · 61 lines
"""Lesson 03 — draw the sweep: throughput up, TTFT knee, capacity line.
Reads results/03-sweep.csv (the latest run per target) and writesresults/03-sweep.png. Without matplotlib it prints a text chart instead.
python plot_sweep.py"""import csvfrom pathlib import Path
from felab.results import RESULTS
path = RESULTS / "03-sweep.csv"if not path.exists(): raise SystemExit("No results yet — run sweep.py first.")
rows = list(csv.DictReader(path.open()))# Keep only the most recent sweep for each target (a sweep restarts at conc=1).latest: dict[str, list[dict]] = {}for r in rows: if r["conc"] == "1": latest[r["target"]] = [] latest.setdefault(r["target"], []).append(r)
try: import matplotlib matplotlib.use("Agg") import matplotlib.pyplot as pltexcept ImportError: for tgt, rs in latest.items(): print(f"\n{tgt} (█ = aggregate tok/s, · = TTFT p95 ms)") top = max(float(r["agg_tok_s"]) for r in rs) for r in rs: bar = "█" * int(40 * float(r["agg_tok_s"]) / top) print(f" C={r['conc']:>3} {bar:<40} {float(r['agg_tok_s']):7.0f} tok/s · {float(r['ttft_p95']):6.0f} ms") raise SystemExit
fig, ax1 = plt.subplots(figsize=(9, 5))ax2 = ax1.twinx()for tgt, rs in latest.items(): c = [int(r["conc"]) for r in rs] ax1.plot(c, [float(r["agg_tok_s"]) for r in rs], "-o", label=f"{tgt} aggregate tok/s") ax2.plot(c, [float(r["ttft_p95"]) for r in rs], "--s", label=f"{tgt} TTFT p95") slo = float(rs[0]["slo_ms"])ax2.axhline(slo, color="red", lw=1, ls=":")ax2.text(c[0], slo * 1.03, f"SLO {slo:.0f} ms", color="red", fontsize=9)ax1.set_xscale("log", base=2)ax1.set_xticks(c)ax1.set_xticklabels([str(x) for x in c])ax1.set_xlabel("concurrent requests")ax1.set_ylabel("aggregate tokens / s (solid)")ax2.set_ylabel("TTFT p95, ms (dashed)")ax1.set_title("Throughput rises, then the queue arrives")h1, l1 = ax1.get_legend_handles_labels()h2, l2 = ax2.get_legend_handles_labels()ax1.legend(h1 + h2, l1 + l2, loc="upper left", fontsize=8)out = Path(RESULTS / "03-sweep.png")fig.tight_layout()fig.savefig(out, dpi=130)print(f"saved → {out}")sweep.py Python · 92 lines
"""Lesson 03 — concurrency sweep: the curve that sizes a deployment.
Single-user speed sells hardware. Concurrency curves size deployments.
For each concurrency level C (1, 2, 4, 8, 16, 32) we keep C requests in flight andmeasure: agg tok/s total tokens per second across everyone → goes UP with C (batching) per-user tok/s what one user sees → goes DOWN slowly TTFT p95 tail wait for the first token → flat, then a KNEE (queueing) goodput requests/s that met the SLO → the number to sell
The capacity of one replica = the highest C where TTFT p95 is still under your SLO.
python sweep.py --target mock python sweep.py --target mock --levels 1,2,4,8,16,32 --slo-ms 1000 python sweep.py --target llamacpp # start it with --parallel 8 for a fair fight python plot_sweep.py # draws the chart from results/"""import argparseimport asyncioimport time
from felab import add_target_args, astream_once, banner, client, percentile, record, resolve, table, user
TOPICS = ["caching", "batching", "quantization", "attention", "tokenizers", "GPUs", "speculation", "MoE"]
async def run_level(cli, model: str, conc: int, total: int, max_tokens: int, slo_ms: float) -> dict: """Fire `total` requests with at most `conc` in flight at any moment.""" sem = asyncio.Semaphore(conc)
async def worker(i: int): async with sem: # the semaphore IS the concurrency level # Unique prompts so the prefix cache doesn't flatter us (see lesson 06). prompt = f"#{i} Summarise the idea of {TOPICS[i % len(TOPICS)]} for a new engineer in 80 words." return await astream_once(cli, model, user(prompt), max_tokens=max_tokens)
t0 = time.perf_counter() samples = await asyncio.gather(*(worker(i) for i in range(total))) wall = time.perf_counter() - t0
ttft = [s.ttft_ms for s in samples] tokens = sum(s.tokens for s in samples) per_user = percentile([s.decode_tps for s in samples if s.decode_tps], 50) ok = sum(1 for s in samples if s.ttft_ms <= slo_ms) return { "conc": conc, "agg_tok_s": tokens / wall, "user_tok_s": per_user, "ttft_p50": percentile(ttft, 50), "ttft_p95": percentile(ttft, 95), "goodput_rps": ok / wall, # only requests that met the SLO count }
async def main_async(args) -> None: t = resolve(args) banner(t) cli = client(t, asynchronous=True) await astream_once(cli, t.model, user("warm up"), max_tokens=4)
rows = [] for conc in [int(x) for x in args.levels.split(",")]: total = max(args.min_requests, conc * 3) # enough requests to reach steady state print(f" running C={conc:>3} ({total} requests)…", flush=True) r = await run_level(cli, t.model, conc, total, args.max_tokens, args.slo_ms) rows.append(r) record("03-sweep", {"target": t.name, "model": t.model.split("/")[-1], "slo_ms": args.slo_ms, **r})
print("\n" + table(rows)) fits = [r["conc"] for r in rows if r["ttft_p95"] <= args.slo_ms] if fits: best = max(fits) print(f"\nCapacity line: TTFT p95 ≤ {args.slo_ms:.0f} ms holds up to C={best} concurrent users on this replica.") else: print(f"\nEven C=1 misses the {args.slo_ms:.0f} ms SLO — the model or hardware is too slow for this target.") print("Plot it: python plot_sweep.py")
def main() -> None: ap = add_target_args(argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)) ap.add_argument("--levels", default="1,2,4,8,16,32") ap.add_argument("--slo-ms", type=float, default=1000, help="TTFT p95 target (default 1000)") ap.add_argument("--max-tokens", type=int, default=128) ap.add_argument("--min-requests", type=int, default=12) asyncio.run(main_async(ap.parse_args()))
if __name__ == "__main__": main()