Benchmarking Properly
file 10 · 2:22 · subtitles burned in
Chapters
In this video Why one speed number misleads, and the table a field engineer hands a customer instead.
What it explains: traffic profiles, concurrency sweeps, percentiles
3 key points
Write down the traffic before you test anything.
Four lines: tokens in, tokens out, requests per second, and the speed promise with its percentile (for example p95: the wait 95 out of 100 requests stay under). The video’s example: 1,024 tokens in and 256 out per request, tested with 200 requests at each crowd size.
Test at 1, 2, 4, 8, 16 and 32 people at once, and report the curve, not one number.
In the video’s table, total output climbs from 24 to 264 tokens per second (11 times). Meanwhile the p95 wait for the first token (95 of 100 wait less) grows from 210 ms to 1,850 ms.
The answer is the biggest crowd whose p95 wait stays under the promise. The average hides the problem.
With a 1-second promise the video sizes at 16 people (720 ms). From 8 to 32 people the average wait only went from 290 to 410 ms, while the p95 went from 380 to 1,850 ms, almost 5 times. An average-only report would have called 32 people fine.
Used in the course
The title card in the video says “Video 10 of 18”: that is the file order. The course plays the videos in the order of its days.