Skip to content

Inference 101

file 01 · 3:02 · subtitles burned in

Chapters

In this video Why every AI reply has two phases, reading the prompt and writing the answer, and the numbers customers feel, starting with the wait for the first word and the pace after it.

What it explains: prefill vs decode, TTFT, ITL, throughput, goodput

3 key points

  1. The model reads your whole prompt in one pass, then writes its answer one token (word piece) at a time.

    Reading is called prefill; writing is called decode. In the video’s support assistant, reading 6,200 tokens takes 0.78 seconds and writing 300 tokens takes 6.6 seconds.

  2. The wait for the first word comes from reading and from waiting in line. The pace after that comes from writing.

    These are TTFT (time to first token) and ITL (inter-token latency, the gap between tokens). At 22 ms (thousandths of a second) per token, a reply types about 45 tokens a second: 1,000 ms ÷ 22 ms = 45.

  3. Before you measure anything, agree with the customer on the request’s shape and on which percentile counts (for example p95: the time 95 of 100 requests stay under).

    The shape is how many tokens go in and come out. A support assistant’s 6,200-token prompt takes 0.78 s to read, and reusing its unchanging part (prompt caching) cuts that to about 0.12 s. A code completion (an editor suggesting the next lines of code) reads its 200 tokens in 25 ms, so the same trick saves almost nothing there.

The title card in the video says “Video 1 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (11.1 MB)