Skip to content

Prefix Caching

file 11 · 2:37 · subtitles burned in

Chapters

In this video Why agent traffic is mostly repeats, how servers reuse the repeated part, and what that does to the wait and the bill.

What it explains: shared prefixes, RadixAttention, prompt caching

3 key points

  1. Most of what an agent sends is the same opening again.

    In the video’s 20-call example, 66,500 of the 72,000 tokens sent repeat the opening: 92%. The video shows 93% because it rounds 66,500 up to 67,000.

  2. Reusing the opening cuts both the wait and the bill.

    In the video’s session, 95 of 100 calls see their first word within 0.35 seconds instead of 1.9. The input costs about a fifth of a cent instead of 2 cents.

  3. Put the parts that never change first, or there is nothing to reuse.

    One changing line at the top, such as the time, makes every call new: calls 2 to 12 took 712 ms instead of 36 ms in the lesson’s sample.

The title card in the video says “Video 11 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (9.6 MB)