Skip to content

The KV Cache

file 02 · 3:02 · subtitles burned in

Chapters

In this video Why the model’s notes on each conversation, not the model itself, decide how many people one GPU (the chip that runs the model) can serve, and two cheap ways to fit more.

What it explains: KV cache maths, PagedAttention, FP8 KV

3 key points

  1. Every token adds the same amount of notes, and the model’s settings file tells you how much.

    For Llama 3.1 8B it is 128 KB per token. A 4,096-token chat (about 3,000 words) needs 128 KB x 4,096 = 512 MiB of notes, about half a GB (MiB: the unit llama.cpp prints, about a million bytes).

  2. Some servers set aside room for the longest possible conversation up front. vLLM (a server program for NVIDIA chips) uses PagedAttention instead: it hands out memory a small page at a time as each conversation grows.

    The video says the same chip then serves 2 to 4 times more users.

  3. Storing the notes in 8 bits instead of 16 about halves them (to 53% in llama.cpp’s 8-bit format, q8_0, which stores a little extra), so about twice as many people fit.

    In 20 GB of spare memory: 4 conversations of 32,768 tokens with 16-bit notes, 8 with 8-bit notes.

The title card in the video says “Video 2 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (11.1 MB)