Skip to content

KV cache and context length

course 7 of 22

lesson 05 · lessons/05-kv-cache

50 min Free Mac

What you will do and why

You watch the model’s notes on one long conversation (the KV cache) outgrow the model itself, then halve them by storing them in 8 bits. Running out of memory for these notes is the most common failure in real AI services.

Why it matters: One 131,072-token conversation (about 98,000 words) needs 17.18 GB of notes at 16 bits, and 9.13 GB at 8 bits.

You are done when: The calculator says about 17 GB (17.18) for one 131,072-token conversation (about 98,000 words). And the real server shows the 8-bit notes at about half the 16-bit size, for example 2,176 MiB (2.28 GB) against 4,096 MiB (4.29 GB) at 32,768 tokens.

The KV Cache · 3:02

Download mp4 (11.1 MB)

Chapters

In this video Why the model’s notes on each conversation, not the model itself, decide how many people one GPU (the chip that runs the model) can serve, and two cheap ways to fit more.

The narrator’s “next video” is the file order. Your next step is below.

3 key points

  1. Every token adds the same amount of notes, and the model’s settings file tells you how much.

    For Llama 3.1 8B it is 128 KB per token. A 4,096-token chat (about 3,000 words) needs 128 KB x 4,096 = 512 MiB of notes, about half a GB (MiB: the unit llama.cpp prints, about a million bytes).

  2. Some servers set aside room for the longest possible conversation up front. vLLM (a server program for NVIDIA chips) uses PagedAttention instead: it hands out memory a small page at a time as each conversation grows.

    The video says the same chip then serves 2 to 4 times more users.

  3. Storing the notes in 8 bits instead of 16 about halves them (to 53% in llama.cpp’s 8-bit format, q8_0, which stores a little extra), so about twice as many people fit.

    In 20 GB of spare memory: 4 conversations of 32,768 tokens with 16-bit notes, 8 with 8-bit notes.

While a model writes, it keeps notes on the conversation so far (the KV cache), so it never has to reread it. The notes share memory with the model and grow with every token and every person, until they can outgrow the model itself.

Picture it

A waiter has one notebook. The menu (the model) takes the first pages; each table’s running order (its notes) takes more as the meal goes on. A few very long dinners and the notebook is full, though the menu itself fits fine.

With real numbersLlama 3.1 8B stored at 4 bits, lesson 05

  • The model itself: 4.9 GB.
  • Notes per token: 2 kinds (keys and values) x 32 layers x 8 note sets x 128 numbers each x 2 bytes per number (16 bits) = 131,072 bytes, which is 128 KB (1 KB = 1,024 bytes here).
  • A 4,096-token conversation (about 3,000 words): 131,072 bytes x 4,096 = 0.537 GB of notes, about half a GB.
  • 32 times longer (131,072 tokens, about 98,000 words): 32 x 0.537 = 17.18 GB, three and a half times the model.
  • The same notes stored in 8 bits: 9.13 GB, about half (llama.cpp prints this as 8,704 MiB).

Words to know

Token
A chunk of text, about three quarters of a word. 1,000 tokens is roughly 750 words.
KV cache
The model’s notes on the conversation so far (its keys and values), so it does not reread everything for each new word. Each person has their own.
Context length
The most tokens one conversation may hold. Example: 131,072 at the model’s maximum.
Layer
One of 32 stacked processing stages each token passes through; each keeps its own notes.

From the lesson

lessons/05-kv-cache/README.md

While a model writes, it keeps notes on every token it has seen so far (the Keys and Values), so it doesn’t reread the whole conversation for each new word. Those notes are the KV cache. They grow with every token and every user:

KV bytes/token = 2 × layers × kv_heads × head_dim × bytes
Llama-3.1-8B = 2 × 32 × 8 × 128 × 2 B = 128 KiB per token
× 131,072 tokens ≈ 17 GB for ONE long conversation

The weights are 4.9 GB. The cache, not the model, is what runs out. It is the most common production failure and the easiest to fix:

  1. Cap the context to what the use case needs. A 4k support chat doesn’t need a 128k window.
  2. Quantize the cache (FP8 or Q8), which roughly halves it and doubles how many users fit.
  3. PagedAttention (vLLM) allocates cache in small pages instead of one worst-case block per user, so less goes to waste.
  • kv_calc.py has the formula, with model shapes taken from config.json.
  • context_ladder.sh starts the server six times and greps the allocation line from the log.
Terminal window
cd lessons/05-kv-cache
python kv_calc.py
python kv_calc.py --ctx 131072
python kv_calc.py --ctx 32768 --free-gb 20 # how many users fit in 20 GB?
bash context_ladder.sh # the real allocations

What each command does

  1. python kv_calc.py

    Run it with the Python environment on, as in step 1. It works out the notes’ size for one 8,192-token conversation (about 6,000 words) on Llama 3.1 8B. Look for 128 KiB per token (the script prints KiB, 1,024 bytes; the videos write 128 KB) and 1.07 GB per conversation.

  2. python kv_calc.py --ctx 131072

    The same sum for the longest conversation the model allows (131,072 tokens, about 98,000 words). Look for 17.18 GB, three and a half times the 4.9 GB model.

  3. python kv_calc.py --ctx 32768 --free-gb 20

    How many 32,768-token conversations (about 25,000 words each) fit in 20 GB of spare memory. Look for 4, then 8 on the line that switches the notes to 8 bits (q8_0).

  4. bash context_ladder.sh

    Starts the real server six times and prints the memory it set aside for notes each time. Look for each 8-bit row at about half the 16-bit row above it. On a Mac with 16 to 24 GB of memory, the longest 16-bit row may say “failed to start”: that is the lesson. If every row is exactly half of the lesson’s, see Stuck? below.

How to read it

Each row is one server start: conversation length in tokens (ctx), note format (f16 = 16-bit, q8_0 = 8-bit) and the memory set aside for notes. From the 4,096 rows to the 32,768 rows, 8 times the tokens means 8 times the memory (512 to 4,096 MiB, 0.54 to 4.29 GB). In each pair the 8-bit row is about half, and it may start even where the 16-bit row says “failed to start”.

ctx cache what llama.cpp allocated
4096 f16 512.00 MiB
4096 q8_0 272.00 MiB
32768 f16 4096.00 MiB
32768 q8_0 2176.00 MiB
131072 f16 16384.00 MiB (or ✗ on a small Mac — the lesson)
131072 q8_0 8704.00 MiB ← the fix

kv_calc.py --ctx 131072 predicts 17.18 GB, which is 16,384 MiB. The log agrees.

3 questions. Say your answer out loud, then tap to check it.

Inside each layer of Llama 3.1 8B, 32 parts look back over the conversation. A design called GQA (grouped-query attention) makes them share 8 sets of notes instead of keeping 32. Why does that matter so much for memory?Show answerHide

In plain words

Sharing makes the notes 4 times smaller. That one design choice is what makes long conversations affordable.

Picture it

32 reporters cover the same meeting. Instead of each writing a full transcript, they work in 8 teams of 4 and each team shares one transcript. Same meeting, a quarter of the paper.

With real numbersthe lesson 05 formula

  • Only one number in the notes formula changes: 8 note sets per layer instead of 32.
  • With 8: 128 KB per token. With 32: 512 KB, 4 times more.
  • One 131,072-token conversation (about 98,000 words): 17.18 GB with sharing, 68.72 GB without.
  • 68.72 GB is 14 times the 4.9 GB model, and most of an 80 GB H100 (NVIDIA’s data-center AI chip).

Words to know

GQA (grouped-query attention)
Groups of the look-back parts (attention heads) share one set of notes. Example: 8 KV heads instead of 32.
KV head (note set)
One set of notes kept in each layer. Example: Llama 3.1 8B keeps 8 per layer.
Layer
One of the stacked processing stages each token passes through; each keeps its own notes. Example: 32 in Llama 3.1 8B.
H100
NVIDIA’s data-center GPU with 80 GB of memory. Example: the GPU in the KV Cache, GPU Bandwidth and Quantization videos.
Go deeper: the engineer version

The kit's question

Why do GQA models (8 KV heads instead of 32) matter so much here?

The kit's answer

The cache is 4× smaller. That architecture choice is what makes long context affordable.

More detail: GQA keeps 32 query heads but shares each K/V head across a group of 4 of them, so attention still looks at the text 32 ways while only 8 K/V heads are cached. Per token: 2 x 32 x 8 x 128 x 2 bytes = 128 KiB with GQA; 2 x 32 x 32 x 128 x 2 bytes = 512 KiB without. The KV Cache video puts the usual saving at 4 to 8 times; the field guide notes that Llama 3 70B needs only about 0.3 MB per token because of it.

A customer’s prompts are almost always short: 95 out of 100 are under 3,000 tokens (about 2,250 words). But they set the limit to 128K tokens (128 x 1,024 = 131,072 tokens, about 98,000 words) “to be safe”. What does that cost them?Show answerHide

In plain words

On a server that sets aside room for the full limit up front, as llama.cpp does, they pay for memory they never use. The same GPU (the chip that runs the model) then serves far fewer people at once.

Picture it

A restaurant sets a 40-seat table for every party “to be safe”, but almost every party is 2 or 3 people. After a few parties the room is full while most chairs sit empty, and new guests wait at the door. Servers with PagedAttention (vLLM) add chairs as guests arrive, so they waste far less, though a few huge parties can still fill the room.

With real numbersLlama 3.1 8B, from lesson 05’s kv_calc.py

  • The model keeps notes on every token it has read: 128 KB per token.
  • Room for 128K tokens: 128 KB x 131,072 = about 17 GB per person. With 20 GB free, that is 1 person.
  • Room for 8K tokens (8,192, still more than double the 3,000 they use): 128 KB x 8,192 = 1.07 GB per person. 20 ÷ 1.07 = 18.7, so 18 people.
  • The fix: set the limit near real use, about 8K here.

Words to know

Token
A chunk of text, about three quarters of a word. 1,000 tokens is roughly 750 words.
KV cache
The model’s notes on the conversation so far, so it does not reread everything for each new word. Each person has their own, and it grows with every token.
p95
The size that 95 out of 100 requests stay under.
People served at once (concurrency)
How many conversations the GPU runs at the same moment.
Go deeper: the engineer version

The kit's question

The customer’s p95 prompt is 3k tokens but they configured 128k “to be safe”. What does that cost?

The kit's answer

Engines that reserve the maximum per slot strand memory. Capping at about 8k frees it for more concurrent users.

More detail: llama.cpp sets aside the whole -c cache at startup, split evenly across its --parallel slots (see lessons/02-three-local-servers/serve_llamacpp.sh). vLLM’s PagedAttention hands out cache in small pages as requests grow, so it wastes far less; a high maximum still lets a few very long requests take most of the memory. Check the sums: python kv_calc.py --ctx 131072 --free-gb 20 prints 17.18 GB and 1 conversation; python kv_calc.py --ctx 8192 --free-gb 20 prints 1.07 GB and 18.

Storing the conversation notes with 8 bits per number instead of 16 (an FP8 KV cache) halves their memory. What does it cost you?Show answerHide

In plain words

A small risk that answers get a little worse on some tasks. You measure it on the customer’s own examples before you switch it on.

Picture it

Saving a photo at lower quality roughly halves the file, and usually nobody can tell. Sometimes the fine print blurs, so you compare both versions before sending it to a client.

With real numbersLlama 3.1 8B, lesson 05’s kv_calc.py

  • 16-bit: 2 bytes per number. 8-bit (q8_0): 1 byte per number plus a small shared scaling number stored with each block of numbers, 1.0625 bytes in all.
  • One 131,072-token conversation (about 98,000 words): 16,384 MiB (17.18 GB) becomes 8,704 MiB (9.13 GB): 53%, not 50%, because of that extra.
  • 20 GB of spare memory at 32,768 tokens: 4 conversations become 8.
  • The price: a small risk to answer quality on some tasks. Check it by running lesson 09’s test (Day 7) on 50 of the customer’s prompts, with 16-bit and with 8-bit notes, and compare.

Words to know

FP8 and q8_0
Two 8-bit formats; FP8 on NVIDIA data-center GPUs, q8_0 in llama.cpp on your Mac.
Bit
A single 0 or 1; 8 bits make a byte. Example: a 16-bit number takes 2 bytes.
Byte
8 bits. Example: one 16-bit number takes 2 bytes.
Eval harness
A repeatable test that scores a model on a fixed set of prompts. Example: lesson 09 (Day 7).
Go deeper: the engineer version

The kit's question

What does FP8 KV cost you?

The kit's answer

A small quality risk on some tasks. Check it with the eval harness in lesson 09.

More detail: q8_0 stores 8-bit integers in blocks that share a scale, which is where 1.0625 bytes per value comes from (kv_calc.py line 25); FP8 is the equivalent on H100-class GPUs and is exactly half of FP16. Run the same eval with the 16-bit and the 8-bit cache and compare before you promise the doubled capacity.

Question The model fits on the GPU, but we still run out of memory under load. Why?

One clear answer

Weights fit; the cache doesn’t. I size KV per token from the config, cap context to the real p95, and quantize the cache. That usually doubles concurrency on the same GPU.

What this means

  • “Weights fit”: The model itself, its learned numbers (the weights), is 4.9 GB and fits easily.
  • “the cache doesn’t”: The per-person notes (the KV cache) run out; one 131,072-token conversation (about 98,000 words) needs 17.18 GB.
  • “I size KV per token from the config”: I work out how much memory each token of notes takes from the model’s settings file. For Llama 3.1 8B it is 128 KB.
  • “cap context to the real p95”: I set the maximum conversation length to what 95 of 100 real requests need, not the model’s maximum. For a customer whose requests are almost all under 3,000 tokens, a limit of 8,192 still leaves room to spare: in 20 GB it serves 18 people instead of 1 at 131,072.
  • “quantize the cache”: I store the notes with 8 bits per number instead of 16, like saving a photo at lower quality. It takes about half the space: 9.13 GB instead of 17.18 GB.
  • “doubles concurrency on the same GPU”: About twice as many people served at once, with no new hardware. In 20 GB at 32,768 tokens: 4 people become 8.

Your numbersSaved on this device and collected in the Day 4 wrap-up.

Hint: The 32768 f16 and 32768 q8_0 rows of context_ladder.sh. Write both, for example “4,096 / 2,176”. If you see exactly half of these (2,048 and 1,088), the script copied only the last number on llama.cpp’s line. In lessons/05-kv-cache, run grep -i "kv buffer size" logs/ctx32768_f16.log (then the same for logs/ctx32768_q8_0.log): the MiB number on that line is the full figure.

Hint: The 131072 f16 row of context_ladder.sh. Expect about 16,384 (17.18 GB). If it failed to start, type “failed”: that is the lesson. If it shows 8,192, exactly half, run grep -i "kv buffer size" logs/ctx131072_f16.log in lessons/05-kv-cache: the MiB number on that line is the full figure.

Hint: The 131072 q8_0 row. Expect about 8,704 (9.13 GB), a little over half of 16,384 (it may also fail on a 16 GB Mac). If it shows 4,352, exactly half, read logs/ctx131072_q8_0.log the same way.

Hint: From python kv_calc.py --ctx 32768 --free-gb 20. Expect 4, and 8 with 8-bit notes: write both, for example “4 / 8”.

The calculator says about 17 GB (17.18) for one 131,072-token conversation (about 98,000 words). And the real server shows the 8-bit notes at about half the 16-bit size, for example 2,176 MiB (2.28 GB) against 4,096 MiB (4.29 GB) at 32,768 tokens.

Every 8-bit (q8_0) row fails, even the short 4,096-token one.
Turn on the server’s flash attention setting. (Flash attention is a faster way to run the look-back step; llama.cpp needs it to store the notes in 8 bits.) Run EXTRA="-fa on" bash context_ladder.sh. On older llama.cpp versions use EXTRA="-fa" bash context_ladder.sh.
The 131,072-token 16-bit row says failed to start.
That is the lesson. The notes (17.18 GB) plus the 4.9 GB model need about 22 GB, but a Mac gives the model only about 70% of its memory (Day 1’s rule): about 17 GB on a 24 GB Mac. Write “failed” in Your numbers; the 8-bit row below it is the fix.
Both 131,072-token rows fail (a 16 GB Mac).
That can happen. The 9.13 GB of 8-bit notes plus the 4.9 GB model is about 14 GB, more than the about 11 GB a 16 GB Mac gives the model (70%, Day 1). Write “failed” for both; the 32,768-token pair shows the same halving.
Your rows are exactly half the lesson’s, for example 2,048 and 1,088 at 32,768 tokens.
llama.cpp printed the notes’ two halves (keys and values) on the same line, and the script copied only the last number. In lessons/05-kv-cache, run grep -i "kv buffer size" logs/ctx32768_f16.log, using the log named after the row (its ctx, then f16 or q8_0). The MiB number on that line is the full figure.
context_ladder.sh Bash · 45 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 05 · step 2 — climb the context ladder and read the KV size from the log.
#
# For each (context, cache type) we start llama-server, wait for it to load,
# grep the line where llama.cpp reports how much KV cache it allocated, then stop it.
# Compare every number with kv_calc.py — they should match within a few %.
#
# f16 : the default. 131k context on an 8B model ≈ 17 GB of cache alone.
# q8_0 : --cache-type-k q8_0 --cache-type-v q8_0 → roughly half.
#
# If the server refuses a quantized V cache, your build needs flash attention on:
# add -fa on (newer builds) or -fa (older builds) to EXTRA below.
# On a 16–24 GB Mac the 131072/f16 rung may fail to load. That IS the lesson.
# ─────────────────────────────────────────────────────────────────────────────
set -uo pipefail
cd "$(dirname "$0")"
REPO=${LLAMACPP_REPO:-bartowski/Meta-Llama-3.1-8B-Instruct-GGUF}:Q4_K_M
EXTRA=${EXTRA:-}
PORT=8090
mkdir -p logs
printf "%-8s %-6s %s\n" ctx cache "what llama.cpp allocated"
for ctx in 4096 32768 131072; do
for kv in f16 q8_0; do
log="logs/ctx${ctx}_${kv}.log"
llama-server -hf "$REPO" -c "$ctx" -ngl 99 --parallel 1 --port $PORT \
--cache-type-k "$kv" --cache-type-v "$kv" $EXTRA >"$log" 2>&1 &
pid=$!
# wait until the server says it is listening, or dies (e.g. out of memory)
for _ in $(seq 1 120); do
grep -q "listening" "$log" && break
kill -0 $pid 2>/dev/null || break
sleep 1
done
size=$(grep -iE "kv.*(size|buffer).*MiB" "$log" | tail -1 | sed 's/^.*: *//')
[[ -z "$size" ]] && size="✗ failed to start — see $log"
printf "%-8s %-6s %s\n" "$ctx" "$kv" "$size"
kill $pid 2>/dev/null; wait $pid 2>/dev/null
done
done
echo
echo "Check: python kv_calc.py --ctx 131072 (f16)"
echo " python kv_calc.py --ctx 131072 --kv q8_0 (q8)"
kv_calc.py Python · 58 lines
"""
Lesson 05 · step 1 — KV cache maths by hand (then check it against the server log).
For every token in every live conversation, each layer stores a Key and a Value vector:
KV bytes per token = 2 (K and V) × layers × kv_heads × head_dim × bytes_per_value
Multiply by context length and concurrent users and you get the memory that is NOT
the model — and it is usually what runs out first.
python kv_calc.py # Llama-3.1-8B, 8k ctx, 1 user
python kv_calc.py --model llama-3.1-70b --ctx 32768 --users 16
python kv_calc.py --ctx 131072 --kv q8_0 --free-gb 20 # how many users fit?
"""
import argparse
# layers, kv_heads (GQA), head_dim — from each model's config.json
MODELS = {
"llama-3.1-8b": (32, 8, 128),
"llama-3.1-70b": (80, 8, 128),
"qwen2.5-7b": (28, 4, 128),
"mistral-7b": (32, 8, 128),
"gpt-oss-20b": (24, 8, 64),
}
BYTES = {"f16": 2.0, "q8_0": 1.0625, "q4_0": 0.5625} # llama.cpp block formats carry a small scale overhead
def main() -> None:
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
ap.add_argument("--model", default="llama-3.1-8b", choices=MODELS)
ap.add_argument("--ctx", type=int, default=8192, help="tokens per conversation")
ap.add_argument("--users", type=int, default=1, help="concurrent conversations")
ap.add_argument("--kv", default="f16", choices=BYTES, help="KV cache precision")
ap.add_argument("--free-gb", type=float, help="memory left after weights → max users")
a = ap.parse_args()
L, H, D = MODELS[a.model]
per_tok = 2 * L * H * D * BYTES[a.kv] # bytes
per_seq = per_tok * a.ctx
total = per_seq * a.users
print(f"{a.model}: {L} layers × {H} KV heads × {D} head dim, KV in {a.kv}\n")
print(f" per token 2 × {L} × {H} × {D} × {BYTES[a.kv]} B = {per_tok / 1024:,.0f} KiB")
print(f" per conversation × {a.ctx:,} tokens = {per_seq / 1e9:,.2f} GB")
print(f" all users × {a.users} users = {total / 1e9:,.2f} GB")
if a.free_gb:
fit = int(a.free_gb * 1e9 // per_seq)
print(f"\n with {a.free_gb} GB free for cache → {fit} concurrent conversations at {a.ctx:,} tokens")
alt = int(a.free_gb * 1e9 // (per_seq / BYTES[a.kv] * BYTES['q8_0'])) if a.kv == "f16" else None
if alt:
print(f" switch KV to q8_0 → {alt} conversations (≈2×) — the cheapest fix there is")
print("\nNow compare with what llama-server reports: bash context_ladder.sh")
if __name__ == "__main__":
main()