Skip to content

Quantization

file 04 · 2:43 · subtitles burned in

Chapters

In this video Why storing a model’s numbers in fewer bits makes it faster and fits more people on one chip, and how to prove the answers stay good.

What it explains: FP16 → FP8 → 4-bit, NVFP4, KV quantization

3 key points

  1. Storing a model at 8 bits instead of 16 halves its size; on NVIDIA’s H100 (a data-center AI chip) and newer chips it also about doubles the math speed.

    A 70-billion-number model is 140 GB at 16 bits (2 bytes per number) and 70 GB at 8 bits. On one 80 GB H100 that leaves 10 GB spare. Each 8K-token conversation (about 6,000 words) needs about 2.6 GB of notes, so only 3 whole ones fit (10 ÷ 2.6 = 3.8).

  2. Storing it at 4 bits cuts it to a quarter of the 16-bit size, but prove quality on the customer’s own tests first.

    The same model is 35 GB at 4 bits, leaving 45 GB spare. At 2.6 GB of notes per conversation that is at most 17 (45 ÷ 2.6 = 17.3); the video plays safe and says about 14. Its rule: if 2 fewer replies in 100 come back in the exact format asked for, but the server writes twice as many tokens per second, that is usually a good trade; 20 fewer in 100 never is.

  3. Storing the conversation notes (the KV cache) in 8 bits is often a bigger win than shrinking the model.

    It doubles how many people fit in the same memory. You measure it in step 2: 20 GB of spare memory holds 4 conversations of 32,768 tokens with 16-bit notes, and 8 with 8-bit notes.

Used in the course

The title card in the video says “Video 4 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (9.7 MB)