Skip to content

GPU Bandwidth

file 03 · 3:10 · subtitles burned in

Chapters

In this video Why one person’s writing speed is set by how fast memory delivers the model, and how to predict it on any chip.

What it explains: roofline, decode ceiling = bandwidth ÷ bytes

3 key points

  1. Top writing speed for one person = memory speed ÷ model size.

    Each token needs the whole model fetched once. A 70-billion-number model at 8 bits (1 byte per number) is 70 GB. An H100 (NVIDIA’s data-center GPU) fetches 3,350 GB a second: 3,350 ÷ 70 = about 48 tokens per second.

  2. Real dense models (ones that use every number for each token) reach about 60 to 85% of that limit.

    The video’s DGX Spark (NVIDIA’s desktop AI computer) moves 273 GB a second. For an 8-billion-number model at 4 bits, which the video counts as 4.5 GB, the limit is 273 ÷ 4.5 = 61 tokens per second. It measured 38.7: 64% of the limit. Today’s 4-bit file is 4.9 GB, which gives 273 ÷ 4.9 = 56.

  3. Batching is how you put the chip’s idle math to work.

    An H100 can do about 989 trillion calculations a second but fetch only 3,350 GB a second: about 300 calculations per byte (989 ÷ 3.35 = 295). Serving one person needs only about 2 per number fetched, so the math sits idle. Serving many people from each fetch puts it to work.

Used in the course

The title card in the video says “Video 3 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (12.2 MB)