Fitting Big Models
file 07 · 3:07 · subtitles burned in
Chapters
In this video How to answer “will it fit?” with a formula, why mixture-of-experts models are big in memory but fast, and the two ways to split a model across GPUs.
What it explains: sizing, offloading, MoE, tensor vs pipeline parallel
3 key points
Memory needed = the model’s weights + its notes + about 10% for the server’s working memory.
The weights are the model’s learned numbers (parameters): count × bytes each. A 405-billion-parameter model at 16 bits (2 bytes each) is 810 GB, more than eight 80 GB H100s (NVIDIA’s data-center chips) hold: 640 GB.
A mixture-of-experts model needs memory for all of itself, but reads only a few experts per token.
The video’s example on a 128 GB desktop (a DGX Spark): a 120-billion-parameter model stored at 4 bits (half a byte each) is about 60 GB on disk but reads about 5 GB per token. It writes 55 tokens a second, against 23 for a dense 14-billion-parameter one.
If it does not fit: store it in fewer bits, add GPUs, or allow shorter conversations and fewer users. Moving part of the model off the GPU is the last resort.
Offloading (moving layers to the CPU or disk) is about 10 times slower (the video’s “order of magnitude”), because the model then crosses a much narrower connection than the GPU’s own memory.
Used in the course
The title card in the video says “Video 7 of 18”: that is the file order. The course plays the videos in the order of its days.