Full Fine-Tuning
file 14 · 4:17 · subtitles burned in
Chapters
In this video Why changing every number in a model is the most powerful and the most expensive way to train it, and when it is worth it.
What it explains: training loop, 16 bytes/param, ZeRO/FSDP, when full FT pays
3 key points
Training needs about 16 bytes of memory per number in the model; running it needs about 2.
Each number brings its gradient (the direction to nudge it), two running averages and a precise master copy. A 7-billion-number model runs in about 14 GB (7 x 2) but needs about 112 GB (7 x 16) to train: 8 times as much.
Memory, not computing speed, decides whether you can train fully at all.
An 8-billion-number (8B) model needs about 130 GB (8 x 16 = 128), so two 80 GB H100s (NVIDIA’s data-center GPU). A 70B needs about 1.1 TB (terabytes: 70 x 16 = 1,120 GB): sixteen H100s plus software that splits the job across them.
On a narrow task, LoRA usually lands within 1 to 2 accuracy points of full training, and forgets less of what the model already knew.
Same 8B model: full training changes 8 billion numbers in about 130 GB and saves a 16 GB file. LoRA trains about 20 million in about 20 GB and saves about 40 MB. Go full for a big shift, such as a new language, with hundreds of thousands of examples.
Used in the course
The title card in the video says “Video 14 of 18”: that is the file order. The course plays the videos in the order of its days.