Skip to content

Parameter-Efficient Fine-Tuning

file 16 · 3:50 · subtitles burned in

Chapters

In this video How LoRA customises a huge model by training two thin grids of numbers beside each frozen one, and when to pick its relatives: QLoRA, DoRA and multi-LoRA.

What it explains: LoRA maths, QLoRA, DoRA, adapters, IA³, multi-LoRA

3 key points

  1. LoRA freezes the model and trains two thin grids, A and B. Multiplied together, they make a full-size correction that is added to the frozen grid.

    For one 4,096 x 4,096 grid (16.8 million numbers) at rank 16 (the add-on’s size setting), A and B hold 65,536 numbers each: 131,072 in all, 0.8% of that grid.

  2. The rank is the dial: a higher rank gives the add-on more room to learn, and costs more memory.

    Rank 8 to 16 suits narrow tasks, 32 to 64 harder ones. Across a whole 8B model, adapting the attention grids, the video puts it at 0.2 to 0.5% of all numbers. Step 2’s calculator gives 0.17% at rank 16, because two of Llama 3.1 8B’s four attention grids are a quarter size (a design called GQA). Four full-size grids would give 0.21%.

  3. Pick QLoRA when memory is short, DoRA when quality lags, multi-LoRA when you have many customers.

    QLoRA is LoRA on a 4-bit copy of the model; DoRA is a LoRA variant that often scores 1 to 2 points higher. Fifty customers with full fine-tunes need 50 files of 16 GB (800 GB) and 50 deployments. Fifty LoRA add-ons of about 40 MB (2 GB in all) fit on one deployment, the right one applied per request.

The title card in the video says “Video 16 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (14.3 MB)