QLoRA On Your Own Box
file 12 · 2:49 · subtitles burned in
Chapters
In this video How to fine-tune on a computer you already own, with the model shrunk to 4 bits and left unchanged, and how that compares with today’s paid cloud run.
What it explains: 4-bit base + adapter on your own box
3 key points
QLoRA (LoRA on a model shrunk to 4 bits) leaves the model unchanged and trains only a small add-on, so an 8-billion-number model trains on a laptop.
The video’s Mac example: about 5 GB of model left unchanged (frozen), a 30 MB add-on, about 11 GB of memory at the peak, 30 to 50 minutes, $0. Day 8 runs a similar one on your Mac: Qwen2.5 7B on the same 800 tickets.
The data and a test set kept aside decide the result, not the hardware.
200 messy, unchecked examples: 97% right on the examples it trained on, but 74% on new ones, so it memorised. 2,000 clean examples: 93% and 89%. Set the test examples aside before training, never after.
Compare doing it yourself with a managed service on time, privacy and cost.
Your own machine: free, and the data stays home. Managed, as in today’s run: the video says a dollar or two of training (today’s smaller dataset costs cents) and a working model in minutes. But serving a LoRA on a hosted platform usually needs a GPU reserved for it.
Used in the course
The title card in the video says “Video 12 of 18”: that is the file order. The course plays the videos in the order of its days.