Skip to content

QLoRA On Your Own Box

file 12 · 2:49 · subtitles burned in

Chapters

In this video How to fine-tune on a computer you already own, with the model shrunk to 4 bits and left unchanged, and how that compares with today’s paid cloud run.

What it explains: 4-bit base + adapter on your own box

3 key points

  1. QLoRA (LoRA on a model shrunk to 4 bits) leaves the model unchanged and trains only a small add-on, so an 8-billion-number model trains on a laptop.

    The video’s Mac example: about 5 GB of model left unchanged (frozen), a 30 MB add-on, about 11 GB of memory at the peak, 30 to 50 minutes, $0. Day 8 runs a similar one on your Mac: Qwen2.5 7B on the same 800 tickets.

  2. The data and a test set kept aside decide the result, not the hardware.

    200 messy, unchecked examples: 97% right on the examples it trained on, but 74% on new ones, so it memorised. 2,000 clean examples: 93% and 89%. Set the test examples aside before training, never after.

  3. Compare doing it yourself with a managed service on time, privacy and cost.

    Your own machine: free, and the data stays home. Managed, as in today’s run: the video says a dollar or two of training (today’s smaller dataset costs cents) and a working model in minutes. But serving a LoRA on a hosted platform usually needs a GPU reserved for it.

The title card in the video says “Video 12 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (10.4 MB)