The same LoRA on your Mac
course 13 of 22
lesson 11 · lessons/11-lora-local-mlx
1 h 25 Free Mac
What you will do and why
You repeat Day 7’s fine-tune (extra training on your own examples) on the same 800 support tickets, but on your own Mac, for free. Doing it both ways gives you the build-versus-buy comparison (do it yourself or pay a service), with your own numbers, that you can explain to a customer.
Why it matters: The lesson’s sample: after training, the Mac’s model names the right ticket category 96% of the time, up from 78% before, at $0. Day 7’s Fireworks sample went from 80% to 97%, plus about $5 for the rented cloud GPU that ran the test.
You are done when: Validation loss has flattened and training has stopped. Your score table has a row for the fine-tune and one for the original model, on the same 60 test tickets. COMPARISON.md has your Mac column filled in next to Day 7’s Fireworks column.
Start this first
Run it from the 03-labs folder with the Python environment on (source .venv/bin/activate). The first time, it downloads the 4-bit starting model (about 4.3 GB by the lesson’s estimate), then trains for about 15 to 40 minutes. Read on while it runs. Keep this window for training and, later, the server. The first command below, python qlora_memory.py, needs a second terminal window: go to 03-labs, run source .venv/bin/activate, then cd lessons/11-lora-local-mlx.
cd lessons/11-lora-local-mlx && bash 1_train_mlx.shOptional: QLoRA On Your Own Box (from Day 7), 2:49ShowHide
In plain words
Section titled “In plain words”Today you train Day 7’s small add-on for a model (a LoRA adapter) again, from scratch, on your own Mac. The model’s own 7.6 billion numbers stay frozen (unchanged), stored at 4 bits, about half a byte each. Only the add-on learns. That is QLoRA, and it is why this model can be trained on a laptop.
Picture it
The model is a heavy printed textbook you cannot write in. The adapter is a pad of sticky notes you can write on. You carry, ship and swap only the notes; the book stays on the shelf. Today’s second script does glue the notes into a photocopy of the book (a fused model), because one self-contained copy is simplest to hand over.
With real numbersQwen2.5 7B Instruct (7.6 billion numbers): lesson 11’s memory script qlora_memory.py and sample run, and the Full Fine-Tuning video (Day 11) for the 16 bytes
- Retraining every number (a full fine-tune) needs about 16 bytes per number. That is 2 for the number, 2 for its correction (gradient), 8 for the optimizer’s two running averages and 4 for a more precise (32-bit) master copy of the number. 7.6 billion x 16 bytes = 121.6 GB, one and a half times what an 80 GB data-center GPU holds (121.6 ÷ 80 = 1.5).
- QLoRA stores the frozen model at 4 bits, half a byte per number, plus small shared helper numbers (scales): 0.5625 bytes per number in all. 7.6 billion x 0.5625 bytes = 4.3 GB.
- The add-on sits on the model’s last 8 layers (processing stages): about 1.8 million numbers by the script’s rough count, 0.024% of the model, about 1 in every 4,000. Training it takes 1.8 million x 16 bytes = 0.03 GB, about 30 MB.
- Plus about 2 GB of working memory (the script’s rough allowance): 4.3 + 0.03 + 2 = 6.3 GB in all, about a nineteenth of 121.6 GB (121.6 ÷ 6.3 = 19).
- The add-on file you keep: the script estimates 3.7 MB (1.8 million x 2 bytes). The lesson’s sample run printed
7.1M, about 7.4 MB (that M is 1,048,576 bytes). The lesson does not say why that is twice the estimate. Either way the model is hundreds of times bigger: 4,300 MB ÷ 7.4 MB = about 580.
Words to know
- Adapter (LoRA adapter)
- The small set of extra numbers LoRA (low-rank adaptation) trains and saves; the base model stays unchanged. Example:
7.1M(about 7.4 MB) in the lesson’s sample run. - QLoRA
- LoRA on a base model stored in 4 bits, so training fits in far less memory. Example: about 6.3 GB for today’s model.
- Gradient
- The correction worked out for each trainable number at each training step. Example: a full fine-tune keeps one for all 7.6 billion numbers.
- Optimizer state
- The two running averages the training method (Adam) keeps per trainable number to size each correction. Example: 8 of the 16 bytes per number in a full fine-tune.
From the lesson
lessons/11-lora-local-mlx/README.md
What and why
Section titled “What and why”This is the same dataset and task as Day 7’s fine-tune on Fireworks, trained locally. Because the base model is already 4-bit, what you are running is QLoRA: frozen 4-bit weights plus small trainable adapters. It is a heavy textbook you can’t write in (the base) with a pad of sticky notes you can write on (the adapter).
full fine-tune 7B ≈ 120 GB weights + gradients + optimizer for every parameterQLoRA 7B ≈ 6 GB 4-bit frozen base + gradients only for ~2M adapter paramsadapter file ≈ 4 MB ← what you ship, version, and hot-swap (multi-LoRA)Doing it both ways gives you the build-versus-buy story with your own numbers.
Read the code first
Section titled “Read the code first”qlora_memory.pyis the memory budget above. Change--paramsand--rankand see what moves.1_train_mlx.shexplains every flag in a comment. Watch validation loss, not training loss.2_fuse_and_serve.shshows the choice between keeping the adapter separate (the multi-LoRA pattern) and fusing it in (simplest to hand over).COMPARISON.mdis the page you fill in to explain your local-versus-managed recommendation.
cd lessons/11-lora-local-mlxpython qlora_memory.pybash 1_train_mlx.sh # 15–40 min; Ctrl-C when val loss flattensbash 2_fuse_and_serve.sh # leave runningbash 3_eval_local.sh # second terminal: activate .venv, cd to this folder first# then fill in COMPARISON.md next to your Day 7 Fireworks numbersWhat each command does
python qlora_memory.pyPrints the memory budget for today’s model (7.6 billion numbers, add-ons on the last 8 layers) in about a second, with no download. Look for full fine-tune about 121.6 GB against QLoRA about 6.3 GB, and an adapter file of about 3.7 MB. Then add
--rank 16to this command. (Rank is how wide the add-on’s small grids are; today’s is 8.) The add-on doubles, from 1.8 to 3.7 million numbers and from 3.7 to 7.3 MB, while the QLoRA total stays at 6.3 GB.bash 1_train_mlx.shTrains the add-on: 300 steps of 2 tickets each, 600 of the 800 training tickets (three quarters of one pass). It takes about 15 to 40 minutes. If you started it at the top of the page, let that run finish instead of starting a second one. Look for
Val lossevery 50 steps: how wrong the model still is on practice tickets it never trains on. It should fall fast, then flatten. If it flattens early, press Ctrl+C. The last save is kept, but the size line does not print: runls -lh adaptersto see the files. At the end, look foradapter saved:with the size of the add-on’s folder (sample:7.1M).bash 2_fuse_and_serve.shIn the first window, once training has finished. It bakes the add-on into a full copy of the model (the
fusedfolder), then starts a server on port 8081 (a numbered door on your Mac). Day 2’s MLX server used the same port, so stop that one first. Leave this window running. Look for the line endingserving on :8081 (Ctrl-C to stop). It prints just before the model loads, so wait a few seconds before the next command.bash 3_eval_local.shRun it in the second window, the one you used for
qlora_memory.py; a new window first needssource .venv/bin/activatefrom03-labs, thencd lessons/11-lora-local-mlx. It tests your fine-tune, then the original model without the add-on, on the same 60 test tickets. It uses Day 7’s test script (the eval harness). The server switches models in between, so the original model’s first answer is slow. Look for themlx fusedrow’scategory_%well above the other row’s (sample: 96.0 against 78.0). Your table has more columns than the sample. Readcategory_%,severity_%andp95_ms;judge_1to5showsnan(not a number) because no judge runs today. One ticket is worth 1.67 points (100 ÷ 60), so real scores land on steps like 95.0 or 96.7. The sample’s 96.0 and 78.0 are not possible with 60 tickets, so treat them as rough examples.
What you should see
Section titled “What you should see”How to read it
Val loss scores how wrong the model still is on practice tickets it never trains on (the validation set); lower is better, and only the trend matters. In the sample it halves from step 50 to 100 (0.412 to 0.198), then drops only 0.077 more by step 300: that slowdown is the flattening. In the score table, category_% and severity_% are the share of 60 test tickets each model got right: in the sample, 96.0 and 94.0 for the fine-tune (fused), 78.0 and 66.0 for the model without it.
Iter 50: Val loss 0.412Iter 100: Val loss 0.198Iter 300: Val loss 0.121 ← flattening: more iterations would memorise▸ adapter saved: 7.1M
candidate category_% severity_%mlx …/fused 96.0 94.0mlx Qwen2.5-7B-Instruct-4bit 78.0 66.0Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
Day 8’s lesson keeps two sets of tickets out of training: 100 validation tickets, checked during training, and 100 test tickets. Why is the final score taken on the test tickets (test.jsonl) and not the validation ones (valid.jsonl)?Show answerHide
In plain words
You used the validation tickets to decide when to stop training, so they helped pick the model and can flatter it. The test tickets played no part in any decision, so their score is honest.
Picture it
A coach decides when the team has practised enough by watching its scores on one practice test. A good score on that same test afterwards proves little. The fair check is a fresh test nobody has looked at.
With real numbersthe kit’s ticket data (data/make_tickets.py) and Day 8’s training script
- 1,000 tickets, split three ways: 800 to train on, 100 to check progress during training (validation), 100 kept back for the final grade (test).
- Training checks the validation tickets every 50 steps, and you stop where their loss flattens: in the sample, 0.198 at step 100 and 0.121 at step 300.
- That stopping point was chosen by looking at the validation tickets, so they are no longer neutral.
- The test tickets never touch training or the stop decision. No test ticket repeats a training ticket, and 30 of the 100 use wordings training never saw.
- The final grade uses the first 60 test tickets, so each one is worth 1.67 points (100 ÷ 60).
Words to know
- Validation set
- Examples kept out of training and checked during it, to decide when to stop. Example:
valid.jsonl, 100 tickets. - Test set (held-out set)
- Examples kept out of training and out of every decision, used only for the final score. Example:
test.jsonl, 100 tickets; the eval uses 60. - Validation loss
- A score of how wrong the model still is on the validation set; lower is better. Example: 0.412, 0.198 and 0.121 at steps 50, 100 and 300.
- Leak (data leakage)
- Data used to judge a model has also shaped it, so the score looks better than it is. Example: the validation tickets, once they set the stopping point.
Go deeper: the engineer version
The kit's question
Why do we evaluate on test.jsonl and not valid.jsonl?
The kit's answer
Validation loss guided when to stop, so it has leaked into the decision. Test is untouched.
More detail: Early stopping on validation loss is a form of model selection, so the validation score is optimistic; the test split stays untouched for the final number. make_tickets.py de-duplicates across all splits and keeps one phrasing per category for the test set only, so the test score measures generalisation, not memory. One caution: the lab book’s standalone snippet builds its validation file with head -50 train.jsonl after copying all of train.jsonl into training, so those 50 rows are also training rows and its validation loss would look better than it is. Day 8’s script uses the separate data/tickets/valid.jsonl.
When would you tell a customer to fine-tune on their own machines (a laptop, or AI chips called GPUs that they own) rather than on a managed platform such as Fireworks?Show answerHide
In plain words
When their data is not allowed to leave their own machines. Or when they want fast, free experiments first, before paying a managed platform to train the final version and serve it to real users.
Picture it
A chef tests new recipes in their own kitchen: it costs nothing, and the secret sauce stays secret. For a 500-guest wedding, they hire a caterer with the staff and ovens to serve everyone at once.
With real numbersDay 8’s lesson, Day 7’s Fireworks run (lesson sample figures and its summary line), and the lab book’s GPU price
- On the Mac: $0 to train, and the 800 tickets never leave the machine. Training takes about 15 to 40 minutes.
- But one Mac serves one person at a time with no uptime promise:
COMPARISON.mdanswers “no” for 100 users at once. - On Fireworks (Day 7’s sample): training cost cents to about a dollar, but answering requests (serving) needs a GPU reserved by the hour. 38 minutes cost $5.07: $5.07 ÷ 38 x 60 = about $8 an hour, the lab book’s price.
- Accuracy is close either way in the samples: category 78% to 96% on the Mac, 80% to 97% on Fireworks.
- The usual pattern, from
COMPARISON.md: try ideas on your own machine first (prototype locally), then run the real service on a managed platform (ship managed).
Words to know
- Managed platform
- A service that runs everything between the GPUs and the customer’s app: the engine, scaling, routing and the models. Example: Fireworks on Day 7.
- Endpoint
- A web address that answers requests for one model. Example:
http://localhost:8081/v1for Day 8’s fine-tune. - Dedicated deployment
- A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not. Example: about $8 an hour.
- SLA (service level agreement)
- A written promise to a customer about uptime or speed. Example:
COMPARISON.mdsays the Mac has none.
Go deeper: the engineer version
The kit's question
When would you tell a customer to fine-tune locally or on their own GPUs?
The kit's answer
When data can’t leave their environment, or for fast iteration before a managed production run.
More detail: Local or self-hosted training fits data that must stay inside the customer’s environment, and cheap, fast iteration. Managed wins on time to a scalable endpoint, autoscaling and an SLA. On Fireworks a LoRA cannot run on serverless (shared, pay-per-token) models; it needs a dedicated deployment, so the serving hours, not the training, are the cost to plan (the lab book). Day 7’s lesson calls training “a few cents” in its sample and “about a dollar” in its summary line; the QLoRA On Your Own Box video says about a dollar. If data residency is the only blocker, the lab book notes that Fireworks can also run inside the customer’s own cloud account (BYOC), or pin a dedicated deployment to a region at 1.5 times the rate. The QLoRA On Your Own Box video puts it this way: “the local versus managed trade-off is a conversation, not a religion.”
The add-on you trained (the adapter) is about 7 MB. Why does that small size matter when the model runs as a service?Show answerHide
In plain words
One copy of the base model can keep hundreds of these small add-ons in memory and apply the right one to each request. So every customer can have their own tuned model for about the price of one shared server.
Picture it
One base recipe, a different spice packet per customer (the field guide’s comparison). The kitchen cooks one big pot and adds a small packet to each order, instead of running a separate kitchen for every customer.
With real numbersDay 7’s lesson (12 business units), Day 8’s sample run and memory script, and the lab book’s $8-an-hour GPU price
- Adapter:
7.1Min Day 8’s sample, about 7.4 MB. Base model: about 4.3 GB, which is 4,300 MB. 4,300 ÷ 7.4 = about 580, so the base is hundreds of times bigger. - 100 adapters: 100 x 7.4 MB = 740 MB, about a sixth of one base model (740 ÷ 4,300 = 0.17).
- Day 7’s example customer wants 12 tuned models, one per business unit (department). Each on its own dedicated GPU at $8 an hour: 12 x $8 = $96 an hour.
- The same 12 as adapters on one shared base: about $8 an hour, if one deployment can carry all their traffic.
- Running all month (730 hours): $8 x 730 = $5,840 for one deployment, against 12 x $5,840 = $70,080 for twelve.
Words to know
- Adapter (LoRA adapter)
- The small set of extra numbers LoRA trains and saves; the base model stays unchanged. Example: about 7 MB in Day 8’s sample.
- Base model
- The model a fine-tune starts from and leaves unchanged. Example: Qwen2.5 7B Instruct at 4 bits, about 4.3 GB.
- Multi-LoRA
- One server keeps one base model and many adapters, applying the right adapter to each request. Example: 12 business units on one deployment.
- Tenant
- One customer, or business unit, sharing a platform with others. Example: per-tenant fine-tunes, one adapter each.
Go deeper: the engineer version
The kit's question
The adapter is 7 MB. Why does that matter for serving?
The kit's answer
One base model can hold hundreds of adapters in memory, which makes per-tenant fine-tunes affordable.
More detail: Multi-LoRA servers keep one base in GPU memory and apply the right adapter per request, batching requests for different adapters together; the field guide says Fireworks serves hundreds of fine-tunes on one base model at the base model’s price. Fusing the adapter into the weights, as Day 8’s 2_fuse_and_serve.sh does for simplicity, gives that up: each fused model is a full copy. The size is the script’s estimate: 8 layers x 4 grids x 2 x 3,584 x rank 8 = 1,835,008 numbers, 3.7 MB at 2 bytes each. It treats all four grids as 3,584 x 3,584; Qwen2.5 7B’s key and value grids are 3,584 x 512 (lesson 17’s peft_params.py), which gives 1,441,792, and mlx-lm’s own defaults decide which grids get an add-on. Lesson 11’s README says about 4 MB, its sample prints 7.1M, and it does not say why they differ. du -sh adapters measures the whole folder, which also keeps copies saved during training, so quote your own ls -lh adapters/adapters.safetensors figure.
Explain what you learned
Section titled “Explain what you learned”Question Should we fine-tune on our own hardware, or pay for a managed platform?
One clear answer
I ran the same LoRA locally and managed. Local is free and private; managed got me to a scalable endpoint in minutes. The adapter is a few megabytes, so multi-LoRA makes per-customer tuning economical.
What this means
- “I ran the same LoRA locally and managed”: I trained the same small add-on (a LoRA adapter) on the same 800 support tickets twice. Once on my own Mac, once on Fireworks, a managed platform that runs the hardware for you.
- “Local is free and private”: Training on my Mac cost $0, and the tickets never left it.
- “managed got me to a scalable endpoint in minutes”: Fireworks put the add-on behind a live web address other programs can call (an endpoint). It was ready within minutes, and it can add copies as more people use it. My Mac serves one person at a time.
- “The adapter is a few megabytes”: The trained add-on is about 7 MB in the lesson’s sample, against about 4.3 GB for the model it sits on.
- “so multi-LoRA makes per-customer tuning economical”: Multi-LoRA: one server keeps one copy of the model and many add-ons, one per customer, and uses the right add-on for each request. Twelve tuned models then cost about one rented GPU, $8 an hour, instead of twelve ($96 an hour).
Your numbersSaved on this device and collected in the Day 8 wrap-up.
Hint: From 3_eval_local.sh, the category_% column. The table shows the mlx fused row (after) first and the mlx Qwen2.5-7B-Instruct-4bit row (before) second. Write before first, for example “78.0 → 96.0”.
Hint: The severity_% column, same two rows: write the second row (before) first.
Hint: The p95_ms column of the mlx fused row.
Hint: The adapter saved: line at the end of 1_train_mlx.sh gives the size of the whole adapters folder. That folder also keeps copies saved during training. For the one file you would ship, run ls -lh adapters/adapters.safetensors in the lesson folder and write its size. The same command works if you stopped training with Ctrl+C.
Hint: From starting 1_train_mlx.sh to the server answering: training, then fusing and loading.
Hint: Pick one kind of customer, for example a bank whose data may not leave its own machines, and write one sentence for them. It is the last line of COMPARISON.md.
Done when
Section titled “Done when”Validation loss has flattened and training has stopped. Your score table has a row for the fine-tune and one for the original model, on the same 60 test tickets. COMPARISON.md has your Mac column filled in next to Day 7’s Fireworks column.
Stuck?
Section titled “Stuck?”- Training stops with an out-of-memory error, or the Mac slows to a crawl.
- Close big apps first. Then in
1_train_mlx.shchange--batch-size 2to--batch-size 1(the script’s own advice) and run it again. If it still fails, also change--num-layers 8to 4: fewer layers with add-ons need less memory, but the add-on can learn less. mlx_lm.lora: command not found- The Python environment from Day 1 is not on. From the
03-labsfolder runsource .venv/bin/activate, thencd lessons/11-lora-local-mlxand run the script again. 2_fuse_and_serve.shfails with ‘Address already in use’.- Another server already uses port 8081, most likely Day 2’s MLX server. Stop it with Ctrl+C in its window, then run the script again.
3_eval_local.shsaysConnection refused.- The server is not up yet: fusing comes first and takes a while. Wait for the line ending
serving on :8081 (Ctrl-C to stop)in the server window, give it a few seconds to load, then run the eval again. Keep that window open. - The eval gets through the fine-tune, then stops with an error at
evaluating mlx mlx-community/Qwen2.5-7B-Instruct-4bit. - Your copy of the MLX tools will not switch models on one server. No table prints, but the fine-tune’s scores are already saved: the last row of
03-labs/results/09-eval.csv, namedmlx fused. Copy itscategory_%,severity_%andp95_msfrom there. Stop the server with Ctrl+C in its window to free memory. In that window, start the original model on a second port:mlx_lm.server --model mlx-community/Qwen2.5-7B-Instruct-4bit --port 8082. Then in the eval window (still in the lesson folder) runMLX_URL=http://localhost:8082/v1 python ../09-eval-harness/evaluate.py "mlx:mlx-community/Qwen2.5-7B-Instruct-4bit" --n 60. It prints the original model’s row. - Validation loss starts rising again before step 300.
- The model has begun memorising the training tickets. Stop with Ctrl+C; the last save is kept. For a rerun, lower
--iters 300in1_train_mlx.shto about the step where the loss was lowest.
Code in this step
Section titled “Code in this step”COMPARISON.md Notes
Build vs buy: the same LoRA, two ways
Section titled “Build vs buy: the same LoRA, two ways”Fill this in after lessons 10 and 11. It helps you compare local and managed training with results you measured yourself, so you can recommend a route to a team.
| Fireworks managed (lesson 10) | MLX on my Mac (lesson 11) | |
|---|---|---|
| base model | ||
| category accuracy: base → tuned | → | → |
| severity accuracy: base → tuned | → | → |
| p95 latency (tuned) | ||
| wall-clock: data → working endpoint | ||
| training cost | $0 | |
| serving cost | $/h dedicated (multi-LoRA can share) | $0, but one user and no SLA |
| data leaves my machine? | yes, to Fireworks | no |
| scales to 100 concurrent users? | yes (replicas, autoscale) | no |
My recommendation for a customer like X: …
One-liner: “I ran the same LoRA locally and managed. Local is free and private; managed got me to a production endpoint in minutes and scales. Most teams prototype locally and ship managed.”
1_train_mlx.sh Bash · 27 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 11 · step 1 — LoRA on your Mac with MLX, on the SAME data as lesson 10.## The base model is already 4-bit, so this is QLoRA: frozen 4-bit weights +# small 16-bit adapters trained on top. That is why a 7B model trains on a laptop.## --data DIR folder with train.jsonl + valid.jsonl (chat "messages" format)# --iters 300 optimizer steps. ~800 rows / batch 2 → 300 iters ≈ 0.75 epoch# --batch-size 2 lower to 1 if you run out of memory# --num-layers 8 adapt only the last 8 layers (fewer = less memory, less capacity)# --adapter-path where the adapter weights (a few MB) are saved## Watch "Val loss" every 50 iters: stop when it stops falling (Ctrl-C keeps the last save).# Time: ~15–40 min on M-series depending on chip. Memory: ~6–10 GB for 7B.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"BASE=${MLX_BASE:-mlx-community/Qwen2.5-7B-Instruct-4bit} # same family as lesson 10's base[[ -f ../../data/tickets/train.jsonl ]] || python ../../data/make_tickets.py
mlx_lm.lora --model "$BASE" --train --data ../../data/tickets \ --iters 300 --batch-size 2 --num-layers 8 \ --steps-per-eval 50 --adapter-path adapters
echo "▸ adapter saved: $(du -sh adapters | cut -f1) — compare that with the base model's size"echo "▸ next: bash 2_fuse_and_serve.sh"2_fuse_and_serve.sh Bash · 17 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 11 · step 2 — serve the fine-tune. Two options, same result:## A) adapter on top of the base at load time (what multi-LoRA servers do)# B) fuse: bake the adapter into the weights (one self-contained model folder)## We fuse (B) because it is the simplest thing to hand over, then serve on :8081.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"BASE=${MLX_BASE:-mlx-community/Qwen2.5-7B-Instruct-4bit}
mlx_lm.fuse --model "$BASE" --adapter-path adapters --save-path fusedecho "▸ fused model in $(pwd)/fused — serving on :8081 (Ctrl-C to stop)"exec mlx_lm.server --model "$(pwd)/fused" --port 8081# option A instead: mlx_lm.server --model "$BASE" --adapter-path adapters --port 80813_eval_local.sh Bash · 16 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 11 · step 3 — base vs local fine-tune: lesson 09's harness, same 60 tickets.## mlx_lm.server loads whichever model a request names (one at a time), so the server# from step 2 can answer for both: first the fused fine-tune, then the base (it reloads# once in between — the first base request is slow, which is fine for an eval).# If your mlx-lm version refuses, run the base on another port and set# MLX_URL=http://localhost:8082/v1 for the second command.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"BASE=${MLX_BASE:-mlx-community/Qwen2.5-7B-Instruct-4bit}
python ../09-eval-harness/evaluate.py "mlx:$(pwd)/fused" "mlx:$BASE" --n 60echo "▸ Compare with lesson 10's Fireworks row in results/09-eval.csv — then fill in COMPARISON.md"qlora_memory.py Python · 37 lines
"""Lesson 11 — why QLoRA fits on a laptop: a back-of-envelope memory budget.
Full fine-tuning keeps, per parameter: the weight (2 B), its gradient (2 B), andAdam's two moments (8 B in fp32) ≈ 12–16 bytes/param → a 7B model needs ~100 GB.
QLoRA freezes the base in 4-bit (~0.56 B/param) and trains only small adapters,so gradients and optimizer state exist ONLY for the adapter.
python qlora_memory.py # 7B, rank 8, last 8 layers python qlora_memory.py --params 70 --rank 16 --layers 80"""import argparse
ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)ap.add_argument("--params", type=float, default=7.6, help="base model size in billions")ap.add_argument("--hidden", type=int, default=3584, help="hidden size (Qwen2.5-7B: 3584)")ap.add_argument("--layers", type=int, default=8, help="layers that get adapters")ap.add_argument("--rank", type=int, default=8)ap.add_argument("--targets", type=int, default=4, help="adapted matrices per layer (q,k,v,o)")ap.add_argument("--act-gb", type=float, default=2.0, help="activations + overhead, rough")a = ap.parse_args()
P = a.params * 1e9full = P * 16 / 1e9 # weights+grads+Adam, mixed precisionlora_params = a.layers * a.targets * 2 * a.hidden * a.rank # A (h×r) + B (r×h) per matrixbase_4bit = P * 0.5625 / 1e9 # 4.5 bits/weight incl. scalesadapter_train = lora_params * 16 / 1e9 # adapter weights+grads+Adamqlora = base_4bit + adapter_train + a.act_gb
print(f"base model {a.params:.1f} B params")print(f"LoRA adapter {lora_params / 1e6:,.1f} M params ({100 * lora_params / P:.3f}% of the model)")print()print(f"full fine-tune ≈ {full:6.1f} GB (multi-GPU territory)")print(f"QLoRA ≈ {qlora:6.1f} GB = 4-bit base {base_4bit:.1f} + adapter training " f"{adapter_train:.2f} + activations {a.act_gb}")print(f"adapter file on disk ≈ {lora_params * 2 / 1e6:6.1f} MB (what you ship / hot-swap in multi-LoRA)")