QuestionStep 2 · Set up your Mac and the spend cap
Your Mac has 36 GB of memory. Roughly how big a model, stored at 4 bits per number, can it run comfortably?
Show answerHide answer
In plain words
About a 40-billion-number model (40B). That fills the whole budget of about 25 GB, so there is little room left for the model’s notes on each conversation.
Picture it
Flatmates share one fridge. You can take about 70% of it before everyone else’s food gets squeezed out. Fill your whole share with one giant cake and there is no room for the leftovers each meal adds (the model’s notes).
With real numberslesson 00, and the 4.9 GB model from Day 1’s setup step
- Budget: 0.7 x 36 GB = 25.2 GB, about 25 GB for the model and its notes.
- Size per billion numbers at 4 bits: Llama 3.1 8B is 4.9 GB, so 4.9 ÷ 8 = about 0.61 GB per billion.
- Biggest model: 25 ÷ 0.61 = about 41 billion numbers, so about a 40B model.
- Room left for notes: almost none. Day 4 shows one long conversation’s notes can outgrow the model itself.
Words to know
- Parameters (8B, 40B)
- The model’s learned numbers, also called weights. Example: 8B = 8 billion.
- 4-bit
- Each number stored in 4 bits, half a byte, plus a little extra. Example: an 8B model is 4.9 GB at 4-bit.
- Dense model
- A model that uses all of its numbers for every token. Example: Llama 3.1 8B.
- Notes (KV cache)
- The model’s memory of each conversation so far, kept in the same memory as the model. Example: Day 4 measures it.
Go deeper: the engineer version
The kit's question
Your Mac has 36 GB. Roughly the biggest 4-bit model you can run comfortably?
The kit's answer
≈ 0.7 × 36 ≈ 25 GB of weights, so about a 40B dense model at 4-bit, with little room left for context.
More detail: 0.7 x 36 = 25.2 GB for weights and KV cache together. Real 4-bit files (Q4_K_M) carry per-block scales and keep some tensors at higher precision, so Llama 3.1 8B at 4 bits is 4.9 GB (about 4.8 bits per weight, lesson 04), not the 4.0 GB that 0.5 bytes x 8B would give. 25.2 ÷ 0.61 = about 41B parameters, so about a 40B dense model, which leaves almost no KV-cache room. check_env.py prints memory in billions of bytes (a 36 GB Mac shows 38.7), so the 70% rule is a rough budget either way.
How did you do?