Day 2 · Local stack on the Mac
Part 1 · Day 2 of 12
0 of 12 days done
About 45 min Free Mac
Today: Run one AI model through three server programs on your Mac: Ollama, llama.cpp and MLX. A server program loads the model and answers requests. Then time all three with Day 1’s timing script, and compare their speed and which settings each lets you see.
Example
Imagine a support chat that needs to remember what a customer said earlier. This server reserves room for four chats at once, but gives each only about 1,500 words of memory. A longer conversation may lose its earlier details. Today you will see where that limit comes from and which program lets you change it.
By the end you’ll have
Three writing speeds from the same script, one per server: how many tokens (word pieces) each writes per second.
Expected: MLX usually writes fastest, and Ollama lands close to llama.cpp. Day 1’s sample Ollama row wrote 44.6 tokens per second, about 33 words a second.
The 3 numbers you write down
Ollama: writing speed
How many tokens (word pieces) one person sees appear each second once the reply starts.
Example: Day 1’s sample Ollama row: 44.6, about 33 words a second.
llama.cpp: writing speed
The same measure on llama.cpp, the engine Ollama runs on.
Example: Close to your Ollama figure. llama.cpp cannot beat your Mac’s speed limit for this 4.9 GB file: memory speed ÷ file size. Day 1’s sample Mac: 273 GB per second ÷ 4.9 GB = about 56 tokens per second.
MLX: writing speed
The same measure on Apple’s own MLX server.
Example: Usually the highest of your three figures.
How you will use this: Your benchmark’s second page: one model timed on three servers, the evidence behind today’s line “Ollama is llama.cpp with the flags hidden.” The rows are saved in results/01-latency.csv, and Day 10’s sizing memo (your written hardware recommendation) lists the latest timing rows from that file.
Before you start
Day 1’s setup is in place.
Today reuses what Day 1 installed: Ollama with the Llama 3.1 8B model (the 4.9 GB download), llama.cpp, Apple’s MLX, and the kit’s Python toolkit (felab). Check it from the
03-labsfolder (use your own path if you put it elsewhere on Day 1):cd ~/Downloads/03-labs && source .venv/bin/activate && make checkLook for
ollamamarked UP. If it says down, start it withollama servein its own terminal.llamacppandmlxsay down until you start them today.About 10 GB of free disk space.
llama.cpp and MLX each download their own 4-bit copy of the model the first time they start. llama.cpp’s is 4.9 GB. MLX’s is at least 4 GB: 8 billion numbers x half a byte (4 bits) each. That is about 9 GB; 10 GB leaves a little spare. Ollama reuses its Day 1 copy.
Four terminal windows.
Three servers each run in their own terminal and keep running; the fourth runs the comparison. In each new terminal, run:
cd ~/Downloads/03-labs && source .venv/bin/activate && cd lessons/02-three-local-serversThis goes to the code folder (use your own path if you put it elsewhere on Day 1), turns on the kit’s Python environment ((.venv) appears at the start of the line), and moves into today’s lesson folder. The MLX server and the comparison script both need the environment.
Warm-up from Day 1
Section titled “Warm-up from Day 1”About 2 minutes. Say your answer out loud, then tap to check it. From Day 1 · Guardrails and first calls.
Your Mac has 36 GB of memory. Roughly how big a model, stored at 4 bits per number, can it run comfortably?Show answerHide
In plain words
About a 40-billion-number model (40B). That fills the whole budget of about 25 GB, so there is little room left for the model’s notes on each conversation.
Picture it
Flatmates share one fridge. You can take about 70% of it before everyone else’s food gets squeezed out. Fill your whole share with one giant cake and there is no room for the leftovers each meal adds (the model’s notes).
With real numberslesson 00, and the 4.9 GB model from Day 1’s setup step
- Budget: 0.7 x 36 GB = 25.2 GB, about 25 GB for the model and its notes.
- Size per billion numbers at 4 bits: Llama 3.1 8B is 4.9 GB, so 4.9 ÷ 8 = about 0.61 GB per billion.
- Biggest model: 25 ÷ 0.61 = about 41 billion numbers, so about a 40B model.
- Room left for notes: almost none. Day 4 shows one long conversation’s notes can outgrow the model itself.
Words to know
- Parameters (8B, 40B)
- The model’s learned numbers, also called weights. Example: 8B = 8 billion.
- 4-bit
- Each number stored in 4 bits, half a byte, plus a little extra. Example: an 8B model is 4.9 GB at 4-bit.
- Dense model
- A model that uses all of its numbers for every token. Example: Llama 3.1 8B.
- Notes (KV cache)
- The model’s memory of each conversation so far, kept in the same memory as the model. Example: Day 4 measures it.
Go deeper: the engineer version
The kit's question
Your Mac has 36 GB. Roughly the biggest 4-bit model you can run comfortably?
The kit's answer
≈ 0.7 × 36 ≈ 25 GB of weights, so about a 40B dense model at 4-bit, with little room left for context.
More detail: 0.7 x 36 = 25.2 GB for weights and KV cache together. Real 4-bit files (Q4_K_M) carry per-block scales and keep some tensors at higher precision, so Llama 3.1 8B at 4 bits is 4.9 GB (about 4.8 bits per weight, lesson 04), not the 4.0 GB that 0.5 bytes x 8B would give. 25.2 ÷ 0.61 = about 41B parameters, so about a 40B dense model, which leaves almost no KV-cache room. check_env.py prints memory in billions of bytes (a 36 GB Mac shows 38.7), so the 70% rule is a rough budget either way.
Why does Day 1’s check script (check_env.py) print memory speed (bandwidth) rather than how many GPU cores (the chip’s calculating units) you have?Show answerHide
In plain words
Because memory speed caps how fast the model writes. For each token (word piece) the chip must fetch the whole model from memory once, and that fetching is slower than the math.
Picture it
A very fast cook needs the whole recipe book carried over on a conveyor belt before every dish. More cooks (more cores) would stand waiting; only a faster belt gets dishes out sooner.
With real numberslesson 00’s sample Mac and felab/hardware.py
- Each token needs the whole 4.9 GB model fetched once.
- At 273 GB per second: 273 ÷ 4.9 = about 56 tokens per second, at most.
- Real runs land at 60 to 85% of that limit (felab/hardware.py); Day 4 checks it on your Mac.
- A Mac whose memory moves 546 GB per second (the faster M4 Max) has twice the limit: 546 ÷ 4.9 = about 111.
Words to know
- Decode
- The writing phase: one token at a time, fetching the whole model for each.
- Bandwidth
- How many GB per second memory delivers to the chip. Example: 273 GB/s.
- GPU cores
- The many small calculating units inside a GPU. Example: they do the math, but on a Mac they mostly wait for memory while writing.
- Compute-bound
- Slowed by calculation speed rather than memory. Example: reading the prompt (prefill).
Go deeper: the engineer version
The kit's question
Why does the script print bandwidth rather than GPU cores?
The kit's answer
Decode reads every weight once per token, so bandwidth, not compute, sets the speed.
More detail: At batch size 1, each decode step streams every weight once and does only about two floating-point operations per weight, far less than the GPU could compute in that time. So the ceiling is bandwidth ÷ bytes read per token (felab/hardware.py), and measured decode typically reaches 60 to 85% of it. Prefill is the opposite: it reuses each weight across many prompt tokens, so compute sets its pace. Where a chip ships in two bandwidth versions (M3 Max, M4 Max), hardware.py lists the faster one.
A customer says their AI feature is ‘slow’. Which number do you ask for first, the wait for the first word or the pace after it, and why?Show answerHide
In plain words
Both, because each points to a different cause. Slow to start means a long prompt or requests waiting in line; slow to type means memory speed, or a model too big for it.
Picture it
A patient says ‘I feel unwell’. The doctor takes both temperature and blood pressure, because each points to different illnesses. ‘Slow’ is the symptom; the two numbers are the tests.
With real numbersthe support assistant in the Inference 101 video
- A customer asks the support bot about a refund. It reads the company rules and question before it can reply: 6,200 tokens (word pieces) take 0.78 s.
- The bot then writes a 300-token answer, one piece every 22 ms: 300 x 22 ms = 6.6 s.
- The customer waits about 7.4 s in all. Of that, 6.6 s is spent watching the reply appear: about 9 of every 10 seconds of this wait.
- Even if reading became instant, the customer would save less than 0.78 s. To shorten this whole reply much more, its writing needs to speed up.
Words to know
- Prefill
- The reading phase: the whole prompt at once. Example: 6,200 tokens in 0.78 s.
- Decode
- The writing phase, one token at a time. Example: 300 tokens x 22 ms = 6.6 s.
- Queueing
- Requests waiting in line because the server is busy with others.
- Bandwidth
- How many GB per second memory delivers to the chip; it caps decode.
Go deeper: the engineer version
The kit's question
The customer says “it’s slow”. Which number do you ask for first, and why?
The kit's answer
Both. Slow to start points at prefill, queueing or a long prompt. Slow to type points at decode, bandwidth or an oversized model.
More detail: TTFT is roughly queue time plus prefill time, so a high TTFT sends you to prompt length, prefix caching, cold starts and queueing. ITL is one decode step, bound by memory bandwidth and the bytes read per token, so a high ITL sends you to model size, quantization, speculative decoding or batch size (bigger batches raise total throughput but slow each user a little). Ask for both, at p50 and p95, on the customer’s real prompt and output lengths.
Today’s steps
Section titled “Today’s steps”1 step, then the wrap-up.
-
Serve one model three ways
Customers serve models with different server programs, and each hides or exposes different settings. You run one model on three of them on your Mac and time each with the same script.
Video: Serving With vLLM 2:56
- Wrap-up and drill Put your numbers in one table, answer 2 questions out loud, explain 1 result and work 2 examples.
Your progress is saved on this device.