Serve one model three ways
course 4 of 22
lesson 02 · lessons/02-three-local-servers
35 min Free Mac
What you will do and why
Customers serve models with different server programs, and each hides or exposes different settings. You run one model on three of them on your Mac and time each with the same script.
Why it matters: A customer asks why their chat bot forgets the start of a long conversation. In today’s llama.cpp setup, four chats share room for about 6,000 words, so each gets only about 1,500. You can see and change that limit in llama.cpp; Ollama chooses settings like this for you.
You are done when: bash compare_backends.sh prints a table for each of the three servers (ollama, llamacpp and mlx), and none of them says ‘skipped’.
Start this first
In a new terminal, go to your 03-labs folder (cd ~/Downloads/03-labs, as on Day 1) and run the command above. Leave it running while you watch: its first start downloads the 4.9 GB model file. This is terminal 2 below. For MLX, open another terminal and run cd ~/Downloads/03-labs && source .venv/bin/activate && bash lessons/02-three-local-servers/serve_mlx.sh. That is terminal 3; its first start downloads the MLX copy.
bash lessons/02-three-local-servers/serve_llamacpp.sh Q4_K_M 8192 4Serving With vLLM · 2:56
Download mp4 (11.5 MB)Chapters
In this video What a server program for AI models does and the five settings that matter most, shown on vLLM, a server built for NVIDIA chips.
3 key points
A server program for AI models does 4 jobs. It takes requests, groups them to share the chip, keeps each conversation’s notes (the model’s memory of it so far), and runs the math fast.
Take one away and you have a demo, not a service. vLLM became the default because of the middle two jobs. People join and leave the group at every step (continuous batching), and notes are handed out a page at a time (PagedAttention).
On NVIDIA AI machines such as the DGX Spark (a desktop AI computer), run the server inside a container: a sealed package with the exact software versions it needs.
The chip’s driver, NVIDIA’s CUDA software and the Python packages must all match, and often do not. Installing without a container often needs a documented workaround, which costs 20 to 30% of total output.
Five settings decide most of a server’s performance: how fast it answers and how many people it serves. Three of them took the same machine from 16 conversations at once to 50.
The three: reuse a shared prompt opening (prefix caching), read long prompts in pieces (chunked prefill), and store the notes in 8 bits (FP8 KV cache). Only after that, the video says, argue about hardware.
In plain words
Section titled “In plain words”A model on its own is a big file of numbers. An inference server is the program that loads that file and answers requests from other programs. Ollama, llama.cpp and MLX are three such servers for your Mac. All three accept the same request format, so one timing script can test them all.
Picture it
Three cars carry the same load (the model) around one test track. Ollama is an automatic: it picks the gears for you. llama.cpp is the same motor with a manual gearbox, so you pick every gear. MLX is a different motor, built by the company that made the road (Apple’s software for Apple chips).
With real numbersLlama 3.1 8B, lesson 02’s scripts and Day 1’s sample results
- One model: Llama 3.1 8B (8 billion learned numbers), stored at about 4 bits per number. Ollama and llama.cpp each keep their own copy, a 4.9 GB file.
- Three servers, three doors (ports) on your Mac: Ollama 11434, llama.cpp 8080, MLX 8081.
- llama.cpp’s settings today: room for notes on 8,192 tokens (word pieces; about 6,000 words), shared by 4 seats (4 conversations served at the same time). 8,192 ÷ 4 = 2,048 tokens per seat, about 1,500 words.
- One test for all three: Day 1’s timing script, 5 requests per server, replies of up to 192 tokens (about 140 words).
- Day 1’s sample Ollama result: 44.6 tokens per second, about 33 words a second. MLX is usually faster; llama.cpp lands close to Ollama.
Words to know
- Inference server
- The program that loads a model and answers requests. Example: llama.cpp’s
llama-server. - Port
- A numbered door on your computer where one server listens for requests. Example: llama.cpp uses 8080.
- Notes (KV cache)
- The model’s memory of each conversation so far, kept in the same memory as the model. Example: today’s llama.cpp keeps room for 8,192 tokens of notes.
- Flag
- An option typed after a command that changes how the program runs. Example:
--parallel 4.
From the lesson
lessons/02-three-local-servers/README.md
What and why
Section titled “What and why”An inference server turns a model file into an API. It owns the GPU memory, batches users together, manages the KV cache and streams tokens back. The three you can run on a Mac sit at different points on the convenience-versus-control line:
| server | analogy | you control | port |
|---|---|---|---|
| Ollama | automatic car | almost nothing — sensible defaults | 11434 |
llama.cpp (llama-server) |
manual gearbox | context, slots, quant, cache type, offload | 8080 |
MLX (mlx_lm.server) |
Apple’s own engine | model and quant; plus local fine-tuning | 8081 |
vLLM and SGLang are the production equivalents. They need an NVIDIA GPU, so they appear in Day 9’s optional Spark step, which needs an NVIDIA DGX Spark (a desktop AI computer). All five speak the same OpenAI API, which is why one harness works everywhere.
Read the code first
Section titled “Read the code first”serve_llamacpp.shhas one comment per flag. Each flag is a concept from a later day: slots (seats) on Day 3, quantization (shrinking the model) and context (conversation length) on Day 4, offload (choosing which chip runs each part of the model) on Day 9.felab/targets.pylists every backend in a single table.
cd lessons/02-three-local-servers# terminal 1ollama serve# terminal 2bash serve_llamacpp.sh Q4_K_M 8192 4# terminal 3bash serve_mlx.sh# terminal 4bash compare_backends.shWhat each command does
ollama serveTerminal 1: starts Ollama at port 11434 (a numbered door on your Mac), ready to serve the
llama3.1:8bmodel from Day 1. Leave it running. Look for the window staying open with no error. If it says the address is already in use, Ollama is already running, so skip this terminal.bash serve_llamacpp.sh Q4_K_M 8192 4Terminal 2. Skip it if ‘Start this first’ already started it. This is llama.cpp’s server at port 8080, with the 4-bit file (Q4_K_M, 4.9 GB, downloaded on the first run). It keeps room for 8,192 tokens, split across 4 seats. Look for
ctx=8192andslots=4on its first line, then a line saying it is listening. 8,192 ÷ 4 = 2,048 tokens per seat: check 1 below is about that split.bash serve_mlx.shTerminal 3. Skip it if ‘Start this first’ already started it; the Python environment must be on (see ‘Before you start’). This is Apple’s MLX server with its own 4-bit copy of the model, at port 8081, downloaded on the first run. Look for a line with
export MLX_MODEL=. You need it only if the comparison later complains about the model name.bash compare_backends.shTerminal 4, with the Python environment on (see ‘Before you start’). It times each running server with Day 1’s script: one uncounted warm-up, then 5 timed requests. It prints one small table per server and saves the rows to
results/01-latency.csv. Look for three tables, and compare theirtok_per_scolumn (writing speed).
What you should see
Section titled “What you should see”How to read it
Each server prints one small table with one row. Read the tok_per_s column, the writing speed in tokens (word pieces) per second: MLX usually leads, and Ollama lands close to llama.cpp because it runs llama.cpp underneath. llama.cpp’s row shows the model as local, because its server ignores the name. If Ollama is well behind, ollama show llama3.1:8b shows its copy’s storage size per number (its quantization, such as Q4_K_M): a different one there is the first thing to rule out.
The same 8B 4-bit model, three rows. MLX is usually fastest on decode. llama.cpp and
Ollama are close, because Ollama is llama.cpp underneath. If Ollama is noticeably
slower, it is usually running a different quant or a smaller context by default. Run
ollama show llama3.1:8b to check.
Check yourself
Section titled “Check yourself”2 questions. Say your answer out loud, then tap to check it.
You start llama.cpp with -c 8192 (room for 8,192 tokens, about 6,000 words) and --parallel 4 (4 seats, so 4 people at once). Why is that a trap?Show answerHide
In plain words
The 8,192 tokens are shared out, not given to each person. Each of the 4 seats gets only 2,048 tokens (about 1,500 words), so a longer conversation does not fit.
Picture it
One large pizza ordered ‘for 4 people’. The box says large, but each person gets a quarter, and anyone hungrier than a quarter goes without.
With real numbersLlama 3.1 8B on llama.cpp with Day 2’s settings, lesson 02
- Room for notes:
-c 8192= 8,192 tokens, about 6,000 words. - Seats:
--parallel 4= 4 conversations at once. - Each seat: 8,192 ÷ 4 = 2,048 tokens, about 1,500 words.
- Day 1’s long test prompt, about 8,000 tokens, is almost 4 times one seat’s room (8,000 ÷ 2,048 = 3.9), so it does not fit.
- The fix: set
-cto seats x tokens each. For 4 seats of 8,192 tokens: 4 x 8,192 = 32,768, which takes 4.29 GB of notes for this model (Day 4 works this out).
Words to know
- Context length (
-c) - The most tokens the server keeps notes for. Example:
-c 8192= 8,192 tokens. - Slot (
--parallel) - One seat on the server: one conversation it serves at the same time. Example:
--parallel 4= 4 seats. - Token
- A chunk of text, about three quarters of a word. Example: 2,048 tokens is about 1,500 words.
- KV cache (notes)
- The model’s notes on each conversation so far, kept in memory. Example: 32,768 tokens of notes take 4.29 GB for Llama 3.1 8B.
Go deeper: the engineer version
The kit's question
Why is --parallel 4 -c 8192 a trap?
The kit's answer
llama.cpp splits the context across slots, so each user gets only 2,048 tokens.
More detail: llama-server sets aside one KV cache of -c tokens at startup and splits it evenly across the --parallel slots (the comment in serve_llamacpp.sh says so), so each slot’s context is c ÷ parallel. Size -c as slots x the longest context one user needs: 4 x 8,192 = 32,768, which is 4,096 MiB of f16 cache for Llama 3.1 8B at 128 KiB per token (lesson 05). llama-server’s startup log lists the context each slot gets, so confirm it there on your build.
A customer runs Ollama, the easy one-command server, for their live service with real users (production). What would you ask them about?Show answerHide
In plain words
Ask how many people use it at the same moment, and what happens when many arrive at once. Ask how they watch its health. Then ask whether a server built for many users, such as vLLM or SGLang (servers for NVIDIA chips), would serve them better.
Picture it
A home espresso machine makes great coffee, one cup at a time. If a busy café runs on one, you ask how many customers come at the morning rush, how long the queue gets, and who notices when it breaks. At some point they need a café machine built to pour hundreds of cups an hour.
With real numbersthe Serving With vLLM video (Day 2) and lesson 02
- Day 2’s llama.cpp script shows its seats:
--parallel 4, so 4 people at once with 2,048 tokens each. Ollama chooses settings like these for you, out of sight. - In the Serving With vLLM video, vLLM (a server built for many users) prints its room for notes at startup: 8,234 blocks of 16 tokens each. A block is a fixed-size piece of note memory.
- 8,234 x 16 = 131,744 tokens of notes. At 8,192 tokens per conversation: 131,744 ÷ 8,192 = 16.08, so 16 people at once if everyone uses the full length.
- In a separate example in the same video, three settings took one vLLM machine from 16 people at once to 50: 50 ÷ 16 = about 3 times as many.
- So ask: how many people at the busiest moment, and do numbers like these cover it?
Words to know
- Production
- A model running as a live service for real users. Example: the customer’s support chat.
- People served at once (concurrency)
- How many conversations the server runs at the same moment. Example: 16 before tuning, 50 after, in the Serving With vLLM video.
- Batching
- Serving several people with one pass over the model. Example: 4 seats share each pass in Day 2’s llama.cpp server.
- Observability
- Seeing what a live service is doing, such as speed, errors and memory, while it runs. Example: llama.cpp’s
--metricsflag.
Go deeper: the engineer version
The kit's question
A customer runs Ollama in production. What would you ask about?
The kit's answer
Concurrency, batching limits, observability, and whether vLLM or SGLang would suit multi-user traffic better.
More detail: Turn each topic into a number. Concurrency: peak simultaneous users. Batching limits: how many requests the engine batches (its parallel slots) and the context each slot gets, which Ollama sets by defaults the operator may never have read. Observability: TTFT and ITL at p50 and p95, errors and memory, live (llama.cpp’s --metrics flag publishes counters at /metrics). Then run the same harness against vLLM or SGLang, which add PagedAttention and the tuning flags from the Serving With vLLM video; Day 9’s optional Spark step does exactly that.
Explain what you learned
Section titled “Explain what you learned”Question Our team already serves our model with Ollama. Is that good enough for customers?
One clear answer
Ollama is llama.cpp with the flags hidden. For a customer workload I want the flags, and for multi-user GPU serving I want vLLM or SGLang.
What this means
- “Ollama is llama.cpp”: Ollama is a friendly layer on top of llama.cpp, the engine (the program doing the work). Given the same 4-bit model (each keeps its own 4.9 GB copy), the two run at about the same speed, and your table should show it.
- “with the flags hidden”: it picks the settings (flags) for you, such as how long a conversation may be and how many people share the server.
ollama show llama3.1:8breveals only some of them, such as the model file’s storage size. - “For a customer workload I want the flags”: for a real customer’s job I set those myself. Example:
-c 8192 --parallel 4quietly gives each of 4 people only 2,048 tokens, and I need to see that. - “and for multi-user GPU serving”: when many people share one NVIDIA GPU (the chip that runs the model) in a live service at the same time,
- “I want vLLM or SGLang”: I use servers built for many users at once. In today’s video, Serving With vLLM, three settings took one machine from 16 conversations at once to 50.
Your numbersSaved on this device and collected in the Day 2 wrap-up.
Hint: The tok_per_s value in the ollama table from bash compare_backends.sh.
Hint: The tok_per_s value in the llamacpp table.
Hint: The tok_per_s value in the mlx table.
Done when
Section titled “Done when”bash compare_backends.sh prints a table for each of the three servers (ollama, llamacpp and mlx), and none of them says ‘skipped’.
Stuck?
Section titled “Stuck?”- The comparison says
ollama is not running(or llamacpp, or mlx) and skips it. - That server is not up. Look in its terminal for an error and start it again. From the
03-labsfolder, with the Python environment on,make checklists which servers answer. - All three are skipped although the servers are running, and the address after ‘at’ is blank.
- The script could not load the kit’s Python toolkit. In that terminal, go to the
03-labsfolder, runsource .venv/bin/activate, thenbash lessons/02-three-local-servers/compare_backends.sh(the script moves into its own folder by itself). llama-serverormlx_lm.server: command not found.- For
mlx_lm.server, the Python environment is usually off in that terminal: from the03-labsfolder runsource .venv/bin/activateand try again. If a command is still missing, rerun Day 1’s setup script, which is safe to run again:bash lessons/00-setup/setup_mac.sh, thensource .venv/bin/activate. ollama servesays the address is already in use.- Ollama is already running, probably still from Day 1. Leave it running and skip terminal 1.
- The mlx row fails with an error about the model.
- The model name the script sends must match the one the MLX server loaded. The MLX server prints the exact line to use: run that
export MLX_MODEL=...line in terminal 4, then compare again. - The lab book card’s
BASE_URL=... python bench_ttft.py --model locallines measure the practice server, or fail to connect, instead of your Mac’s servers. - Those lines belong to the lab book’s own short script. Day 1’s script picks a server with
--targetinstead (--target llamacpp,mlxorollama), or runbash compare_backends.sh.
Code in this step
Section titled “Code in this step”compare_backends.sh Bash · 24 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 02 — one harness, three backends, one table.# Runs lesson 01's bench_ttft.py against every local server that is up.# Start the servers first (three terminals):# ollama serve# bash serve_llamacpp.sh# bash serve_mlx.sh# ─────────────────────────────────────────────────────────────────────────────set -uo pipefailcd "$(dirname "$0")"BENCH=../01-latency-harness/bench_ttft.py
for target in ollama llamacpp mlx; do url=$(python -c "from felab import TARGETS; print(TARGETS['$target'].base_url)") if curl -sf "$url/models" >/dev/null; then python "$BENCH" --target "$target" --runs 5 | sed -n '/^target/,/^$/p' else echo " ✗ $target is not running at $url — skipped" fidone
echoecho "Full history: results/01-latency.csv (compare the tok_per_s column)"serve_llamacpp.sh Bash · 24 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 02 — llama.cpp's llama-server: the engine Ollama wraps, with the flags visible.## Every flag here maps to a concept you will be asked about:# -hf repo:QUANT download a GGUF from Hugging Face at that quantization (lesson 04)# -c 8192 context window in tokens → sizes the KV cache (lesson 05)# -ngl 99 offload all layers to the GPU (Metal). Fewer = CPU offload (video 7)# --parallel 4 4 slots = up to 4 sequences batched together (lesson 03)# NOTE: the context is SPLIT across slots → 8192/4 = 2048 each# --port 8080 OpenAI-compatible API at http://localhost:8080/v1## Usage: bash serve_llamacpp.sh [QUANT] [CTX] [SLOTS]# bash serve_llamacpp.sh Q4_K_M 8192 4# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailQUANT=${1:-Q4_K_M}CTX=${2:-8192}SLOTS=${3:-4}REPO=${LLAMACPP_REPO:-bartowski/Meta-Llama-3.1-8B-Instruct-GGUF}
echo "▸ llama-server $REPO:$QUANT ctx=$CTX slots=$SLOTS → http://localhost:8080/v1"exec llama-server -hf "$REPO:$QUANT" -c "$CTX" -ngl 99 --parallel "$SLOTS" --port 8080 \ --metrics # Prometheus metrics at /metrics — handy in lesson 03serve_mlx.sh Bash · 18 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 02 — mlx_lm.server: Apple's MLX runtime, usually the fastest path on M-series.## MLX models on Hugging Face live under mlx-community/, already converted and# quantized (…-4bit, …-8bit). The same library trains LoRA adapters in lesson 11.## The `model` field in API requests must match what you serve here; felab reads# it from MLX_MODEL in .env so the harness sends the right name.## Usage: bash serve_mlx.sh [HF_REPO]# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailMODEL=${1:-${MLX_MODEL:-mlx-community/Meta-Llama-3.1-8B-Instruct-4bit}}
echo "▸ mlx_lm.server $MODEL → http://localhost:8081/v1"echo " (tell the harness: export MLX_MODEL=$MODEL)"exec mlx_lm.server --model "$MODEL" --port 8081