Skip to content

Serve one model three ways

course 4 of 22

lesson 02 · lessons/02-three-local-servers

35 min Free Mac

What you will do and why

Customers serve models with different server programs, and each hides or exposes different settings. You run one model on three of them on your Mac and time each with the same script.

Why it matters: A customer asks why their chat bot forgets the start of a long conversation. In today’s llama.cpp setup, four chats share room for about 6,000 words, so each gets only about 1,500. You can see and change that limit in llama.cpp; Ollama chooses settings like this for you.

You are done when: bash compare_backends.sh prints a table for each of the three servers (ollama, llamacpp and mlx), and none of them says ‘skipped’.

Start this first

In a new terminal, go to your 03-labs folder (cd ~/Downloads/03-labs, as on Day 1) and run the command above. Leave it running while you watch: its first start downloads the 4.9 GB model file. This is terminal 2 below. For MLX, open another terminal and run cd ~/Downloads/03-labs && source .venv/bin/activate && bash lessons/02-three-local-servers/serve_mlx.sh. That is terminal 3; its first start downloads the MLX copy.

Terminal window
bash lessons/02-three-local-servers/serve_llamacpp.sh Q4_K_M 8192 4

Serving With vLLM · 2:56

Download mp4 (11.5 MB)

Chapters

In this video What a server program for AI models does and the five settings that matter most, shown on vLLM, a server built for NVIDIA chips.

3 key points

  1. A server program for AI models does 4 jobs. It takes requests, groups them to share the chip, keeps each conversation’s notes (the model’s memory of it so far), and runs the math fast.

    Take one away and you have a demo, not a service. vLLM became the default because of the middle two jobs. People join and leave the group at every step (continuous batching), and notes are handed out a page at a time (PagedAttention).

  2. On NVIDIA AI machines such as the DGX Spark (a desktop AI computer), run the server inside a container: a sealed package with the exact software versions it needs.

    The chip’s driver, NVIDIA’s CUDA software and the Python packages must all match, and often do not. Installing without a container often needs a documented workaround, which costs 20 to 30% of total output.

  3. Five settings decide most of a server’s performance: how fast it answers and how many people it serves. Three of them took the same machine from 16 conversations at once to 50.

    The three: reuse a shared prompt opening (prefix caching), read long prompts in pieces (chunked prefill), and store the notes in 8 bits (FP8 KV cache). Only after that, the video says, argue about hardware.

A model on its own is a big file of numbers. An inference server is the program that loads that file and answers requests from other programs. Ollama, llama.cpp and MLX are three such servers for your Mac. All three accept the same request format, so one timing script can test them all.

Picture it

Three cars carry the same load (the model) around one test track. Ollama is an automatic: it picks the gears for you. llama.cpp is the same motor with a manual gearbox, so you pick every gear. MLX is a different motor, built by the company that made the road (Apple’s software for Apple chips).

With real numbersLlama 3.1 8B, lesson 02’s scripts and Day 1’s sample results

  • One model: Llama 3.1 8B (8 billion learned numbers), stored at about 4 bits per number. Ollama and llama.cpp each keep their own copy, a 4.9 GB file.
  • Three servers, three doors (ports) on your Mac: Ollama 11434, llama.cpp 8080, MLX 8081.
  • llama.cpp’s settings today: room for notes on 8,192 tokens (word pieces; about 6,000 words), shared by 4 seats (4 conversations served at the same time). 8,192 ÷ 4 = 2,048 tokens per seat, about 1,500 words.
  • One test for all three: Day 1’s timing script, 5 requests per server, replies of up to 192 tokens (about 140 words).
  • Day 1’s sample Ollama result: 44.6 tokens per second, about 33 words a second. MLX is usually faster; llama.cpp lands close to Ollama.

Words to know

Inference server
The program that loads a model and answers requests. Example: llama.cpp’s llama-server.
Port
A numbered door on your computer where one server listens for requests. Example: llama.cpp uses 8080.
Notes (KV cache)
The model’s memory of each conversation so far, kept in the same memory as the model. Example: today’s llama.cpp keeps room for 8,192 tokens of notes.
Flag
An option typed after a command that changes how the program runs. Example: --parallel 4.

From the lesson

lessons/02-three-local-servers/README.md

An inference server turns a model file into an API. It owns the GPU memory, batches users together, manages the KV cache and streams tokens back. The three you can run on a Mac sit at different points on the convenience-versus-control line:

server analogy you control port
Ollama automatic car almost nothing — sensible defaults 11434
llama.cpp (llama-server) manual gearbox context, slots, quant, cache type, offload 8080
MLX (mlx_lm.server) Apple’s own engine model and quant; plus local fine-tuning 8081

vLLM and SGLang are the production equivalents. They need an NVIDIA GPU, so they appear in Day 9’s optional Spark step, which needs an NVIDIA DGX Spark (a desktop AI computer). All five speak the same OpenAI API, which is why one harness works everywhere.

  • serve_llamacpp.sh has one comment per flag. Each flag is a concept from a later day: slots (seats) on Day 3, quantization (shrinking the model) and context (conversation length) on Day 4, offload (choosing which chip runs each part of the model) on Day 9.
  • felab/targets.py lists every backend in a single table.
Terminal window
cd lessons/02-three-local-servers
# terminal 1
ollama serve
# terminal 2
bash serve_llamacpp.sh Q4_K_M 8192 4
# terminal 3
bash serve_mlx.sh
# terminal 4
bash compare_backends.sh

What each command does

  1. ollama serve

    Terminal 1: starts Ollama at port 11434 (a numbered door on your Mac), ready to serve the llama3.1:8b model from Day 1. Leave it running. Look for the window staying open with no error. If it says the address is already in use, Ollama is already running, so skip this terminal.

  2. bash serve_llamacpp.sh Q4_K_M 8192 4

    Terminal 2. Skip it if ‘Start this first’ already started it. This is llama.cpp’s server at port 8080, with the 4-bit file (Q4_K_M, 4.9 GB, downloaded on the first run). It keeps room for 8,192 tokens, split across 4 seats. Look for ctx=8192 and slots=4 on its first line, then a line saying it is listening. 8,192 ÷ 4 = 2,048 tokens per seat: check 1 below is about that split.

  3. bash serve_mlx.sh

    Terminal 3. Skip it if ‘Start this first’ already started it; the Python environment must be on (see ‘Before you start’). This is Apple’s MLX server with its own 4-bit copy of the model, at port 8081, downloaded on the first run. Look for a line with export MLX_MODEL=. You need it only if the comparison later complains about the model name.

  4. bash compare_backends.sh

    Terminal 4, with the Python environment on (see ‘Before you start’). It times each running server with Day 1’s script: one uncounted warm-up, then 5 timed requests. It prints one small table per server and saves the rows to results/01-latency.csv. Look for three tables, and compare their tok_per_s column (writing speed).

How to read it

Each server prints one small table with one row. Read the tok_per_s column, the writing speed in tokens (word pieces) per second: MLX usually leads, and Ollama lands close to llama.cpp because it runs llama.cpp underneath. llama.cpp’s row shows the model as local, because its server ignores the name. If Ollama is well behind, ollama show llama3.1:8b shows its copy’s storage size per number (its quantization, such as Q4_K_M): a different one there is the first thing to rule out.

The same 8B 4-bit model, three rows. MLX is usually fastest on decode. llama.cpp and Ollama are close, because Ollama is llama.cpp underneath. If Ollama is noticeably slower, it is usually running a different quant or a smaller context by default. Run ollama show llama3.1:8b to check.

2 questions. Say your answer out loud, then tap to check it.

You start llama.cpp with -c 8192 (room for 8,192 tokens, about 6,000 words) and --parallel 4 (4 seats, so 4 people at once). Why is that a trap?Show answerHide

In plain words

The 8,192 tokens are shared out, not given to each person. Each of the 4 seats gets only 2,048 tokens (about 1,500 words), so a longer conversation does not fit.

Picture it

One large pizza ordered ‘for 4 people’. The box says large, but each person gets a quarter, and anyone hungrier than a quarter goes without.

With real numbersLlama 3.1 8B on llama.cpp with Day 2’s settings, lesson 02

  • Room for notes: -c 8192 = 8,192 tokens, about 6,000 words.
  • Seats: --parallel 4 = 4 conversations at once.
  • Each seat: 8,192 ÷ 4 = 2,048 tokens, about 1,500 words.
  • Day 1’s long test prompt, about 8,000 tokens, is almost 4 times one seat’s room (8,000 ÷ 2,048 = 3.9), so it does not fit.
  • The fix: set -c to seats x tokens each. For 4 seats of 8,192 tokens: 4 x 8,192 = 32,768, which takes 4.29 GB of notes for this model (Day 4 works this out).

Words to know

Context length (-c)
The most tokens the server keeps notes for. Example: -c 8192 = 8,192 tokens.
Slot (--parallel)
One seat on the server: one conversation it serves at the same time. Example: --parallel 4 = 4 seats.
Token
A chunk of text, about three quarters of a word. Example: 2,048 tokens is about 1,500 words.
KV cache (notes)
The model’s notes on each conversation so far, kept in memory. Example: 32,768 tokens of notes take 4.29 GB for Llama 3.1 8B.
Go deeper: the engineer version

The kit's question

Why is --parallel 4 -c 8192 a trap?

The kit's answer

llama.cpp splits the context across slots, so each user gets only 2,048 tokens.

More detail: llama-server sets aside one KV cache of -c tokens at startup and splits it evenly across the --parallel slots (the comment in serve_llamacpp.sh says so), so each slot’s context is c ÷ parallel. Size -c as slots x the longest context one user needs: 4 x 8,192 = 32,768, which is 4,096 MiB of f16 cache for Llama 3.1 8B at 128 KiB per token (lesson 05). llama-server’s startup log lists the context each slot gets, so confirm it there on your build.

A customer runs Ollama, the easy one-command server, for their live service with real users (production). What would you ask them about?Show answerHide

In plain words

Ask how many people use it at the same moment, and what happens when many arrive at once. Ask how they watch its health. Then ask whether a server built for many users, such as vLLM or SGLang (servers for NVIDIA chips), would serve them better.

Picture it

A home espresso machine makes great coffee, one cup at a time. If a busy café runs on one, you ask how many customers come at the morning rush, how long the queue gets, and who notices when it breaks. At some point they need a café machine built to pour hundreds of cups an hour.

With real numbersthe Serving With vLLM video (Day 2) and lesson 02

  • Day 2’s llama.cpp script shows its seats: --parallel 4, so 4 people at once with 2,048 tokens each. Ollama chooses settings like these for you, out of sight.
  • In the Serving With vLLM video, vLLM (a server built for many users) prints its room for notes at startup: 8,234 blocks of 16 tokens each. A block is a fixed-size piece of note memory.
  • 8,234 x 16 = 131,744 tokens of notes. At 8,192 tokens per conversation: 131,744 ÷ 8,192 = 16.08, so 16 people at once if everyone uses the full length.
  • In a separate example in the same video, three settings took one vLLM machine from 16 people at once to 50: 50 ÷ 16 = about 3 times as many.
  • So ask: how many people at the busiest moment, and do numbers like these cover it?

Words to know

Production
A model running as a live service for real users. Example: the customer’s support chat.
People served at once (concurrency)
How many conversations the server runs at the same moment. Example: 16 before tuning, 50 after, in the Serving With vLLM video.
Batching
Serving several people with one pass over the model. Example: 4 seats share each pass in Day 2’s llama.cpp server.
Observability
Seeing what a live service is doing, such as speed, errors and memory, while it runs. Example: llama.cpp’s --metrics flag.
Go deeper: the engineer version

The kit's question

A customer runs Ollama in production. What would you ask about?

The kit's answer

Concurrency, batching limits, observability, and whether vLLM or SGLang would suit multi-user traffic better.

More detail: Turn each topic into a number. Concurrency: peak simultaneous users. Batching limits: how many requests the engine batches (its parallel slots) and the context each slot gets, which Ollama sets by defaults the operator may never have read. Observability: TTFT and ITL at p50 and p95, errors and memory, live (llama.cpp’s --metrics flag publishes counters at /metrics). Then run the same harness against vLLM or SGLang, which add PagedAttention and the tuning flags from the Serving With vLLM video; Day 9’s optional Spark step does exactly that.

Question Our team already serves our model with Ollama. Is that good enough for customers?

One clear answer

Ollama is llama.cpp with the flags hidden. For a customer workload I want the flags, and for multi-user GPU serving I want vLLM or SGLang.

What this means

  • “Ollama is llama.cpp”: Ollama is a friendly layer on top of llama.cpp, the engine (the program doing the work). Given the same 4-bit model (each keeps its own 4.9 GB copy), the two run at about the same speed, and your table should show it.
  • “with the flags hidden”: it picks the settings (flags) for you, such as how long a conversation may be and how many people share the server. ollama show llama3.1:8b reveals only some of them, such as the model file’s storage size.
  • “For a customer workload I want the flags”: for a real customer’s job I set those myself. Example: -c 8192 --parallel 4 quietly gives each of 4 people only 2,048 tokens, and I need to see that.
  • “and for multi-user GPU serving”: when many people share one NVIDIA GPU (the chip that runs the model) in a live service at the same time,
  • “I want vLLM or SGLang”: I use servers built for many users at once. In today’s video, Serving With vLLM, three settings took one machine from 16 conversations at once to 50.

Your numbersSaved on this device and collected in the Day 2 wrap-up.

Hint: The tok_per_s value in the ollama table from bash compare_backends.sh.

Hint: The tok_per_s value in the llamacpp table.

Hint: The tok_per_s value in the mlx table.

bash compare_backends.sh prints a table for each of the three servers (ollama, llamacpp and mlx), and none of them says ‘skipped’.

The comparison says ollama is not running (or llamacpp, or mlx) and skips it.
That server is not up. Look in its terminal for an error and start it again. From the 03-labs folder, with the Python environment on, make check lists which servers answer.
All three are skipped although the servers are running, and the address after ‘at’ is blank.
The script could not load the kit’s Python toolkit. In that terminal, go to the 03-labs folder, run source .venv/bin/activate, then bash lessons/02-three-local-servers/compare_backends.sh (the script moves into its own folder by itself).
llama-server or mlx_lm.server: command not found.
For mlx_lm.server, the Python environment is usually off in that terminal: from the 03-labs folder run source .venv/bin/activate and try again. If a command is still missing, rerun Day 1’s setup script, which is safe to run again: bash lessons/00-setup/setup_mac.sh, then source .venv/bin/activate.
ollama serve says the address is already in use.
Ollama is already running, probably still from Day 1. Leave it running and skip terminal 1.
The mlx row fails with an error about the model.
The model name the script sends must match the one the MLX server loaded. The MLX server prints the exact line to use: run that export MLX_MODEL=... line in terminal 4, then compare again.
The lab book card’s BASE_URL=... python bench_ttft.py --model local lines measure the practice server, or fail to connect, instead of your Mac’s servers.
Those lines belong to the lab book’s own short script. Day 1’s script picks a server with --target instead (--target llamacpp, mlx or ollama), or run bash compare_backends.sh.
compare_backends.sh Bash · 24 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 02 — one harness, three backends, one table.
# Runs lesson 01's bench_ttft.py against every local server that is up.
# Start the servers first (three terminals):
# ollama serve
# bash serve_llamacpp.sh
# bash serve_mlx.sh
# ─────────────────────────────────────────────────────────────────────────────
set -uo pipefail
cd "$(dirname "$0")"
BENCH=../01-latency-harness/bench_ttft.py
for target in ollama llamacpp mlx; do
url=$(python -c "from felab import TARGETS; print(TARGETS['$target'].base_url)")
if curl -sf "$url/models" >/dev/null; then
python "$BENCH" --target "$target" --runs 5 | sed -n '/^target/,/^$/p'
else
echo " ✗ $target is not running at $url — skipped"
fi
done
echo
echo "Full history: results/01-latency.csv (compare the tok_per_s column)"
serve_llamacpp.sh Bash · 24 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 02 — llama.cpp's llama-server: the engine Ollama wraps, with the flags visible.
#
# Every flag here maps to a concept you will be asked about:
# -hf repo:QUANT download a GGUF from Hugging Face at that quantization (lesson 04)
# -c 8192 context window in tokens → sizes the KV cache (lesson 05)
# -ngl 99 offload all layers to the GPU (Metal). Fewer = CPU offload (video 7)
# --parallel 4 4 slots = up to 4 sequences batched together (lesson 03)
# NOTE: the context is SPLIT across slots → 8192/4 = 2048 each
# --port 8080 OpenAI-compatible API at http://localhost:8080/v1
#
# Usage: bash serve_llamacpp.sh [QUANT] [CTX] [SLOTS]
# bash serve_llamacpp.sh Q4_K_M 8192 4
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
QUANT=${1:-Q4_K_M}
CTX=${2:-8192}
SLOTS=${3:-4}
REPO=${LLAMACPP_REPO:-bartowski/Meta-Llama-3.1-8B-Instruct-GGUF}
echo "▸ llama-server $REPO:$QUANT ctx=$CTX slots=$SLOTS → http://localhost:8080/v1"
exec llama-server -hf "$REPO:$QUANT" -c "$CTX" -ngl 99 --parallel "$SLOTS" --port 8080 \
--metrics # Prometheus metrics at /metrics — handy in lesson 03
serve_mlx.sh Bash · 18 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 02 — mlx_lm.server: Apple's MLX runtime, usually the fastest path on M-series.
#
# MLX models on Hugging Face live under mlx-community/, already converted and
# quantized (…-4bit, …-8bit). The same library trains LoRA adapters in lesson 11.
#
# The `model` field in API requests must match what you serve here; felab reads
# it from MLX_MODEL in .env so the harness sends the right name.
#
# Usage: bash serve_mlx.sh [HF_REPO]
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
MODEL=${1:-${MLX_MODEL:-mlx-community/Meta-Llama-3.1-8B-Instruct-4bit}}
echo "▸ mlx_lm.server $MODEL → http://localhost:8081/v1"
echo " (tell the harness: export MLX_MODEL=$MODEL)"
exec mlx_lm.server --model "$MODEL" --port 8081