Skip to content

Day 2 wrap-up and drill

  1. Overview
  2. Step 1
  3. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

The same timing script has measured all three servers on your Mac, and your table shows each one’s writing speed in tokens per second.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 1 · Serve one model three ways

How many tokens (word pieces) one person sees appear each second once the reply starts.

Example: Day 1’s sample Ollama row: 44.6, about 33 words a second.

The same measure on llama.cpp, the engine Ollama runs on.

Example: Close to your Ollama figure. llama.cpp cannot beat your Mac’s speed limit for this 4.9 GB file: memory speed ÷ file size. Day 1’s sample Mac: 273 GB per second ÷ 4.9 GB = about 56 tokens per second.

The same measure on Apple’s own MLX server.

Example: Usually the highest of your three figures.

Every question from today’s steps. Say your answer out loud, then show the answer and rate yourself. 2 cards

QuestionStep 1 · Serve one model three ways

You start llama.cpp with -c 8192 (room for 8,192 tokens, about 6,000 words) and --parallel 4 (4 seats, so 4 people at once). Why is that a trap?

Show answerHide answer

In plain words

The 8,192 tokens are shared out, not given to each person. Each of the 4 seats gets only 2,048 tokens (about 1,500 words), so a longer conversation does not fit.

Picture it

One large pizza ordered ‘for 4 people’. The box says large, but each person gets a quarter, and anyone hungrier than a quarter goes without.

With real numbersLlama 3.1 8B on llama.cpp with Day 2’s settings, lesson 02

  • Room for notes: -c 8192 = 8,192 tokens, about 6,000 words.
  • Seats: --parallel 4 = 4 conversations at once.
  • Each seat: 8,192 ÷ 4 = 2,048 tokens, about 1,500 words.
  • Day 1’s long test prompt, about 8,000 tokens, is almost 4 times one seat’s room (8,000 ÷ 2,048 = 3.9), so it does not fit.
  • The fix: set -c to seats x tokens each. For 4 seats of 8,192 tokens: 4 x 8,192 = 32,768, which takes 4.29 GB of notes for this model (Day 4 works this out).

Words to know

Context length (-c)
The most tokens the server keeps notes for. Example: -c 8192 = 8,192 tokens.
Slot (--parallel)
One seat on the server: one conversation it serves at the same time. Example: --parallel 4 = 4 seats.
Token
A chunk of text, about three quarters of a word. Example: 2,048 tokens is about 1,500 words.
KV cache (notes)
The model’s notes on each conversation so far, kept in memory. Example: 32,768 tokens of notes take 4.29 GB for Llama 3.1 8B.
Go deeper: the engineer version

The kit's question

Why is --parallel 4 -c 8192 a trap?

The kit's answer

llama.cpp splits the context across slots, so each user gets only 2,048 tokens.

More detail: llama-server sets aside one KV cache of -c tokens at startup and splits it evenly across the --parallel slots (the comment in serve_llamacpp.sh says so), so each slot’s context is c ÷ parallel. Size -c as slots x the longest context one user needs: 4 x 8,192 = 32,768, which is 4,096 MiB of f16 cache for Llama 3.1 8B at 128 KiB per token (lesson 05). llama-server’s startup log lists the context each slot gets, so confirm it there on your build.

How did you do?

QuestionStep 1 · Serve one model three ways

A customer runs Ollama, the easy one-command server, for their live service with real users (production). What would you ask them about?

Show answerHide answer

In plain words

Ask how many people use it at the same moment, and what happens when many arrive at once. Ask how they watch its health. Then ask whether a server built for many users, such as vLLM or SGLang (servers for NVIDIA chips), would serve them better.

Picture it

A home espresso machine makes great coffee, one cup at a time. If a busy café runs on one, you ask how many customers come at the morning rush, how long the queue gets, and who notices when it breaks. At some point they need a café machine built to pour hundreds of cups an hour.

With real numbersthe Serving With vLLM video (Day 2) and lesson 02

  • Day 2’s llama.cpp script shows its seats: --parallel 4, so 4 people at once with 2,048 tokens each. Ollama chooses settings like these for you, out of sight.
  • In the Serving With vLLM video, vLLM (a server built for many users) prints its room for notes at startup: 8,234 blocks of 16 tokens each. A block is a fixed-size piece of note memory.
  • 8,234 x 16 = 131,744 tokens of notes. At 8,192 tokens per conversation: 131,744 ÷ 8,192 = 16.08, so 16 people at once if everyone uses the full length.
  • In a separate example in the same video, three settings took one vLLM machine from 16 people at once to 50: 50 ÷ 16 = about 3 times as many.
  • So ask: how many people at the busiest moment, and do numbers like these cover it?

Words to know

Production
A model running as a live service for real users. Example: the customer’s support chat.
People served at once (concurrency)
How many conversations the server runs at the same moment. Example: 16 before tuning, 50 after, in the Serving With vLLM video.
Batching
Serving several people with one pass over the model. Example: 4 seats share each pass in Day 2’s llama.cpp server.
Observability
Seeing what a live service is doing, such as speed, errors and memory, while it runs. Example: llama.cpp’s --metrics flag.
Go deeper: the engineer version

The kit's question

A customer runs Ollama in production. What would you ask about?

The kit's answer

Concurrency, batching limits, observability, and whether vLLM or SGLang would suit multi-user traffic better.

More detail: Turn each topic into a number. Concurrency: peak simultaneous users. Batching limits: how many requests the engine batches (its parallel slots) and the context each slot gets, which Ollama sets by defaults the operator may never have read. Observability: TTFT and ITL at p50 and p95, errors and memory, live (llama.cpp’s --metrics flag publishes counters at /metrics). Then run the same harness against vLLM or SGLang, which add PagedAttention and the tuning flags from the Serving With vLLM video; Day 9’s optional Spark step does exactly that.

How did you do?

Under 20 seconds each. Record yourself once and listen back.

Step 1 · Serve one model three ways

Question Our team already serves our model with Ollama. Is that good enough for customers?

One clear answer

Ollama is llama.cpp with the flags hidden. For a customer workload I want the flags, and for multi-user GPU serving I want vLLM or SGLang.

What this means
  • “Ollama is llama.cpp”: Ollama is a friendly layer on top of llama.cpp, the engine (the program doing the work). Given the same 4-bit model (each keeps its own 4.9 GB copy), the two run at about the same speed, and your table should show it.
  • “with the flags hidden”: it picks the settings (flags) for you, such as how long a conversation may be and how many people share the server. ollama show llama3.1:8b reveals only some of them, such as the model file’s storage size.
  • “For a customer workload I want the flags”: for a real customer’s job I set those myself. Example: -c 8192 --parallel 4 quietly gives each of 4 people only 2,048 tokens, and I need to see that.
  • “and for multi-user GPU serving”: when many people share one NVIDIA GPU (the chip that runs the model) in a live service at the same time,
  • “I want vLLM or SGLang”: I use servers built for many users at once. In today’s video, Serving With vLLM, three settings took one machine from 16 conversations at once to 50.

2 worked examples from today’s videos. Work each on paper, then reveal.

From the video: Serving With vLLM 2:56

Whiteboard 1 of 2

A server program (vLLM) starts up. It prints how much memory the model took (8.4 GB) and how many blocks are left for conversation notes (8,234, each holding 16 tokens). Each conversation may hold up to 8,192 tokens. How many conversations fit at once? If that is too few, what do you change?

Given, in plain words

A block is a fixed-size piece of memory for notes: it holds the notes on 16 tokens (word pieces). The server was started with --max-model-len 8192, so no conversation may be longer than 8,192 tokens (about 6,000 words). Plan for the worst case: every conversation uses all 8,192.

Reveal the answerHide the answer

Answer · in plain words

About 16 conversations fit at once. To fit more, store the notes in 8 bits, or halve the longest allowed conversation to 4,096 tokens. Each change doubles it to about 32; both together give about 64. You change settings, not the GPU (the chip).

Picture it

A storage building has 8,234 lockers, and each locker holds 16 boxes. Every tenant may need up to 8,192 boxes, which is 512 lockers, so the building can promise space to only 16 tenants. Pack the boxes twice as tight, or cap each tenant at half as many boxes, and it can promise 32.

Careful

The video does not name the model or the GPU. Before you lower the limit, check how long the customer’s real conversations are: a limit below what they need breaks their product.

Worked answer, step by step

  1. Tokens of notes the memory holds: 8,234 blocks x 16 tokens = 131,744 tokens (the video rounds this to about 131,000).
  2. Blocks one full conversation needs: 8,192 ÷ 16 = 512 blocks.
  3. Conversations at once, worst case: 131,744 ÷ 8,192 = 16.08, so 16 (the same as 8,234 ÷ 512).
  4. Change 1, notes in 8 bits (--kv-cache-dtype fp8): each token’s notes take half the space, so the same memory holds 131,744 x 2 = 263,488 tokens. 263,488 ÷ 8,192 = 32 conversations.
  5. Change 2, cap conversations at 4,096 tokens (--max-model-len 4096): 131,744 ÷ 4,096 = 32 conversations.
  6. Both together: 263,488 ÷ 4,096 = 64 conversations, four times the 16, on the same GPU.
Go deeper: the engineer version

The kit's question

Reading the startup log · vLLM prints this within the first ten seconds · model weights 8.4 GB · # GPU blocks 8,234 · block size 16 tokens

The kit's answer

cache capacity 8,234 × 16 ≈ 131,000 tokens · at 8K context ≈ 16 concurrent requests · Too low? Change max-model-len or KV dtype — not the GPU

More detail: vLLM loads the weights, then turns what is left of the memory it may claim (--gpu-memory-utilization, 0.8 in the video’s command), after the weights and its own working memory, into KV blocks of 16 tokens. Capacity in tokens = blocks x block size; worst-case concurrency = capacity ÷ max-model-len = 131,744 ÷ 8,192 = 16.08. With PagedAttention, blocks are handed out as conversations grow, so real concurrency with shorter conversations is higher; 16 is what you can promise if everyone fills the window. FP8 KV halves the bytes per token, so the same memory holds twice as many blocks. 8.4 GB of weights is about what an 8-billion-parameter model takes at 1 byte per parameter, the format of the kit’s own vLLM model (nvidia/Llama-3.1-8B-Instruct-FP8 in felab/targets.py), but the video does not say which model it is.

Words to know

Block
A fixed-size piece of note memory. Example: 16 tokens per block in vLLM.
Notes (KV cache)
The model’s memory of each conversation so far. Example: 512 blocks for one full 8,192-token conversation here.
Max model length
The longest conversation the server allows. Example: --max-model-len 8192.
Startup log
What a server prints as it starts. Example: vLLM’s shows the weights (8.4 GB) and the blocks left (8,234).

From the video: Serving With vLLM 2:56

Whiteboard 2 of 2

In Serving With vLLM, one model on default settings handles 16 conversations at once. When prompts share the same opening, 95 of 100 requests wait up to 1.6 seconds for the first token (word piece). Which three settings do you try first, and what did the video measure after?

Given, in plain words

Same machine and model before and after; only three settings change. The prompts share a common opening (a shared prefix), such as the same instructions at the top of every request. The 1.6 seconds is p95 TTFT: the time to first token that 95 of 100 requests stay under.

Reveal the answerHide the answer

Answer · in plain words

Reuse the notes for the opening that prompts share, read long prompts in small pieces, and store the notes in 8 bits. In the video that took one machine from 16 conversations at once to 50. The wait for the first word piece that 95 of 100 requests stay under fell from 1.6 to 0.5 seconds.

Picture it

A busy kitchen. It makes the sauce every dish shares once, instead of fresh for each order (prefix caching). It preps a huge banquet order in batches so regular diners still get served (chunked prefill). And it writes order tickets in shorthand, so twice as many fit on the rail (8-bit notes). Same kitchen, three times the diners.

Careful

The narration says the tuning ‘halved’ time to first token. Its own table, 1.6 to 0.5 seconds, shows a cut of more than two thirds: quote the table.

Worked answer, step by step

  1. The three settings: --enable-prefix-caching (reuse the notes for a shared opening), --enable-chunked-prefill (read long prompts in pieces) and --kv-cache-dtype fp8 (notes in 8 bits).
  2. Conversations at once: 16 before, 50 after. 50 ÷ 16 = 3.1 times as many, with no new hardware.
  3. 8-bit notes alone about double it: 16 x 2 = 32 (the video’s settings list: ‘roughly doubles concurrency’). The video does not say how the rest splits between the other two settings.
  4. Wait for the first token that 95 of 100 requests stay under (p95 TTFT, shared opening): 1.6 s before, 0.5 s after. 0.5 ÷ 1.6 = 0.31, about a third of the wait: over 3 times faster.
  5. Why the wait drops: with prefix caching the shared opening is not read again for every request, and chunked prefill stops one long prompt from freezing everyone else.
  6. Only then compare hardware. The video’s own words: ‘Then, and only then, argue about hardware.’
Go deeper: the engineer version

The kit's question

Three flags, measured

The kit's answer

Here is a five minute tuning pass on one model. Turning on prefix caching and chunked prefill, and moving the cache to eight bits, took the same hardware from sixteen concurrent sessions to fifty, and halved time to first token on a shared prompt. Those three flags are the first thing to try in any customer deployment. · Then, and only then, argue about hardware.

More detail: In vLLM, requests that share a prefix share its KV blocks and skip its prefill, which cuts TTFT and the memory each session needs. Chunked prefill splits long prefills into chunks interleaved with decode steps, trading a little TTFT on the long prompt for smoother streaming for everyone (Batching & Scheduling, Day 3). FP8 KV halves the bytes per token. The video reports only the combined result: FP8 KV alone accounts for about 2x (16 to 32), and the rest most plausibly comes from prefix sharing. 1.6 s to 0.5 s is a 69% cut (3.2 times faster), more than the ‘halved’ in the narration. On recent vLLM releases (the V1 engine), prefix caching and chunked prefill are already on by default, so check what the customer’s version turns on before promising those two as wins; the FP8 cache is still opt-in.

Words to know

Prefix caching
Reusing the notes for a prompt opening that many requests share. Example: the same instructions at the top of every chat.
Chunked prefill
Reading a long prompt in pieces, between other people’s writing steps, so it does not freeze their replies.
FP8 KV cache
Notes stored with 8 bits per number instead of 16: half the memory. Example: --kv-cache-dtype fp8.
p95 TTFT
The wait for the first token that 95 of 100 requests stay under. Example: 1.6 s before tuning, 0.5 s after.

How you will use this

Your benchmark’s second page: one model timed on three servers, the evidence behind today’s line “Ollama is llama.cpp with the flags hidden.” The rows are saved in results/01-latency.csv, and Day 10’s sizing memo (your written hardware recommendation) lists the latest timing rows from that file.