QuestionStep 1 · Serve one model three ways
You start llama.cpp with -c 8192 (room for 8,192 tokens, about 6,000 words) and --parallel 4 (4 seats, so 4 people at once). Why is that a trap?
Show answerHide answer
In plain words
The 8,192 tokens are shared out, not given to each person. Each of the 4 seats gets only 2,048 tokens (about 1,500 words), so a longer conversation does not fit.
Picture it
One large pizza ordered ‘for 4 people’. The box says large, but each person gets a quarter, and anyone hungrier than a quarter goes without.
With real numbersLlama 3.1 8B on llama.cpp with Day 2’s settings, lesson 02
- Room for notes:
-c 8192= 8,192 tokens, about 6,000 words. - Seats:
--parallel 4= 4 conversations at once. - Each seat: 8,192 ÷ 4 = 2,048 tokens, about 1,500 words.
- Day 1’s long test prompt, about 8,000 tokens, is almost 4 times one seat’s room (8,000 ÷ 2,048 = 3.9), so it does not fit.
- The fix: set
-cto seats x tokens each. For 4 seats of 8,192 tokens: 4 x 8,192 = 32,768, which takes 4.29 GB of notes for this model (Day 4 works this out).
Words to know
- Context length (
-c) - The most tokens the server keeps notes for. Example:
-c 8192= 8,192 tokens. - Slot (
--parallel) - One seat on the server: one conversation it serves at the same time. Example:
--parallel 4= 4 seats. - Token
- A chunk of text, about three quarters of a word. Example: 2,048 tokens is about 1,500 words.
- KV cache (notes)
- The model’s notes on each conversation so far, kept in memory. Example: 32,768 tokens of notes take 4.29 GB for Llama 3.1 8B.
Go deeper: the engineer version
The kit's question
Why is --parallel 4 -c 8192 a trap?
The kit's answer
llama.cpp splits the context across slots, so each user gets only 2,048 tokens.
More detail: llama-server sets aside one KV cache of -c tokens at startup and splits it evenly across the --parallel slots (the comment in serve_llamacpp.sh says so), so each slot’s context is c ÷ parallel. Size -c as slots x the longest context one user needs: 4 x 8,192 = 32,768, which is 4,096 MiB of f16 cache for Llama 3.1 8B at 128 KiB per token (lesson 05). llama-server’s startup log lists the context each slot gets, so confirm it there on your build.
How did you do?