Skip to content

Measure TTFT and ITL

course 3 of 22

lesson 01 · lessons/01-latency-harness

30 min About $0.02 Mac, then Fireworks

What you will do and why

Build the stopwatch you will use all course: it times how long a model takes to start answering, and how fast it writes after that. The two have different causes, so you always measure them apart.

Why it matters: In the lesson’s example run, an 8,000-token prompt instead of a 12-token one made the wait for the first word 29 times longer (145 to 4,210 ms). The gap between tokens stayed about the same (22.4 to 24.1 ms).

You are done when: results/01-latency.csv has your Ollama rows (short and long prompt) and a row for each Fireworks model that answered (two, with gpt-oss-20b’s expected failure), and make fw-check lists no deployments.

Rewatch Inference 101, 0:42 to 1:51ShowHide

Every reply has two phases. First the model reads your whole prompt, and you wait for the first word. Then it writes one token (word piece) at a time, and you watch the text appear. Each phase has its own number and its own causes, so you report both.

Picture it

In a restaurant, the wait for your first dish is the kitchen reading your order and starting to cook: a longer order, a longer wait. The pace of the next dishes is how fast the kitchen plates them, one by one. A complaint about ‘slow service’ does not say which of the two to fix.

With real numbersthe lesson’s sample run: Llama 3.1 8B through Ollama on a Mac

  • Short prompt (12 tokens): the first word after 145 ms (thousandths of a second), then a new token every 22.4 ms.
  • Long prompt (8,000 tokens, about 6,000 words): the first word after 4,210 ms. 4,210 ÷ 145 = 29 times longer.
  • Gap between tokens with the long prompt: 24.1 ms, only 8% more (24.1 ÷ 22.4 = 1.08).
  • Writing speed for one person: a second is 1,000 ms, so 1,000 ms ÷ 22.4 ms = 44.6 tokens per second, about 33 words.
  • So a long prompt delays the start; it barely slows the typing.

Words to know

TTFT (time to first token)
How long until the first token of the reply appears. Example: 145 ms in the sample.
ITL (inter-token latency)
The gap between one token of the reply and the next. Example: 22.4 ms, 44.6 tokens per second.
Prefill
The reading phase: the model reads the whole prompt at once. It sets most of TTFT.
Decode
The writing phase: one token at a time. It sets ITL.

From the lesson

lessons/01-latency-harness/README.md

Every LLM request has two phases:

phase what the GPU does what the user feels metric
prefill reads the whole prompt at once (compute-heavy) “is it thinking?” TTFT, time to first token
decode writes one token at a time (memory-bandwidth-heavy) “how fast does it type?” ITL, inter-token latency

Because the two phases have different bottlenecks, you always report them separately and at p50 and p95. bench_ttft.py is the tool for that. You will point this same script at Ollama, llama.cpp, MLX, Fireworks and the Spark.

  • felab/measure.py → stream_once() timestamps every streamed token. TTFT is the first stamp; ITL is the gaps between stamps.
  • bench_ttft.py → build_prompt() pads the prompt to put the context first and the question last, which is the shape of a real RAG prompt.
  • A warm-up call runs first and is never counted. Cold starts are a separate conversation.
Terminal window
cd lessons/01-latency-harness
# 01a — free
python bench_ttft.py --target mock
python bench_ttft.py --target mock --prompt-tokens 8000
python bench_ttft.py --target ollama
python bench_ttft.py --target ollama --prompt-tokens 8000
# 01b: Fireworks (about 1 to 2 cents). Needs your key in .env and the spend cap from step 2
bash compare_models.sh

What each command does

  1. python bench_ttft.py --target mock

    Times 10 requests to the practice server, one after another. A first warm-up request is never counted. Look for a first word after roughly 20 to 30 ms and a gap of about 18 ms: the practice server is set that way.

  2. python bench_ttft.py --target mock --prompt-tokens 8000

    The same, with filler text making the prompt about 8,000 tokens long. Look for a first word after about 3,200 ms (the practice server reads 2,500 tokens a second). The gap stays about 18 ms.

  3. python bench_ttft.py --target ollama

    The same stopwatch on a real model: Llama 3.1 8B through Ollama, still running from step 2. The uncounted warm-up is when Ollama loads the model, which can take seconds. The lesson’s example run: 145 ms to the first word, 22.4 ms between tokens.

  4. python bench_ttft.py --target ollama --prompt-tokens 8000

    The real model with the long prompt. The lesson’s example run: 4,210 ms to the first word (29 times longer) and 24.1 ms between tokens (about the same). Your Mac’s numbers will differ; the pattern is the lesson. If Ollama’s log says ‘truncating input prompt’, see Stuck? below.

  5. bash compare_models.sh

    Paid, about 1 to 2 cents: times three Fireworks models (5 requests each), then the first one with the long prompt. Expect one failed row, and that is fine. The second model in the list, gpt-oss-20b, is no longer offered per token: on 27 September 2026 its model page said “Serverless: Not supported”. So the script prints failed — check the id in the model library for it and carries on. The other rows still count: gpt-oss-120b (short and long prompt) and, if its id is still current, the Llama 3.1 8B row. Optional: for a third row, open your copy of compare_models.sh and, in the MODELS list at the top, replace the gpt-oss-20b line with another model the Fireworks model library (app.fireworks.ai/models) marks as serverless. Before you run it, add your key: create it at app.fireworks.ai, on the API keys page. Run open -e ../../.env to open your settings file in TextEdit (../.. is the 03-labs folder; Finder hides files whose names start with a dot). Replace the placeholder after FIREWORKS_API_KEY= with your key and save. Afterwards, go back with cd ../.. and run make fw-check.

How to read it

Each command prints its timed runs, then one summary row, also saved to results/01-latency.csv. In the row, ttft_p50_ms and ttft_p95_ms are the typical and the slowest waits for the first word, itl_ms is the gap between tokens and tok_per_s is 1,000 ÷ that gap. Compare the short and long rows: the first-word columns jump (145 to 4,210 ms in the sample) while the gap barely moves.

target model prompt_tok ttft_p50_ms ttft_p95_ms itl_ms tok_per_s
------ -------- ---------- ----------- ----------- ------ ---------
ollama llama3.1:8b 12 145.0 190.0 22.4 44.6
ollama llama3.1:8b 8000 4,210.0 4,480.0 24.1 41.5 ← TTFT ×29, ITL about the same

Exact numbers depend on your chip. The shape is the lesson.

3 questions. Say your answer out loud, then tap to check it.

A customer says their AI feature is ‘slow’. Which number do you ask for first, the wait for the first word or the pace after it, and why?Show answerHide

In plain words

Both, because each points to a different cause. Slow to start means a long prompt or requests waiting in line; slow to type means memory speed, or a model too big for it.

Picture it

A patient says ‘I feel unwell’. The doctor takes both temperature and blood pressure, because each points to different illnesses. ‘Slow’ is the symptom; the two numbers are the tests.

With real numbersthe support assistant in the Inference 101 video

  • A customer asks the support bot about a refund. It reads the company rules and question before it can reply: 6,200 tokens (word pieces) take 0.78 s.
  • The bot then writes a 300-token answer, one piece every 22 ms: 300 x 22 ms = 6.6 s.
  • The customer waits about 7.4 s in all. Of that, 6.6 s is spent watching the reply appear: about 9 of every 10 seconds of this wait.
  • Even if reading became instant, the customer would save less than 0.78 s. To shorten this whole reply much more, its writing needs to speed up.

Words to know

Prefill
The reading phase: the whole prompt at once. Example: 6,200 tokens in 0.78 s.
Decode
The writing phase, one token at a time. Example: 300 tokens x 22 ms = 6.6 s.
Queueing
Requests waiting in line because the server is busy with others.
Bandwidth
How many GB per second memory delivers to the chip; it caps decode.
Go deeper: the engineer version

The kit's question

The customer says “it’s slow”. Which number do you ask for first, and why?

The kit's answer

Both. Slow to start points at prefill, queueing or a long prompt. Slow to type points at decode, bandwidth or an oversized model.

More detail: TTFT is roughly queue time plus prefill time, so a high TTFT sends you to prompt length, prefix caching, cold starts and queueing. ITL is one decode step, bound by memory bandwidth and the bytes read per token, so a high ITL sends you to model size, quantization, speculative decoding or batch size (bigger batches raise total throughput but slow each user a little). Ask for both, at p50 and p95, on the customer’s real prompt and output lengths.

Why report p95 (the time 95 out of 100 requests stay under) and not the average?Show answerHide

In plain words

An average hides the few very slow requests, and those are the ones users remember and complain about.

Picture it

A bus that is on time 19 days out of 20 and an hour late on the 20th has an average delay of only 3 minutes. The rider remembers the day they missed a meeting.

With real numbersDay 1’s stopwatch script (felab/measure.py) and its sample run

  • The stopwatch script times 10 requests. p50 is the middle one; with only 10, p95 is the slowest one.
  • Lesson sample, short prompt: p50 is 145 ms, p95 is 190 ms.
  • So the worst request waited 31% longer than the typical one (190 ÷ 145 = 1.31).
  • Inference 101: a fine p50 with an ugly p95 almost always means waiting in line or servers waking up (cold starts), not the model.

Words to know

p50 (median)
The middle value: half the requests are faster, half slower. Example: 145 ms.
p95
The value 95 out of 100 requests stay under. Example: 190 ms.
Tail
The slowest few requests, the ones p95 describes.
Cold start
A slow request while a server that was idle or newly started loads the model.
Go deeper: the engineer version

The kit's question

Why p95 and not the average?

The kit's answer

Averages hide the tail, and the tail is what users complain about.

More detail: felab.percentile uses the nearest-rank method, so with 10 samples p95 is the maximum: deliberately pessimistic, because tails are what customers complain about. Agree the percentile and the traffic profile with the customer before you benchmark (Inference 101 recap).

With the long 8,000-token prompt, the gap between tokens grows a little (22.4 to 24.1 ms in the lesson’s sample). Why?Show answerHide

In plain words

For each new token the model also rereads its notes on everything it has read so far. A longer prompt means more notes to fetch, so each step takes a little longer.

Picture it

Before every dish, the cook glances at the notes on this table’s order. For a long order the notes run to several pages, so each glance takes a moment longer. The recipe book, fetched every time too, is still the big load.

With real numbersthe lesson’s sample run, and Llama 3.1 8B’s notes size from Day 4

  • The model: 4.9 GB, fetched once for every token written.
  • The notes (the KV cache) for Llama 3.1 8B: 128 KB per token.
  • 12-token prompt: 12 x 128 KB = about 1.6 MB (million bytes) of notes, next to nothing.
  • 8,000-token prompt: 8,000 x 128 KB = about 1.05 GB of notes, about a fifth more to fetch on top of the 4.9 GB model (1.05 ÷ 4.9 = 0.21).
  • Counting bytes alone, each step could take up to a fifth longer. In the sample the gap grew 8% (22.4 to 24.1 ms): a little, because the model is still most of what is fetched.

Words to know

KV cache
The model’s notes on the conversation so far (keys and values), so it does not reread everything for each new word.
Context
Everything the model has read and written so far in one conversation. Example: an 8,000-token prompt.
KB
1,024 bytes in this course. Example: 128 KB = 131,072 bytes.
Decode step
One pass through the model that writes one token; it fetches the weights and the notes.
Go deeper: the engineer version

The kit's question

Why does ITL creep up a little with the long prompt?

The kit's answer

Each decode step also reads the KV cache, and a longer context means a bigger cache.

More detail: Each decode step reads all the weights plus the K and V entries of every earlier token (attention). At 128 KiB per token for Llama 3.1 8B (2 x 32 layers x 8 KV heads x 128 x 2 bytes, lesson 05), 8,000 tokens add about 1.05 GB per step against 4.9 GB of weights, so ITL rises. Bytes alone predict about 21% ((4.9 + 1.05) ÷ 4.9 = 1.21); the sample’s 8% is lower, and the kit does not explain the gap. Its sample figures are illustrative, the padded prompt may hold fewer than 8,000 real tokens (the script assumes 4 characters per token), or Ollama may have cut it to its context setting (see Stuck?). With many users at long context those reads add up, which is why Day 4 caps context and quantizes the cache.

Question The customer says ‘it’s slow’. What do you measure first?

One clear answer

I measure TTFT and ITL separately, at p50 and p95, on the customer’s own prompt shape. One latency number hides two different bottlenecks.

What this means

  • “I measure TTFT and ITL separately”: I time how long until the first token appears (time to first token, 145 ms in the sample), and separately the gap between the tokens after it (inter-token latency, 22.4 ms).
  • “at p50 and p95”: For each, I report the typical request (p50: half are faster) and the slow tail (p95: 95 of 100 are faster).
  • “on the customer’s own prompt shape”: I test with prompts and answers as long as theirs: 12 tokens gave 145 ms to the first word, 8,000 tokens gave 4,210 ms.
  • “One latency number hides two different bottlenecks”: Latency is how long someone waits; a bottleneck is the slowest part, the one that holds everything up. A single ‘response time’ mixes reading and writing, which have different bottlenecks and different fixes.

Your numbersSaved on this device and collected in the Day 1 wrap-up.

Hint: Your ollama row with prompt_tok 12, column ttft_p50_ms.

Hint: Your ollama row with prompt_tok 8000, column ttft_p50_ms.

Hint: itl_ms on your short-prompt ollama row; tok_per_s next to it is 1,000 ÷ this.

Hint: The fireworks rows with prompt_tok 12 from compare_models.sh, one value per model. Expect two: gpt-oss-20b fails and saves no row.

Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.

results/01-latency.csv has your Ollama rows (short and long prompt) and a row for each Fireworks model that answered (two, with gpt-oss-20b’s expected failure), and make fw-check lists no deployments.

The script prints ‘failed — check the id in the model library’ for gpt-oss-20b.
Expected, nothing is wrong with your setup. On 27 September 2026 gpt-oss-20b’s model page said “Serverless: Not supported”: Fireworks no longer runs it per token, so the call fails at once. The script carries on, and your other Fireworks rows still count. Optional: replace the gpt-oss-20b line in the MODELS list at the top of your copy of compare_models.sh with another model the model library (app.fireworks.ai/models) marks as serverless, and run it again.
Every model says ‘failed — check the id in the model library’, and the error above mentions 401 or an invalid API key.
.env still holds the placeholder key fw_xxxxxxxxxxxxxxxx. From the lesson folder run open -e ../../.env, replace the placeholder after FIREWORKS_API_KEY= with your key from app.fireworks.ai (API keys page), save, and run bash compare_models.sh again.
‘FIREWORKS_API_KEY is not set’.
The FIREWORKS_API_KEY= line is missing from .env, or has nothing after the =. From the lesson folder run open -e ../../.env, add the line with your key after the = (create the key at app.fireworks.ai, API keys page), and save. Never commit or share that file.
The script prints ‘failed — check the id in the model library’ for another model too.
Model names change. Copy a current id from the Fireworks model library (app.fireworks.ai/models) into the list at the top of compare_models.sh, or pass it to bench_ttft.py with --model. If all three fail and the error mentions 401, it is the key, not the ids: see the first row.
Fireworks errors with code 429 (too many requests).
Until a payment method is added, the account allows 10 requests a minute. compare_models.sh sends about 17: 6 per model that answers (5 timed plus a warm-up), 1 for gpt-oss-20b’s failed call and 4 for the long prompt (6 + 1 + 6 + 4). Running it again hits the same limit. Add a small prepaid credit (spend-cap checklist item 5), or time one model at a time, a minute apart, for example python bench_ttft.py --target fireworks --model accounts/fireworks/models/gpt-oss-120b --runs 5 (6 requests, counting the warm-up).
During the long-prompt Ollama run, Ollama’s log (in the window where it runs) says ‘truncating input prompt’.
Ollama cut the prompt to its context setting (the most tokens it accepts), so the row timed a shorter prompt. Stop it with pkill ollama, start it with room for the long prompt, OLLAMA_CONTEXT_LENGTH=16384 ollama serve &, and run the long-prompt line again.
‘Connection refused’.
That server is not running. From the 03-labs folder, make check shows which are up. Start the practice server with make mock &, and Ollama with ollama serve &.
make fw-check says ‘No rule to make target’.
You are still in the lesson folder. Run cd ../.. to get back to 03-labs, then run it again.
bench_ttft.py Python · 76 lines
"""
Lesson 01 — the latency harness you will reuse in every later lesson.
It sends the same prompt N times, one at a time, and reports:
TTFT p50 / p95 how long until the first token (prefill + queue)
ITL median gap between tokens (one decode step)
tok/s what ONE user sees once text is flowing (= 1000 / ITL)
Try it:
python bench_ttft.py --target mock
python bench_ttft.py --target mock --prompt-tokens 8000 # TTFT jumps, ITL doesn't
python bench_ttft.py --target mock --prompt-tokens 8000 --warm-cache # sneak peek of lesson 06
python bench_ttft.py --target ollama
python bench_ttft.py --target fireworks --runs 5 # ~ $0.01
The lesson hides in the second line: a longer prompt makes TTFT worse (more to
prefill) but leaves ITL flat (each new token costs about the same). Two different
bottlenecks, so two different numbers.
"""
import argparse
from felab import add_target_args, banner, client, percentile, record, resolve, stream_once, table, user
QUESTION = "Explain KV caching to a CFO in about 120 words."
FILLER = ("Context paragraph about a customer's support history, product catalogue and policies. " * 400)
def build_prompt(prompt_tokens: int) -> str:
"""Pad the question with filler until it is roughly `prompt_tokens` long (~4 chars/token).
Padding goes FIRST and the question LAST, which is how real RAG prompts look."""
if prompt_tokens <= 50:
return QUESTION
pad = FILLER[: prompt_tokens * 4]
return f"{pad}\n\nUsing the context above if useful: {QUESTION}"
def main() -> None:
ap = add_target_args(argparse.ArgumentParser(description=__doc__,
formatter_class=argparse.RawDescriptionHelpFormatter))
ap.add_argument("--runs", type=int, default=10, help="sequential requests (default 10)")
ap.add_argument("--prompt-tokens", type=int, default=0, help="pad prompt to ~N tokens (e.g. 8000)")
ap.add_argument("--max-tokens", type=int, default=192)
ap.add_argument("--warm-cache", action="store_true",
help="reuse the identical prompt so the server's prefix cache can hit (see lesson 06)")
args = ap.parse_args()
t = resolve(args)
banner(t)
cli = client(t)
prompt = build_prompt(args.prompt_tokens)
# 1 warm-up request: loads the model / opens connections. Never count the first call.
stream_once(cli, t.model, user("Say hi."), max_tokens=8)
samples = []
for i in range(args.runs):
# Servers cache prompt prefixes (lesson 06). A unique tag at the very START makes
# every run a cold prefill, so TTFT measures real prefill — unless --warm-cache.
p = prompt if args.warm_cache else f"[run {i}-{id(samples)}] {prompt}"
s = stream_once(cli, t.model, user(p), max_tokens=args.max_tokens)
samples.append(s)
print(f" run {i + 1:2d}: TTFT {s.ttft_ms:7.0f} ms ITL {s.itl_median:5.1f} ms {s.tokens} tokens")
ttfts = [s.ttft_ms for s in samples]
itl = percentile([x for s in samples for x in s.itl_ms], 50)
row = {
"target": t.name, "model": t.model.split("/")[-1], "prompt_tok": args.prompt_tokens or 12,
"ttft_p50_ms": percentile(ttfts, 50), "ttft_p95_ms": percentile(ttfts, 95),
"itl_ms": itl, "tok_per_s": 1000 / itl if itl else 0.0,
}
print("\n" + table([row]))
print(f"\nsaved → {record('01-latency', row)}")
if __name__ == "__main__":
main()
compare_models.sh Bash · 27 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 01b — the same harness against three Fireworks serverless models.
# Cost: 3 models × 6 requests × ~200 tokens ≈ one cent.
#
# Model ids change as the library changes. Look up current ones at
# https://app.fireworks.ai/models and edit the list below (full "accounts/..." ids).
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
cd "$(dirname "$0")"
MODELS=(
accounts/fireworks/models/gpt-oss-120b # large open MoE
accounts/fireworks/models/gpt-oss-20b # small MoE
accounts/fireworks/models/llama-v3p1-8b-instruct # small dense — may have moved; swap in any 7–8B
)
for m in "${MODELS[@]}"; do
python bench_ttft.py --target fireworks --model "$m" --runs 5 || echo " ✗ $m failed — check the id in the model library"
done
echo
echo "Now the long-prompt run on one model — watch TTFT move and ITL stay put:"
python bench_ttft.py --target fireworks --model "${MODELS[0]}" --runs 3 --prompt-tokens 8000
echo
echo "All rows are in results/01-latency.csv"