Skip to content

Day 1 wrap-up and drill

  1. Overview
  2. Step 1
  3. Step 2
  4. Step 3
  5. Wrap-up

About 10 minutes, free, and it works on your phone. Lock in today before you move on.

Done when

You have your own timing table: the wait for the first word and the gap between tokens. It covers Llama 3.1 8B on your Mac, with a short and a long prompt, and the Fireworks models that answered (gpt-oss-20b’s row is expected to fail). And make fw-check lists nothing still billing.

Fields you filled in on the step pages are already here. Nothing leaves this browser.

Step 2 · Set up your Mac and the spend cap

How much model fits: plan on about 70% of it for the model and its notes.

Example: 36

How fast memory feeds the chip. It caps how fast any model can write.

Example: 273

The fastest your Mac can write with the 4.9 GB model. Real runs reach 60 to 85% of it.

Example: 56 (273 ÷ 4.9)

Step 3 · Measure TTFT and ITL

How long a typical person waits before the reply starts, after a one-line question.

Example: 145

The same wait when the prompt is about 6,000 words long: reading time grows with the prompt.

Example: 4,210

The typing pace after the first word. It barely changes with prompt length.

Example: 22.4 (44.6 tokens per second)

How hosted models compare on the same question, timed by your own stopwatch.

Example: The lesson has no example here: write the model, then ms, for each model that answered.

Every question from today’s steps. Say your answer out loud, then show the answer and rate yourself. 5 cards

QuestionStep 2 · Set up your Mac and the spend cap

Your Mac has 36 GB of memory. Roughly how big a model, stored at 4 bits per number, can it run comfortably?

Show answerHide answer

In plain words

About a 40-billion-number model (40B). That fills the whole budget of about 25 GB, so there is little room left for the model’s notes on each conversation.

Picture it

Flatmates share one fridge. You can take about 70% of it before everyone else’s food gets squeezed out. Fill your whole share with one giant cake and there is no room for the leftovers each meal adds (the model’s notes).

With real numberslesson 00, and the 4.9 GB model from Day 1’s setup step

  • Budget: 0.7 x 36 GB = 25.2 GB, about 25 GB for the model and its notes.
  • Size per billion numbers at 4 bits: Llama 3.1 8B is 4.9 GB, so 4.9 ÷ 8 = about 0.61 GB per billion.
  • Biggest model: 25 ÷ 0.61 = about 41 billion numbers, so about a 40B model.
  • Room left for notes: almost none. Day 4 shows one long conversation’s notes can outgrow the model itself.

Words to know

Parameters (8B, 40B)
The model’s learned numbers, also called weights. Example: 8B = 8 billion.
4-bit
Each number stored in 4 bits, half a byte, plus a little extra. Example: an 8B model is 4.9 GB at 4-bit.
Dense model
A model that uses all of its numbers for every token. Example: Llama 3.1 8B.
Notes (KV cache)
The model’s memory of each conversation so far, kept in the same memory as the model. Example: Day 4 measures it.
Go deeper: the engineer version

The kit's question

Your Mac has 36 GB. Roughly the biggest 4-bit model you can run comfortably?

The kit's answer

≈ 0.7 × 36 ≈ 25 GB of weights, so about a 40B dense model at 4-bit, with little room left for context.

More detail: 0.7 x 36 = 25.2 GB for weights and KV cache together. Real 4-bit files (Q4_K_M) carry per-block scales and keep some tensors at higher precision, so Llama 3.1 8B at 4 bits is 4.9 GB (about 4.8 bits per weight, lesson 04), not the 4.0 GB that 0.5 bytes x 8B would give. 25.2 ÷ 0.61 = about 41B parameters, so about a 40B dense model, which leaves almost no KV-cache room. check_env.py prints memory in billions of bytes (a 36 GB Mac shows 38.7), so the 70% rule is a rough budget either way.

How did you do?

QuestionStep 2 · Set up your Mac and the spend cap

Why does Day 1’s check script (check_env.py) print memory speed (bandwidth) rather than how many GPU cores (the chip’s calculating units) you have?

Show answerHide answer

In plain words

Because memory speed caps how fast the model writes. For each token (word piece) the chip must fetch the whole model from memory once, and that fetching is slower than the math.

Picture it

A very fast cook needs the whole recipe book carried over on a conveyor belt before every dish. More cooks (more cores) would stand waiting; only a faster belt gets dishes out sooner.

With real numberslesson 00’s sample Mac and felab/hardware.py

  • Each token needs the whole 4.9 GB model fetched once.
  • At 273 GB per second: 273 ÷ 4.9 = about 56 tokens per second, at most.
  • Real runs land at 60 to 85% of that limit (felab/hardware.py); Day 4 checks it on your Mac.
  • A Mac whose memory moves 546 GB per second (the faster M4 Max) has twice the limit: 546 ÷ 4.9 = about 111.

Words to know

Decode
The writing phase: one token at a time, fetching the whole model for each.
Bandwidth
How many GB per second memory delivers to the chip. Example: 273 GB/s.
GPU cores
The many small calculating units inside a GPU. Example: they do the math, but on a Mac they mostly wait for memory while writing.
Compute-bound
Slowed by calculation speed rather than memory. Example: reading the prompt (prefill).
Go deeper: the engineer version

The kit's question

Why does the script print bandwidth rather than GPU cores?

The kit's answer

Decode reads every weight once per token, so bandwidth, not compute, sets the speed.

More detail: At batch size 1, each decode step streams every weight once and does only about two floating-point operations per weight, far less than the GPU could compute in that time. So the ceiling is bandwidth ÷ bytes read per token (felab/hardware.py), and measured decode typically reaches 60 to 85% of it. Prefill is the opposite: it reuses each weight across many prompt tokens, so compute sets its pace. Where a chip ships in two bandwidth versions (M3 Max, M4 Max), hardware.py lists the faster one.

How did you do?

QuestionStep 3 · Measure TTFT and ITL

A customer says their AI feature is ‘slow’. Which number do you ask for first, the wait for the first word or the pace after it, and why?

Show answerHide answer

In plain words

Both, because each points to a different cause. Slow to start means a long prompt or requests waiting in line; slow to type means memory speed, or a model too big for it.

Picture it

A patient says ‘I feel unwell’. The doctor takes both temperature and blood pressure, because each points to different illnesses. ‘Slow’ is the symptom; the two numbers are the tests.

With real numbersthe support assistant in the Inference 101 video

  • A customer asks the support bot about a refund. It reads the company rules and question before it can reply: 6,200 tokens (word pieces) take 0.78 s.
  • The bot then writes a 300-token answer, one piece every 22 ms: 300 x 22 ms = 6.6 s.
  • The customer waits about 7.4 s in all. Of that, 6.6 s is spent watching the reply appear: about 9 of every 10 seconds of this wait.
  • Even if reading became instant, the customer would save less than 0.78 s. To shorten this whole reply much more, its writing needs to speed up.

Words to know

Prefill
The reading phase: the whole prompt at once. Example: 6,200 tokens in 0.78 s.
Decode
The writing phase, one token at a time. Example: 300 tokens x 22 ms = 6.6 s.
Queueing
Requests waiting in line because the server is busy with others.
Bandwidth
How many GB per second memory delivers to the chip; it caps decode.
Go deeper: the engineer version

The kit's question

The customer says “it’s slow”. Which number do you ask for first, and why?

The kit's answer

Both. Slow to start points at prefill, queueing or a long prompt. Slow to type points at decode, bandwidth or an oversized model.

More detail: TTFT is roughly queue time plus prefill time, so a high TTFT sends you to prompt length, prefix caching, cold starts and queueing. ITL is one decode step, bound by memory bandwidth and the bytes read per token, so a high ITL sends you to model size, quantization, speculative decoding or batch size (bigger batches raise total throughput but slow each user a little). Ask for both, at p50 and p95, on the customer’s real prompt and output lengths.

How did you do?

QuestionStep 3 · Measure TTFT and ITL

Why report p95 (the time 95 out of 100 requests stay under) and not the average?

Show answerHide answer

In plain words

An average hides the few very slow requests, and those are the ones users remember and complain about.

Picture it

A bus that is on time 19 days out of 20 and an hour late on the 20th has an average delay of only 3 minutes. The rider remembers the day they missed a meeting.

With real numbersDay 1’s stopwatch script (felab/measure.py) and its sample run

  • The stopwatch script times 10 requests. p50 is the middle one; with only 10, p95 is the slowest one.
  • Lesson sample, short prompt: p50 is 145 ms, p95 is 190 ms.
  • So the worst request waited 31% longer than the typical one (190 ÷ 145 = 1.31).
  • Inference 101: a fine p50 with an ugly p95 almost always means waiting in line or servers waking up (cold starts), not the model.

Words to know

p50 (median)
The middle value: half the requests are faster, half slower. Example: 145 ms.
p95
The value 95 out of 100 requests stay under. Example: 190 ms.
Tail
The slowest few requests, the ones p95 describes.
Cold start
A slow request while a server that was idle or newly started loads the model.
Go deeper: the engineer version

The kit's question

Why p95 and not the average?

The kit's answer

Averages hide the tail, and the tail is what users complain about.

More detail: felab.percentile uses the nearest-rank method, so with 10 samples p95 is the maximum: deliberately pessimistic, because tails are what customers complain about. Agree the percentile and the traffic profile with the customer before you benchmark (Inference 101 recap).

How did you do?

QuestionStep 3 · Measure TTFT and ITL

With the long 8,000-token prompt, the gap between tokens grows a little (22.4 to 24.1 ms in the lesson’s sample). Why?

Show answerHide answer

In plain words

For each new token the model also rereads its notes on everything it has read so far. A longer prompt means more notes to fetch, so each step takes a little longer.

Picture it

Before every dish, the cook glances at the notes on this table’s order. For a long order the notes run to several pages, so each glance takes a moment longer. The recipe book, fetched every time too, is still the big load.

With real numbersthe lesson’s sample run, and Llama 3.1 8B’s notes size from Day 4

  • The model: 4.9 GB, fetched once for every token written.
  • The notes (the KV cache) for Llama 3.1 8B: 128 KB per token.
  • 12-token prompt: 12 x 128 KB = about 1.6 MB (million bytes) of notes, next to nothing.
  • 8,000-token prompt: 8,000 x 128 KB = about 1.05 GB of notes, about a fifth more to fetch on top of the 4.9 GB model (1.05 ÷ 4.9 = 0.21).
  • Counting bytes alone, each step could take up to a fifth longer. In the sample the gap grew 8% (22.4 to 24.1 ms): a little, because the model is still most of what is fetched.

Words to know

KV cache
The model’s notes on the conversation so far (keys and values), so it does not reread everything for each new word.
Context
Everything the model has read and written so far in one conversation. Example: an 8,000-token prompt.
KB
1,024 bytes in this course. Example: 128 KB = 131,072 bytes.
Decode step
One pass through the model that writes one token; it fetches the weights and the notes.
Go deeper: the engineer version

The kit's question

Why does ITL creep up a little with the long prompt?

The kit's answer

Each decode step also reads the KV cache, and a longer context means a bigger cache.

More detail: Each decode step reads all the weights plus the K and V entries of every earlier token (attention). At 128 KiB per token for Llama 3.1 8B (2 x 32 layers x 8 KV heads x 128 x 2 bytes, lesson 05), 8,000 tokens add about 1.05 GB per step against 4.9 GB of weights, so ITL rises. Bytes alone predict about 21% ((4.9 + 1.05) ÷ 4.9 = 1.21); the sample’s 8% is lower, and the kit does not explain the gap. Its sample figures are illustrative, the padded prompt may hold fewer than 8,000 real tokens (the script assumes 4 characters per token), or Ollama may have cut it to its context setting (see Stuck?). With many users at long context those reads add up, which is why Day 4 caps context and quantizes the cache.

How did you do?

Under 20 seconds each. Record yourself once and listen back.

Step 2 · Set up your Mac and the spend cap

Question How big a model can this machine run, and how fast?

One clear answer

On unified memory, model size competes with everything else on the box. I size to about 70% of RAM and check the bandwidth before I promise a tokens-per-second number.

What this means
  • “On unified memory”: On a Mac, the main processor (CPU) and the graphics processor (GPU) share one pool of memory; the GPU has none of its own.
  • “model size competes with everything else on the box”: The model sits in the same memory as your browser and other apps (‘the box’ is the machine), so it cannot have all of it.
  • “I size to about 70% of RAM”: RAM is the computer’s working memory. I plan for the model and its notes to use at most 70% of it: about 25 GB on a 36 GB Mac.
  • “and check the bandwidth”: I look up how fast memory delivers data, for example 273 GB per second.
  • “before I promise a tokens-per-second number”: Only then do I quote a writing speed. The ceiling is bandwidth ÷ model size: 273 ÷ 4.9 = about 56 tokens per second.

Step 3 · Measure TTFT and ITL

Question The customer says ‘it’s slow’. What do you measure first?

One clear answer

I measure TTFT and ITL separately, at p50 and p95, on the customer’s own prompt shape. One latency number hides two different bottlenecks.

What this means
  • “I measure TTFT and ITL separately”: I time how long until the first token appears (time to first token, 145 ms in the sample), and separately the gap between the tokens after it (inter-token latency, 22.4 ms).
  • “at p50 and p95”: For each, I report the typical request (p50: half are faster) and the slow tail (p95: 95 of 100 are faster).
  • “on the customer’s own prompt shape”: I test with prompts and answers as long as theirs: 12 tokens gave 145 ms to the first word, 8,000 tokens gave 4,210 ms.
  • “One latency number hides two different bottlenecks”: Latency is how long someone waits; a bottleneck is the slowest part, the one that holds everything up. A single ‘response time’ mixes reading and writing, which have different bottlenecks and different fixes.

2 worked examples from today’s videos. Work each on paper, then reveal.

From the video: Inference 101 3:02

Whiteboard 1 of 2

A support assistant sends the model 6,000 tokens of standing instructions (the system prompt) plus a 200-token customer question, and gets a 300-token answer. The model reads about 8,000 tokens per second and writes one token every 22 ms. How long does the customer wait, and where does the time go?

Given, in plain words

Tokens are word pieces, about three quarters of a word each: 6,200 tokens in is about 4,650 words, and 300 out is about 225 words. Reading speed (prefill): about 8,000 tokens per second. Writing pace (inter-token latency): 22 ms (thousandths of a second) per token. One request at a time, no waiting in line.

Reveal the answerHide the answer

Answer · in plain words

Picture a customer waiting for a support answer. The bot first reads its instructions and the question; the customer sees nothing for about 0.78 seconds. Then the answer appears piece by piece for 6.6 more seconds. The whole wait is about 7.4 seconds, and most of it happens after the reply has started.

Picture it

A kitchen reads a long order in one glance but plates the dishes one at a time. Here the glance takes under a second and the plating over six. The long order looks like the problem; the one-by-one plating is.

Careful

With many users at once, requests can also wait in line before reading starts, which adds to the wait for the first word.

Worked answer, step by step

  1. Tokens in: 6,000 (instructions) + 200 (question) = 6,200.
  2. Wait for the first word (TTFT): 6,200 ÷ 8,000 tokens per second = 0.775 s, about 0.78 s.
  3. Writing time: 300 tokens x 22 ms = 6,600 ms = 6.6 s.
  4. Total: 0.775 + 6.6 = 7.375 s, about 7.4 s.
  5. Share that is writing: 6.6 ÷ 7.375 = 0.89, so 89%.
  6. The video’s fix is for the start: the 6,000-token instructions never change, so prompt caching reuses them and the first word comes after about 120 ms instead of 780. The writing (6.6 s) is still most of the wait; the lever for it is faster memory or speculative decoding (a small model drafts, the big one checks).
Go deeper: the engineer version

The kit's question

A support assistant sends a six thousand token system prompt, a two hundred token question, and gets a three hundred token answer.

The kit's answer

Total 7.4 s — and 89% of it is decode

More detail: TTFT here is prefill alone (no queue): 6,200 ÷ 8,000 = 0.775 s. Decode: 300 x 22 ms = 6.6 s. Total 7.375 s, of which 6.6 ÷ 7.375 = 89.5% is decode. The video files this case under ‘long prompt, short answer: prefill dominates’, but its own numbers put 89% of the time in decode: prompt caching fixes the first-word wait (p95 TTFT 780 ms to 120 ms) and the input bill, not the total. The narration also lists batching as a decode fix; batching raises total throughput but can slow each user (field guide: ‘Bigger batches raise throughput but can slow each user’), so it does not shorten this customer’s wait. The screen shows input cost per 1,000 calls falling from $0.93 to $0.09. $0.93 for 6.2 million tokens implies $0.15 per million, the lab book’s gpt-oss 120B input price. At its cached price ($0.015 per million), $0.09 matches either the 6 million cached tokens alone ($0.09) or all 6.2 million at the cached price ($0.093); charging each call’s 200 fresh tokens at the full price gives about $0.12.

Words to know

System prompt
Standing instructions sent at the start of every request. Example: the support assistant’s 6,000 tokens.
TTFT
Time to first token: how long until the reply starts. Example: 0.78 s here.
Decode
The writing phase, one token at a time. Example: 300 x 22 ms = 6.6 s.
Prompt caching
Reusing the server’s work on a prompt opening that repeats in every request. Example: 780 ms to about 120 ms here.

From the video: Inference 101 3:02

Whiteboard 2 of 2

Same model, same hardware, a different job: code completion, where a code editor suggests the next few lines. 200 tokens go in and 30 come out. Reading runs at the same 8,000 tokens per second, writing at 22 ms per token. Which phase takes the time, and what would you speed up?

Given, in plain words

200 tokens in (about 150 words of code around the cursor), 30 tokens out. Reading (prefill): about 8,000 tokens per second. Writing: 22 ms per token.

Reveal the answerHide the answer

Answer · in plain words

Writing still takes almost all the time: 25 ms to read, 660 ms to write, 96% of the 685 ms total. So you speed up the writing, and you always ask for the request’s shape before you promise a speed.

Picture it

A short order for a few dishes. The kitchen reads it instantly, so any delay is at the plating station. Reading orders faster would change nothing.

The video's lesson

Ask for the token shape (how many tokens in and out) before you promise a latency number.

Worked answer, step by step

  1. Reading (prefill): 200 ÷ 8,000 = 0.025 s = 25 ms.
  2. Writing (decode): 30 x 22 ms = 660 ms = 0.66 s.
  3. Total: 25 + 660 = 685 ms, about 0.7 s.
  4. Share that is writing: 660 ÷ 685 = 96%.
  5. Compare the support assistant: its reading took 780 ms, 31 times this one’s 25 ms (6,200 ÷ 200 = 31).
  6. So the fix is on the writing side: speculative decoding (a small model drafts, the big one checks; Day 9) or faster memory. Prompt caching, the video’s fix for the support assistant’s slow start, would save almost nothing here.
Go deeper: the engineer version

The kit's question

Now change the shape. A code completion request: two hundred tokens in, thirty tokens out.

The kit's answer

Prefill is twenty five milliseconds, decode is under a second. Same model, same GPU, a completely different performance problem, and a completely different fix.

More detail: Prefill 200 ÷ 8,000 = 25 ms; decode 30 x 22 = 660 ms; decode is 96% of 685 ms. The screen lists ‘speculation, faster memory’ as what to optimise, and its takeaway is ‘Ask for the token shape before you promise a latency number’. The ‘completely different fix’ is relative to the video’s fix for the support assistant, prompt caching; both workloads are mostly decode.

Words to know

Code completion
An editor feature that suggests the next few lines of code. Example: 200 tokens in, 30 out.
Traffic profile (token shape)
How many tokens go in and come out per request, and how many requests arrive at once.
Speculative decoding
A small, fast model drafts a few tokens and the big model checks them in one pass. Example: Day 9.
Prefill
The reading phase: the whole prompt at once. Example: 25 ms for 200 tokens.

How you will use this

When someone says an AI feature is slow, your timing table helps you ask whether the wait is before or after the reply starts. Day 3 adds how many people one server can handle. Together those measurements become evidence for the hardware recommendation you will write on Day 10.