Day 1 · Guardrails and first calls
Part 1 · Day 1 of 12
0 of 12 days done
About 1 h 25 4.9 GB model download · See Step 2 About $0.02 Mac and Fireworks
Today: Get the course code running on your Mac with a practice server: a free stand-in for a real AI model that gives made-up replies with realistic timing. Then put a spending cap on Fireworks, the cloud service that runs AI models for a fee. Finally, time how fast models answer: the wait for the first word, and the pace after it.
Example
Imagine asking a support assistant for help with a refund. Before answering, it reads the company rules and your question; then it writes its reply piece by piece. In the video’s example, you wait less than a second for the reply to start, but more than six seconds for it to finish. Today you will measure those two waits separately, so you can tell a customer which part needs speeding up.
By the end you’ll have
Your own timing table, saved in results/01-latency.csv: how long a reply takes to start and how quickly it appears after that, for a model on your Mac and for models on Fireworks. Step 3 explains a model in the script that is no longer available there.
Expected: In the lesson’s Mac test, a long message makes the reply take much longer to start than a short question. Once the reply starts, its words appear at almost the same pace. Your numbers will differ, but that split tells you whether to investigate the reading stage or the writing stage.
The 7 numbers you write down
Your Mac’s memory (GB)
How much model fits: plan on about 70% of it for the model and its notes.
Example: 36
Memory speed, bandwidth (GB/s)
How fast memory feeds the chip. It caps how fast any model can write.
Example: 273
Writing-speed limit for Llama 3.1 8B at 4 bits (tokens/s)
The fastest your Mac can write with the 4.9 GB model. Real runs reach 60 to 85% of it.
Example: 56 (273 ÷ 4.9)
Wait for the first word, short question (ms, p50)
How long a typical person waits before the reply starts, after a one-line question.
Example: 145
Wait for the first word, 8,000-token prompt (ms, p50)
The same wait when the prompt is about 6,000 words long: reading time grows with the prompt.
Example: 4,210
Gap between tokens while writing (ms)
The typing pace after the first word. It barely changes with prompt length.
Example: 22.4 (44.6 tokens per second)
Wait for the first word on each Fireworks model (ms, p50)
How hosted models compare on the same question, timed by your own stopwatch.
Example: The lesson has no example here: write the model, then ms, for each model that answered.
How you will use this: When someone says an AI feature is slow, your timing table helps you ask whether the wait is before or after the reply starts. Day 3 adds how many people one server can handle. Together those measurements become evidence for the hardware recommendation you will write on Day 10.
Before you start
An Apple-silicon Mac (M1 or later) with more than 5 GB of free disk space.
Step 2 downloads your first real model, Llama 3.1 8B (8 billion learned numbers, each stored in about 4 bits, half a byte): about 4.9 GB. The tools installed in steps 1 and 2 need room on top of that. The setup script is for Macs with Apple’s own chips; on an older Intel Mac, MLX (Apple’s tool for running models on its chips) will not work.
Python 3.10 or newer, and Homebrew.
The course code needs Python 3.10 or later. Step 2’s setup script installs its tools with Homebrew (the Mac’s usual installer for developer tools) and stops if Homebrew is missing.
From the
03-labsfolder:python3 --version && brew --versionIf
brewis not found, install Homebrew from brew.sh, then run the ‘Next steps’ lines it prints at the end; without thembrewis still not found. If Python says 3.9 or older, runbrew install python, open a new Terminal window and check again.A Fireworks account, for step 3.
Until you add a payment method, Fireworks gives you $1 of credit and at most 10 requests a minute. Step 3 sends about 17 requests in quick succession (22 if all three of its models answered), so step 2’s checklist has you buy a small prepaid credit (item 5). Step 2 also sets the spending cap first; step 3 costs about 1 to 2 cents.
Today’s steps
Section titled “Today’s steps”3 steps, then the wrap-up.
-
Get the code, run it offline
Download the course code and check that it works on your Mac. A practice server (a free stand-in for a real AI model) lets you test everything before any real model or money is involved.
Video: Inference 101 3:02
- Set up your Mac and the spend cap Install the programs that run models on your Mac, and read its two limits: memory size and memory speed. Then cap your Fireworks spending before anything can bill.
-
Measure TTFT and ITL
Build the stopwatch you will use all course: it times how long a model takes to start answering, and how fast it writes after that. The two have different causes, so you always measure them apart.
- Wrap-up and drill Put your numbers in one table, answer 5 questions out loud, explain 2 results and work 2 examples.
Your progress is saved on this device.