Get the code, run it offline
course 1 of 22
15 min Free Mac
What you will do and why
Download the course code and check that it works on your Mac. A practice server (a free stand-in for a real AI model) lets you test everything before any real model or money is involved.
Why it matters: A quick check of everything (the smoke test) runs the code of 16 of the kit’s 19 lessons in about 2 minutes. It is offline and free: the practice server stands in for a real model.
You are done when: make smoke finishes, after about 2 minutes, with the line ‘smoke test passed’.
Inference 101 · 3:02
Download mp4 (11.1 MB)Watch while pip installs.
Chapters
In this video Why every AI reply has two phases, reading the prompt and writing the answer, and the numbers customers feel, starting with the wait for the first word and the pace after it.
The narrator’s “next video” is the file order. Your next step is below.
3 key points
The model reads your whole prompt in one pass, then writes its answer one token (word piece) at a time.
Reading is called prefill; writing is called decode. In the video’s support assistant, reading 6,200 tokens takes 0.78 seconds and writing 300 tokens takes 6.6 seconds.
The wait for the first word comes from reading and from waiting in line. The pace after that comes from writing.
These are TTFT (time to first token) and ITL (inter-token latency, the gap between tokens). At 22 ms (thousandths of a second) per token, a reply types about 45 tokens a second: 1,000 ms ÷ 22 ms = 45.
Before you measure anything, agree with the customer on the request’s shape and on which percentile counts (for example p95: the time 95 of 100 requests stay under).
The shape is how many tokens go in and come out. A support assistant’s 6,200-token prompt takes 0.78 s to read, and reusing its unchanging part (prompt caching) cuts that to about 0.12 s. A code completion (an editor suggesting the next lines of code) reads its 200 tokens in 25 ms, so the same trick saves almost nothing there.
In plain words
Section titled “In plain words”The download is the course’s code: a small shared toolkit and one folder of scripts per lesson. It includes a practice server, a stand-in for a real AI model that runs on your Mac. It replies with made-up words but copies a real model’s timing: a wait before the first word, then a steady pace. So you can check everything before you download a model or pay for one.
Picture it
A flight simulator. The controls, the checklists and the timings work like the real plane, but nothing leaves the ground and a crash costs nothing. Where it differs: its timings are set by hand, so it teaches the pattern, not your Mac’s real speed.
With real numbersthe practice server’s settings, felab/mock_server.py, and scripts/smoke_test.sh
- It runs on your own Mac at the address
localhost:9000(localhost means this computer; 9000 is the numbered door it listens on). Nothing leaves the machine and nothing is billed. - It writes one token (word piece) every 18 ms (thousandths of a second). A second is 1,000 ms, so 1,000 ÷ 18 = about 56 tokens per second.
- It reads prompts at 2,500 tokens per second, plus a fixed 15 ms per request. An 8,000-token prompt: 8,000 ÷ 2,500 = 3.2 s = 3,200 ms, plus 15 ms, so about 3,200 ms before the first word.
- It has 8 seats (slots): 8 requests are served at once, and a 9th waits in line. Day 3 uses that.
- The smoke test runs 16 of the kit’s 19 lessons offline with it, about 2 minutes in all.
Words to know
- Practice server (mock)
- The kit’s stand-in server with realistic behaviour and made-up numbers. Example: it writes a token every 18 ms.
- Virtual environment (venv)
- A private set of Python add-ons for one folder, so nothing else on your Mac changes. Example: the
.venvfolder inside03-labs. - Token
- A chunk of text, about three quarters of a word. 1,000 tokens is roughly 750 words.
- Smoke test
- A quick run of everything to prove the setup works before real use. Example:
make smoke, about 2 minutes.
Everything below runs in Terminal, the Mac app for typing commands (Applications, then Utilities). Copy each block, paste it and press Return.
Part 1. Download and unzip the code
Section titled “Part 1. Download and unzip the code”Download the code (03-labs.zip). Double-click the file in your Downloads folder: the Mac unzips it into a folder named 03-labs. If a 03-labs folder is already there (Safari often unzips downloads for you), skip the double-click. It holds code and small data files only. Models download later, in the steps that need them.
Open Terminal and move into the folder. cd (change directory) moves Terminal into a folder:
cd ~/Downloads/03-labsEvery command in this course runs from this folder, unless a lesson’s own commands move into a lesson folder with cd. If you put the folder somewhere else, cd there instead.
Part 2. Create a virtual environment and install
Section titled “Part 2. Create a virtual environment and install”python3 -m venv .venv && source .venv/bin/activatepip install -r requirements.txt && pip install -e .The first line creates a virtual environment: a private set of Python add-ons for this folder only, so nothing else on your Mac changes. While it is on, (.venv) shows at the start of the line where you type commands. The second line installs what the course needs:
openai: code that sends requests to model servers.jsonschema: checks data against a format.matplotlib: draws charts.mlx-lm: Apple’s tool for running models on its chips.
Then pip install -e . installs felab, the course’s own toolkit that every lesson uses.
While it installs, watch Inference 101 at the top of this page (3 minutes).
In every new Terminal window, turn the environment back on before you run anything:
cd ~/Downloads/03-labs && source .venv/bin/activatePart 3. Create your settings file
Section titled “Part 3. Create your settings file”cp .env.example .env.env holds your settings. It already has the line FIREWORKS_API_KEY=fw_xxxxxxxxxxxxxxxx, where the x’s are a placeholder. In today’s third step, Measure TTFT and ITL, you replace everything after the = with your own Fireworks key. The file stays on your Mac: never share it or upload it anywhere.
Part 4. Start the practice server and run the smoke test
Section titled “Part 4. Start the practice server and run the smoke test”make mock & # the practice server, on port 9000 of this Macmake smoke # every offline lesson, end to end, about 2 minutesmake mock & starts the practice server and keeps it running in the background: the & lets you type the next command in the same window. It prints a line starting mock LLM on http://localhost:9000/v1.
make smoke then runs 16 of the kit’s 19 lessons, offline and free, with the practice server standing in for a real model. It prints a heading for each stage and should end with smoke test passed.
Leave the practice server running: the next step, Set up your Mac and the spend cap, uses it too. To stop it later, run kill %1 in the same window, or close the window.
Done when
Section titled “Done when”Stuck?
Section titled “Stuck?”make smokestops partway with an error.- Scroll up to the last heading it printed (for example ‘05 kv’): the error under it names the problem, and the rows below cover the usual ones.
make smokestarts the practice server itself when none is running. Ifmake checkthen shows mock down, start it withmake mock &and read the error it prints. - ‘No module named felab’, ‘No module named openai’ or ‘python: command not found’.
- The virtual environment is off, or the install did not finish. From the
03-labsfolder runsource .venv/bin/activate, thenpip install -r requirements.txt && pip install -e .again. Every new Terminal window needs thesourceline. - pip says the package ‘requires a different Python’.
- Your
python3is older than 3.10, the minimum the code needs. Install a current one withbrew install python, open a new Terminal window, delete the old environment withrm -rf .venv, and repeat part 2 on this page (Create a virtual environment and install). - ‘Address already in use’ when you start the practice server.
- One is already running, from earlier or in another window, and it keeps serving. Check with
make check: if mock says UP, carry on. makeis not found, or a window asks to install command line developer tools.- Accept the install: Apple’s command line tools include
make. Or runxcode-select --install, then try again.