Skip to content

Day 1 · Guardrails and first calls

Part 1 · Day 1 of 12

0 of 12 days done

About 1 h 25 4.9 GB model download · See Step 2 About $0.02 Mac and Fireworks

Today: Get the course code running on your Mac with a practice server: a free stand-in for a real AI model that gives made-up replies with realistic timing. Then put a spending cap on Fireworks, the cloud service that runs AI models for a fee. Finally, time how fast models answer: the wait for the first word, and the pace after it.

Example

Imagine asking a support assistant for help with a refund. Before answering, it reads the company rules and your question; then it writes its reply piece by piece. In the video’s example, you wait less than a second for the reply to start, but more than six seconds for it to finish. Today you will measure those two waits separately, so you can tell a customer which part needs speeding up.

By the end you’ll have

Your own timing table, saved in results/01-latency.csv: how long a reply takes to start and how quickly it appears after that, for a model on your Mac and for models on Fireworks. Step 3 explains a model in the script that is no longer available there.

Expected: In the lesson’s Mac test, a long message makes the reply take much longer to start than a short question. Once the reply starts, its words appear at almost the same pace. Your numbers will differ, but that split tells you whether to investigate the reading stage or the writing stage.

The 7 numbers you write down

  1. Your Mac’s memory (GB)

    How much model fits: plan on about 70% of it for the model and its notes.

    Example: 36

  2. Memory speed, bandwidth (GB/s)

    How fast memory feeds the chip. It caps how fast any model can write.

    Example: 273

  3. Writing-speed limit for Llama 3.1 8B at 4 bits (tokens/s)

    The fastest your Mac can write with the 4.9 GB model. Real runs reach 60 to 85% of it.

    Example: 56 (273 ÷ 4.9)

  4. Wait for the first word, short question (ms, p50)

    How long a typical person waits before the reply starts, after a one-line question.

    Example: 145

  5. Wait for the first word, 8,000-token prompt (ms, p50)

    The same wait when the prompt is about 6,000 words long: reading time grows with the prompt.

    Example: 4,210

  6. Gap between tokens while writing (ms)

    The typing pace after the first word. It barely changes with prompt length.

    Example: 22.4 (44.6 tokens per second)

  7. Wait for the first word on each Fireworks model (ms, p50)

    How hosted models compare on the same question, timed by your own stopwatch.

    Example: The lesson has no example here: write the model, then ms, for each model that answered.

How you will use this: When someone says an AI feature is slow, your timing table helps you ask whether the wait is before or after the reply starts. Day 3 adds how many people one server can handle. Together those measurements become evidence for the hardware recommendation you will write on Day 10.

Before you start

  1. An Apple-silicon Mac (M1 or later) with more than 5 GB of free disk space.

    Step 2 downloads your first real model, Llama 3.1 8B (8 billion learned numbers, each stored in about 4 bits, half a byte): about 4.9 GB. The tools installed in steps 1 and 2 need room on top of that. The setup script is for Macs with Apple’s own chips; on an older Intel Mac, MLX (Apple’s tool for running models on its chips) will not work.

  2. Python 3.10 or newer, and Homebrew.

    The course code needs Python 3.10 or later. Step 2’s setup script installs its tools with Homebrew (the Mac’s usual installer for developer tools) and stops if Homebrew is missing.

    From the 03-labs folder:

    python3 --version && brew --version

    If brew is not found, install Homebrew from brew.sh, then run the ‘Next steps’ lines it prints at the end; without them brew is still not found. If Python says 3.9 or older, run brew install python, open a new Terminal window and check again.

  3. A Fireworks account, for step 3.

    Until you add a payment method, Fireworks gives you $1 of credit and at most 10 requests a minute. Step 3 sends about 17 requests in quick succession (22 if all three of its models answered), so step 2’s checklist has you buy a small prepaid credit (item 5). Step 2 also sets the spending cap first; step 3 costs about 1 to 2 cents.

Start step 1: Get the code, run it offline (15 min)

Your progress is saved on this device.