Set up your Mac and the spend cap
course 2 of 22
lesson 00 · lessons/00-setup
30 min Free Mac
What you will do and why
Install the programs that run models on your Mac, and read its two limits: memory size and memory speed. Then cap your Fireworks spending before anything can bill.
Why it matters: A dedicated GPU on Fireworks (a chip that runs AI models, reserved for you alone) bills about $8 an hour while it is up, busy or idle. By default it stops billing only after a full hour with no requests. That idle hour alone, $8, is more than half of the whole course’s Fireworks budget (well under $15). All of step 3 costs about 1 to 2 cents.
You are done when: The check shows mock UP and ollama UP. All 7 spend-cap items are ticked (item 5 means buying a small prepaid credit, which step 3 and later days spend from). fireworks_guardrails.sh shows an empty deployment list.
Start this first
Run it from the 03-labs folder. The script installs Ollama, llama.cpp and MLX (three programs that run models on a Mac) and is safe to run twice. The last line starts Ollama and downloads Llama 3.1 8B, about 4.9 GB. This window stays busy until the download ends: open a second Terminal window (Command-N) for the spend-cap checklist on this page, and in it first run cd ~/Downloads/03-labs && source .venv/bin/activate. The lesson’s own commands below repeat these lines: skip the ones you have already run.
bash lessons/00-setup/setup_mac.shsource .venv/bin/activateollama serve & sleep 2 && ollama pull llama3.1:8bIn plain words
Section titled “In plain words”On a Mac, the model shares one pool of memory with every app you have open, so plan on it using about 70% at most. How fast that memory delivers data sets how fast the model can write. The check script prints both numbers. On Fireworks, a GPU reserved for you bills by the hour even when nobody uses it, so the spending cap goes on first.
Picture it
A kitchen with one storeroom shared by everyone. The cook’s recipe book (the model) must fit there next to everyone else’s supplies (your other apps). Before every dish (every token), the whole book comes over on a conveyor belt. So the belt’s speed caps the pace, however quick the cook is.
With real numberslesson 00’s sample Mac (36 GB, 273 GB per second) and the lab book’s spend-control section
- Memory budget: 0.7 x 36 GB = 25.2 GB, about 25 GB for the model plus its notes (the running record it keeps of each conversation, which Day 4 measures).
- Your first model, Llama 3.1 8B stored at 4 bits: 4.9 GB, about a fifth of that budget.
- Writing-speed limit: each token needs the whole 4.9 GB model moved from memory once. 273 GB per second ÷ 4.9 GB = about 56 tokens (word pieces) per second at most, about 42 words.
- Fireworks, paid per token (serverless): all of step 3 costs about 1 to 2 cents.
- Fireworks, a GPU reserved for you (a dedicated deployment): about $8 an hour while it is up, used or not. By default it stops after an hour with no requests; never count on that, delete it in the same sitting.
- The whole course costs well under $15 on Fireworks if nothing is left running. One forgotten dedicated GPU, if it stays up, passes that in under 2 hours (15 ÷ 8 = 1.9).
Words to know
- Unified memory
- One pool of memory shared by the Mac’s main processor (CPU) and its graphics processor (GPU). Example: the model and your browser share the same 36 GB.
- Bandwidth (memory bandwidth)
- How many GB per second memory delivers to the chip. Example: 273 GB/s in the lesson’s sample.
- Serverless
- Fireworks’ shared, always-on models, billed per token you send and receive. Example: step 3 costs about 1 to 2 cents.
- Dedicated deployment
- A GPU reserved for you, billed by the second at an hourly rate while it is up, used or not. Example: about $8 an hour.
From the lesson
lessons/00-setup/README.md
What and why
Section titled “What and why”You are building a small inference lab on your Mac: three local servers, one shared
Python toolkit (felab) and a fake server for offline practice. You also set the
Fireworks handbrake before spending anything.
On a Mac, RAM is your VRAM. The GPU and CPU share one pool of memory, so the model
competes with your browser for space. Memory bandwidth is your speed limit, and
check_env.py prints both.
bash lessons/00-setup/setup_mac.sh # installs ollama, llama.cpp, venv, felabsource .venv/bin/activatepython lessons/00-setup/check_env.py # chip, RAM, bandwidth, which servers are up
# first real model (≈4.9 GB download)ollama serve & # leave runningollama pull llama3.1:8bpython lessons/00-setup/check_env.py # ollama should now say UP ✓
# the offline practice server (most lessons can rehearse against it)python -m felab.mock_server & # :9000
# Fireworks: run after the spend-cap checklist on this page (before step 3)bash lessons/00-setup/fireworks_guardrails.shWhat each command does
bash lessons/00-setup/setup_mac.shInstalls Ollama and llama.cpp with Homebrew, then the same virtual environment and add-ons as step 1 (MLX among them). It skips anything already installed. Look for ‘Done. Next:’ at the end.
source .venv/bin/activateTurns the virtual environment on in this window:
(.venv)shows at the start of the line where you type. Needed in every new Terminal window.python lessons/00-setup/check_env.pyPrints your chip, memory and memory speed, then which model servers answer. Look for your bandwidth and the writing-speed limit after it (bandwidth ÷ 4.9 GB). Ollama says down until it runs.
ollama serve &Starts Ollama, the easiest of the three programs, in the background; its log lines may scroll by. If it says the address is already in use, Ollama is already running: carry on.
ollama pull llama3.1:8bDownloads Llama 3.1 8B stored at 4 bits, about 4.9 GB. Do the spend-cap checklist in a second Terminal window while it runs.
python lessons/00-setup/check_env.pyThe same check, after the download. Look for UP at the end of the ollama line.
python -m felab.mock_server &Starts the practice server from step 1 in the background. If it is still running from step 1, you see ‘Address already in use’: that is fine, the first one keeps serving.
bash lessons/00-setup/fireworks_guardrails.shChecks your Fireworks account: who you are signed in as, any dedicated deployments (they bill by the hour) and any training jobs. It needs firectl, Fireworks’ command-line tool, installed and signed in (checklist item 2). Look for an empty deployment list.
make fw-checkruns the same script.
What you should see
Section titled “What you should see”How to read it
You also see your chip’s name and one row per model server. llamacpp and mlx say down until Day 2, and spark (a separate NVIDIA machine) matters only on Day 9’s optional step. fireworks already says ‘key set’, but that is the placeholder in .env: your real key goes in at step 3. The bandwidth line ends with your writing-speed limit, bandwidth ÷ 4.9 GB (56 tokens per second at 273 GB/s). A Mac sold as 36 GB shows about 38.7 GB: the script counts a GB as 1,000,000,000 bytes, and Apple counts it as 1,073,741,824.
memory 36.0 GB ← model weights + KV cache must fit in ~70% of thisbandwidth ~273 GB/s ← decode ceiling for an 8B model at 4-bit (≈4.9 GB): ~56 tok/smock http://localhost:9000/v1 UP ✓ollama http://localhost:11434/v1 UP ✓The spend-cap checklist
Section titled “The spend-cap checklist”From the lab book. Tick each one as you do it; the lab book shows the same ticks.
The trap: a LoRA fine-tune can only be served on a dedicated deployment, which bills by the hour whether or not you use it, so delete it in the same sitting.
Check yourself
Section titled “Check yourself”2 questions. Say your answer out loud, then tap to check it.
Your Mac has 36 GB of memory. Roughly how big a model, stored at 4 bits per number, can it run comfortably?Show answerHide
In plain words
About a 40-billion-number model (40B). That fills the whole budget of about 25 GB, so there is little room left for the model’s notes on each conversation.
Picture it
Flatmates share one fridge. You can take about 70% of it before everyone else’s food gets squeezed out. Fill your whole share with one giant cake and there is no room for the leftovers each meal adds (the model’s notes).
With real numberslesson 00, and the 4.9 GB model from Day 1’s setup step
- Budget: 0.7 x 36 GB = 25.2 GB, about 25 GB for the model and its notes.
- Size per billion numbers at 4 bits: Llama 3.1 8B is 4.9 GB, so 4.9 ÷ 8 = about 0.61 GB per billion.
- Biggest model: 25 ÷ 0.61 = about 41 billion numbers, so about a 40B model.
- Room left for notes: almost none. Day 4 shows one long conversation’s notes can outgrow the model itself.
Words to know
- Parameters (8B, 40B)
- The model’s learned numbers, also called weights. Example: 8B = 8 billion.
- 4-bit
- Each number stored in 4 bits, half a byte, plus a little extra. Example: an 8B model is 4.9 GB at 4-bit.
- Dense model
- A model that uses all of its numbers for every token. Example: Llama 3.1 8B.
- Notes (KV cache)
- The model’s memory of each conversation so far, kept in the same memory as the model. Example: Day 4 measures it.
Go deeper: the engineer version
The kit's question
Your Mac has 36 GB. Roughly the biggest 4-bit model you can run comfortably?
The kit's answer
≈ 0.7 × 36 ≈ 25 GB of weights, so about a 40B dense model at 4-bit, with little room left for context.
More detail: 0.7 x 36 = 25.2 GB for weights and KV cache together. Real 4-bit files (Q4_K_M) carry per-block scales and keep some tensors at higher precision, so Llama 3.1 8B at 4 bits is 4.9 GB (about 4.8 bits per weight, lesson 04), not the 4.0 GB that 0.5 bytes x 8B would give. 25.2 ÷ 0.61 = about 41B parameters, so about a 40B dense model, which leaves almost no KV-cache room. check_env.py prints memory in billions of bytes (a 36 GB Mac shows 38.7), so the 70% rule is a rough budget either way.
Why does Day 1’s check script (check_env.py) print memory speed (bandwidth) rather than how many GPU cores (the chip’s calculating units) you have?Show answerHide
In plain words
Because memory speed caps how fast the model writes. For each token (word piece) the chip must fetch the whole model from memory once, and that fetching is slower than the math.
Picture it
A very fast cook needs the whole recipe book carried over on a conveyor belt before every dish. More cooks (more cores) would stand waiting; only a faster belt gets dishes out sooner.
With real numberslesson 00’s sample Mac and felab/hardware.py
- Each token needs the whole 4.9 GB model fetched once.
- At 273 GB per second: 273 ÷ 4.9 = about 56 tokens per second, at most.
- Real runs land at 60 to 85% of that limit (felab/hardware.py); Day 4 checks it on your Mac.
- A Mac whose memory moves 546 GB per second (the faster M4 Max) has twice the limit: 546 ÷ 4.9 = about 111.
Words to know
- Decode
- The writing phase: one token at a time, fetching the whole model for each.
- Bandwidth
- How many GB per second memory delivers to the chip. Example: 273 GB/s.
- GPU cores
- The many small calculating units inside a GPU. Example: they do the math, but on a Mac they mostly wait for memory while writing.
- Compute-bound
- Slowed by calculation speed rather than memory. Example: reading the prompt (prefill).
Go deeper: the engineer version
The kit's question
Why does the script print bandwidth rather than GPU cores?
The kit's answer
Decode reads every weight once per token, so bandwidth, not compute, sets the speed.
More detail: At batch size 1, each decode step streams every weight once and does only about two floating-point operations per weight, far less than the GPU could compute in that time. So the ceiling is bandwidth ÷ bytes read per token (felab/hardware.py), and measured decode typically reaches 60 to 85% of it. Prefill is the opposite: it reuses each weight across many prompt tokens, so compute sets its pace. Where a chip ships in two bandwidth versions (M3 Max, M4 Max), hardware.py lists the faster one.
Explain what you learned
Section titled “Explain what you learned”Question How big a model can this machine run, and how fast?
One clear answer
On unified memory, model size competes with everything else on the box. I size to about 70% of RAM and check the bandwidth before I promise a tokens-per-second number.
What this means
- “On unified memory”: On a Mac, the main processor (CPU) and the graphics processor (GPU) share one pool of memory; the GPU has none of its own.
- “model size competes with everything else on the box”: The model sits in the same memory as your browser and other apps (‘the box’ is the machine), so it cannot have all of it.
- “I size to about 70% of RAM”: RAM is the computer’s working memory. I plan for the model and its notes to use at most 70% of it: about 25 GB on a 36 GB Mac.
- “and check the bandwidth”: I look up how fast memory delivers data, for example 273 GB per second.
- “before I promise a tokens-per-second number”: Only then do I quote a writing speed. The ceiling is bandwidth ÷ model size: 273 ÷ 4.9 = about 56 tokens per second.
Your numbersSaved on this device and collected in the Day 1 wrap-up.
Hint: Apple menu, About This Mac, the Memory line. check_env.py prints a slightly bigger number (38.7 for a 36 GB Mac) because it counts a GB as a billion bytes.
Hint: The bandwidth line of check_env.py. M3 Max and M4 Max chips come in two memory speeds and the script shows the faster one, so confirm your exact model on Apple’s tech specs page and write that figure. If it says unknown, look it up there too.
Hint: Work it out from the bandwidth you wrote down: bandwidth ÷ 4.9. The script prints it at the end of the bandwidth line, using the faster speed on an M3 Max or M4 Max.
Done when
Section titled “Done when”The check shows mock UP and ollama UP. All 7 spend-cap items are ticked (item 5 means buying a small prepaid credit, which step 3 and later days spend from). fireworks_guardrails.sh shows an empty deployment list.
Stuck?
Section titled “Stuck?”setup_mac.shstops with ‘Install Homebrew first’.- Install Homebrew from brew.sh, run the ‘Next steps’ lines it prints at the end, then run the script again. It skips anything already installed.
check_env.pysays ‘bandwidth unknown (not an Apple chip)’ on an Apple Mac.- Your chip is newer than the kit’s list (felab/hardware.py knows M1 to M4 and the base M5). Look up your chip’s memory bandwidth on Apple’s tech specs page and work out the limit yourself: bandwidth ÷ 4.9.
ollama pullsays it could not connect.- The Ollama server is not running yet. Run
ollama serve &, wait two seconds, then pull again. fireworks_guardrails.shsays ‘firectl is not installed’.- Install it with
brew tap fw-ai/firectl && brew install firectl, then runfirectl signin, and run the script again. firectl whoamifails, or the script says ‘Run: firectl signin’.- Run
firectl signin(it opens a browser to log in), then run the script again. - A firectl command rejects its options.
- Commands change. Run it with
--help(for examplefirectl quota update --help) and follow what it shows. The Fireworks console’s billing page also has usage limits.
Code in this step
Section titled “Code in this step”check_env.py Python · 47 lines
"""Lesson 00 · step 2 — what have I got?
Prints your chip, RAM (= your "VRAM" on a Mac: unified memory is shared by theGPU and everything else), memory bandwidth (= your decode speed limit), andwhich model servers are answering right now.
python lessons/00-setup/check_env.py"""import osimport urllib.request
from felab import TARGETSfrom felab.hardware import detect
def up(url: str) -> bool: """A server is 'up' if GET /models answers within a second.""" try: with urllib.request.urlopen(url.rstrip("/") + "/models", timeout=1) as r: return r.status == 200 except Exception: return False
hw = detect()print("── machine ─────────────────────────────────────────────")print(f"chip {hw['chip']}")print(f"memory {hw['ram_gb']} GB ← model weights + KV cache must fit in ~70% of this")if hw["bw_gbs"]: bw = hw["bw_gbs"] print(f"bandwidth ~{bw} GB/s ← decode ceiling for an 8B model at 4-bit (≈4.9 GB): " f"~{bw / 4.9:.0f} tok/s")else: print("bandwidth unknown (not an Apple chip) — see felab/hardware.py")
print("\n── servers ─────────────────────────────────────────────")for name, t in TARGETS.items(): if name in ("fireworks",): state = "key set ✓" if t.api_key else "no FIREWORKS_API_KEY (fine until lesson 01b)" elif name == "spark" and not os.environ.get("SPARK_IP") and not os.environ.get("SPARK_URL"): state = "SPARK_IP not set (fine until lesson 14)" else: state = "UP ✓" if up(t.base_url) else "down" print(f"{name:10s} {t.base_url:45s} {state}")
print("\nStart the offline mock any time: python -m felab.mock_server")fireworks_guardrails.sh Bash · 47 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 00 · step 3 — Fireworks, with the handbrake on.## The whole course costs well under $15 on Fireworks IF you follow one rule:## Serverless is pay-per-token (cents). Dedicated deployments bill per GPU-hour# from the moment they start, whether or not you send a request.## Lessons 10 and 13 create deployments. They end with teardown. This script is# the "is anything still running?" check — run it at the end of every session.## Commands follow the firectl docs at time of writing; if a flag has moved,# `firectl <command> --help` is the source of truth.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefail
if ! command -v firectl >/dev/null; then cat <<'EOF'firectl is not installed. On a Mac: brew tap fw-ai/firectl && brew install firectl firectl signin(or see https://docs.fireworks.ai/tools-sdks/firectl/firectl)EOF exit 1fi
echo "▸ Who am I?"firectl whoami || { echo "Run: firectl signin"; exit 1; }
echoecho "▸ Deployments (anything listed here is billing by the hour):"firectl deployment list
echoecho "▸ Fine-tuning jobs (a running job bills training tokens):"firectl sftj list 2>/dev/null || true
cat <<'EOF'
Checklist [ ] A monthly budget / spend alert is set in the console (Billing → usage limits). [ ] The deployments list above is empty, unless you are mid-lesson. [ ] Your key is in .env, not in any file you commit.
To delete a leftover deployment: firectl deployment delete <DEPLOYMENT_ID>EOFsetup_mac.sh Bash · 41 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 00 · step 1 — install the local toolchain on an Apple-silicon Mac.## Ollama convenience server (hides the flags) → :11434# llama.cpp the engine under Ollama, every flag exposed → :8080# MLX Apple's own array framework; fastest on M-series, can fine-tune## Safe to re-run: each step checks before installing.# Run from the repo root: bash lessons/00-setup/setup_mac.sh# ─────────────────────────────────────────────────────────────────────────────set -euo pipefail
say() { printf "\n\033[1;36m▸ %s\033[0m\n" "$*"; }
[[ "$(uname -s)" == "Darwin" ]] || { echo "This script is for macOS. On Linux use the Spark track (lesson 14)."; exit 1; }[[ "$(uname -m)" == "arm64" ]] || echo "⚠ Intel Mac detected: MLX will not work; Ollama/llama.cpp will be slow."
say "Homebrew"command -v brew >/dev/null || { echo "Install Homebrew first: https://brew.sh"; exit 1; }
say "Ollama + llama.cpp (Metal builds — no CUDA on a Mac)"brew list ollama >/dev/null 2>&1 || brew install ollamabrew list llama.cpp >/dev/null 2>&1 || brew install llama.cpp
say "Python venv with the repo's helper package (felab) + openai SDK + mlx-lm"python3 -m venv .venv# shellcheck disable=SC1091source .venv/bin/activatepip install -q --upgrade pippip install -q -r requirements.txtpip install -q -e . # makes `import felab` work from any lesson folder
say "A .env file for your settings (git-ignored)"if [[ ! -f .env ]]; then cp .env.example .env; echo " created .env — edit it later for Fireworks / Spark"; fi
say "Done. Next:"cat <<'EOF' source .venv/bin/activate python lessons/00-setup/check_env.pyEOF