Glossary
404 terms from the course, each with an example.
- .env file
The settings file in the
03-labsfolder, holding defaults and your Fireworks key; it never leaves your Mac.Example:
cp .env.example .envcreates it in Day 1’s first step, with a placeholder key.- 4-bit, 8-bit, 16-bit
How many bits store each number; fewer bits, smaller model.
Example: an 8B model is 8.5 GB at 8-bit, 4.9 GB at 4-bit.
- β (beta)
The leash in RLHF and DPO: how far training may pull the model from its starting copy; smaller means a longer leash.
Example: lesson 16’s toy loosens it from 0.50 to 0.02.
- Acceptance rate
The share of a draft model’s guessed tokens the big model keeps; speculative decoding only pays when it is high.
Example: α (alpha) = 0.78 on structured prompts and 0.41 on prose in lesson 13’s sample.
- Accuracy (category accuracy)
The share of replies with the right answer, checked against each test example’s known label.
Example: 78% of 50 tickets on the practice server (
category_acc_%).- Activations
The working numbers a model computes for each example during training, kept for the correction step; they grow with batch size and example length.
Example: they come on top of the 16 bytes per number.
- Active parameters
The numbers a mixture-of-experts model actually uses for each token, as opposed to all the numbers it stores.
Example: 3.6 of gpt-oss:20b’s 21 billion (17%).
- Adam
The usual training method for adjusting a model’s numbers; it keeps two running averages per trainable number.
Example: its two averages are part of the 16 bytes per number a full fine-tune needs.
- Adapter (LoRA adapter)
The small set of extra numbers LoRA trains and saves; the base model stays unchanged.
Example: about 30 MB in the Training Techniques video.
- Adaptive speculative decoding
Fireworks’ option that trains the draft model on a customer’s own traffic, so more guesses are kept.
Example: the Batching & Scheduling video: the drafter trained on each customer’s own traffic.
- Agent
A program that calls a model many times, using tools, to finish a task.
Example: Day 5’s Acme Cloud support agent, with 5 tools.
- All-reduce
A network step where every GPU combines its partial results with all the others, so each ends up with the total.
Example: tensor parallelism does this inside every layer.
- All-to-all
A network step where every GPU sends a different piece of data to every other GPU.
Example: moving each token to the GPUs that hold its experts.
- API
The agreed format one program uses to send requests to another and get answers back.
Example:
/v1/chat/completionsis where a chat request goes.- API key
A secret code that proves a request comes from your account, so it can be billed.
Example:
FIREWORKS_API_KEYin.env.- Arithmetic intensity
How many calculations the chip does for each byte it fetches from memory.
Example: about 1 to 2 for one person’s decode; an H100 needs about 300 to keep its math busy.
- arm64
The chip design used by Apple silicon and by the DGX Spark’s processor; software must be built for it.
Example: container images tagged for GB10/arm64.
- Attention (the look-back step)
The step where each new token looks back at the notes on earlier tokens.
Example: each writing step reads every earlier token’s notes, which is why a long prompt slows it a little.
- Attention head (query head)
One of the parallel parts of a layer that looks back over the conversation.
Example: Llama 3.1 8B has 32 per layer; with GQA they share 8 note sets.
- Auto Reload
Fireworks’ setting that tops up your prepaid credit automatically when it runs low.
Example: leave it off while you are learning.
- Automatic prefix caching
vLLM’s built-in prefix cache: it reuses saved notes for identical blocks at the start of a prompt.
Example:
vllm serve $MODEL --enable-prefix-caching.- Autoscaling
Adding or removing copies of a model as traffic rises and falls.
Example: one of the layers a managed platform runs for you.
- B200
NVIDIA’s Blackwell data-center GPU, 192 GB of memory at 8,000 GB/s.
Example: 8,000 ÷ 70 = 114 tokens/s limit for a 70B model at 8 bits.
- Back-fill
Running a job over all your existing records at once, for example after adding a new field.
Example: re-sorting last year’s tickets into categories.
- Bandwidth (memory bandwidth)
How many GB per second memory delivers to the chip.
Example: M4 Max 546 GB/s, H100 3,350 GB/s.
- Base model
The model a fine-tune starts from and leaves unchanged.
Example: Qwen2.5 7B Instruct on Day 7.
- Batch API
Fireworks’ way to send many requests as one file and collect the answers later, at about half the serverless price.
Example: Day 6’s 100 tickets in
batch_input.jsonl.- Batch job (batch-inference job)
One run of the Batch API: your uploaded file of requests, processed when GPUs are free.
Example:
triage-batch-…, checked every 30 seconds until COMPLETED.- Batch size
How many requests share one pass over the model.
Example: 20.5 tokens a second at batch 1 and 368 in total at batch 32 on a Spark (lab book).
- Batch size (training)
How many examples the model learns from in one training step.
Example: 2 in
1_train_mlx.sh; lower it to 1 if memory runs out.- Batching
Serving several people with one pass over the model.
Example: 8 people share each fetch of the weights.
- Benchmark (public benchmark)
A standard test that model makers report scores on, used to compare models at launch.
Example: it says little about one customer’s tickets.
- bf16 (bfloat16)
A 16-bit number format used for training, 2 bytes per number.
Example: multi-LoRA on Fireworks needs a BF16 deployment shape.
- Bias (judge bias)
A judge’s habit of favouring something that is not correctness.
Example: longer, more polished answers.
- Bit
A single 0 or 1; 8 bits make a byte.
Example: a 16-bit number is 2 bytes.
- Blackwell
The NVIDIA GPU generation after H100, with built-in 4-bit math.
Example: B200, 8,000 GB/s.
- Block (KV block)
A fixed-size piece of note memory holding the notes for a few tokens; vLLM hands blocks out as conversations grow.
Example: 16 tokens per block; 8,234 blocks in the Serving With vLLM startup log.
- Block hash
A fingerprint of one fixed-size block of a prompt together with everything before it; a matching fingerprint means the saved notes can be reused.
Example: 64-token blocks in the practice server.
- Bottleneck
The slowest part of a process, the one that holds everything up.
Example: memory speed while the model writes.
- Bradley-Terry model
The rule that turns two scores into the chance a person prefers one: the sigmoid of the gap between them.
Example: scores 2.1 and -0.4 give about 92%.
- Break-even
The point where two options cost the same.
Example: about 800 million tokens a day for serverless against one dedicated H100, in the Platform Map example.
- Build versus buy
Doing a job with your own hardware and people, or paying a service to do it.
Example:
COMPARISON.md: the same fine-tune on your Mac and on Fireworks.- Bursty traffic
Traffic that comes in sudden peaks with quiet spells between.
Example: the Platform Map video matches it to serverless.
- Business case
The argument, in money and results, for making a change.
Example: 88% off the input bill for an agent whose input is 90% repeats, at the video’s prices.
- BYOC (bring your own cloud)
A managed service’s software running inside the customer’s own cloud account, so the data stays there.
Example: the lab book’s answer to “Can we keep data in-region?”
- Byte
8 bits.
Example: one FP16 number takes 2 bytes.
- Cache hit (hit rate)
A request that finds its opening already saved; the hit rate is the share of requests that do.
Example: the Prefix Caching video: measure the hit rate, because an unused cache only takes memory.
- Cached input
Input tokens the server reused from an earlier request, billed at a discount.
Example: $0.015 instead of $0.15 per million for gpt-oss 120B.
- Cached tokens (cached_tokens)
Input tokens the server reused instead of reading again; Fireworks reports them and bills them at a discount.
Example:
usage: prompt_tokens=… cached_tokens=…from Day 5’s--usagerun.- Capacity (of one replica)
The most people one copy of the server serves at once while keeping the speed promise.
Example: 8 on the practice server within 500 ms.
- Capacity modes
The ways a platform sells computing: serverless, dedicated, batch and reserved.
Example: the four in the Platform Map video.
- Ceiling (speed limit)
The fastest a model can write: memory speed ÷ model size (for a mixture-of-experts model, only the part it reads per token).
Example: 546 ÷ 4.9 = 111 tokens/s.
- Changelog
A provider’s dated list of changes to its product.
Example: where the lab book found adaptive rate limits described.
- Chatty answer (chatty_%)
A reply that wraps the JSON in chat or writes prose, so a program cannot read it.
Example: about 20% before DPO, about 1% after, as the lesson expects.
- Checkpoint (saved model file)
The file a training run saves: the whole changed model for a full fine-tune, only the adapter for LoRA.
Example: about 990 MB against about 9 MB in lesson 17’s expected table.
- Chip
The processor that does the model’s math.
Example: an Apple M4 Max, or an NVIDIA H100 (a GPU).
- Chosen and rejected
The two answers in a preference pair: the one a person approved, and the one they edited away or turned down.
Example: chosen: the clean JSON; rejected: the same JSON wrapped in chat.
- Chunked prefill
Reading a long prompt in pieces, between other people’s writing steps, so one long prompt does not freeze everyone else.
Example:
--enable-chunked-prefillin vLLM.- Closed model
A model you can only rent through its maker’s service; you never get its weights.
Example: where the Platform Map video’s customer journey starts.
- Cluster
Several GPUs, often across machines, working as one.
Example: what a full fine-tune of an 8B model needs, at about 128 GB.
- Code completion
An editor feature that suggests the next few lines of code.
Example: 200 tokens in, 30 out, in the Inference 101 video.
- Cold and warm
Cold: nothing is saved yet, so the whole prompt is read; warm: saved notes are reused. Not the same as a cold start (Day 1), when a server is still loading the model.
Example: call 1 is cold (740 ms) and calls 2 to 12 are warm (36 ms) in Day 5’s sample.
- Cold start
A slow request while a server that was idle or newly started loads the model.
Example: why Day 1’s stopwatch never counts its first call.
- Completion
The text the model writes back.
Example: the JSON reply to one ticket.
- Compute-bound
Slowed by calculation speed rather than memory.
Example: prefill.
- Concurrency (people served at once)
How many conversations the server runs at the same moment.
Example: 8 on the practice server.
- Concurrency sweep
The same test run at 1, 2, 4, 8, 16 and 32 people at once, then drawn as a curve.
Example: lesson 03’s
sweep.py.- Config (settings file)
The file shipped with a model that lists its shape.
Example: 32 layers, 8 KV heads, head dim 128.
- Constrained decoding
At each step the server blocks every token that would break the schema or grammar, so the reply cannot come out malformed.
Example:
json_schemamode: 50 of 50 valid in lesson 07’s practice run. A reply cut off by the length limit (max_tokens) can still break.- Container
A sealed package holding a program and the exact software versions it needs, so it runs the same on any machine.
Example: Serving With vLLM starts the
vllm/vllm-openai:latestcontainer withdocker run(Docker is the usual container tool).- Container image (tag)
The packaged software a container starts from; the tag names its version.
Example:
nvcr.io/nvidia/vllm:25.09-py3in1_vllm.sh.- Context length (context, ctx)
The most tokens one conversation may hold.
Example: 8K = 8,192 tokens.
- Continuous batching
Letting requests join and leave the shared group at every writing step, instead of waiting for the whole group to finish.
Example: a finished reply’s seat is refilled at the next step, like a taxi rank rather than a tour bus.
- CPU
The computer’s general-purpose processor, slower than the GPU for model math.
Example: model layers not placed on the GPU run here.
- Crossover
The monthly volume at which two ways of serving cost the same; the memo’s name for the break-even.
Example: about 744 million tickets a month for the memo’s example customer, if its current 18 copies could carry that much.
- CSV file
A plain table file: one row per line, commas between the columns.
Example:
results/01-latency.csv.- CUDA
NVIDIA’s software layer that lets programs run on its GPUs.
Example: a Mac has none; llama.cpp uses Metal instead.
- CUDA graphs
A GPU speed-up that records a sequence of work once and replays it, instead of launching each piece separately.
Example: the documented no-container workaround gives it up and loses 20 to 30% of throughput.
- custom_id
The label on each request in a batch file, used to match every answer back to its question.
Example:
t-000tot-099.- Customer journey (maturity curve)
The usual path: rent a closed model, improve the prompt and context, move to open models, then train on your own data.
Example: the Platform Map video’s four stages.
- Dataset (Fireworks dataset)
A file stored in your Fireworks account for jobs to read or write.
Example:
firectl dataset createuploads the batch file; the job writes its results to another.- Decode
The writing phase, one token at a time, fetching the whole model for each (a mixture-of-experts model fetches only the part it uses).
Example: 86 tokens/s at 4-bit.
- Dedicated deployment
A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not.
Example: about $8 an hour for one H100 (lab book); by default it stops only after a full hour with no requests.
- Dense model
A model that uses all of its numbers for every token.
Example: Llama 3.1 8B.
- Deployment (production)
A model running as a live service for real users.
Example: running out of note memory is the most common production failure.
- Deployment shape
A preset for a dedicated deployment on Fireworks: the GPU and the number format it runs.
Example:
--deployment-shape defaultin3_deploy.sh.- DGX Spark
NVIDIA’s desktop AI computer, with 128 GB of memory shared by its CPU and GPU.
Example: about 273 GB/s of memory speed; Day 9’s optional step runs vLLM and SGLang on it.
- Discovery call
The first meeting where you learn a customer’s task, volume, speed target, quality bar and data limits.
Example: the memo is what you send after it.
- Docker
The common tool for building and running containers.
Example: Serving With vLLM starts vLLM with
docker run.- Document classification
Sorting documents into a fixed set of categories.
Example: ticket triage is one kind.
- DoRA
A LoRA variant that splits each weight into a size (magnitude) and a direction, and trains LoRA on the direction.
Example: often 1 to 2 points above LoRA, a little slower (the Parameter-Efficient Fine-Tuning video).
- DPO (direct preference optimization)
Training from pairs of answers where one is marked better, to teach taste or judgement.
Example: Day 12 runs it on your Mac.
- Draft model
A small, fast model that guesses the next few tokens for a big model to check.
Example: Llama 3.2 1B drafting for Llama 3.1 8B in lesson 13.
- Driver
The software that lets the operating system use a piece of hardware such as a GPU.
Example: the driver, CUDA and the Python packages must agree on an NVIDIA machine.
- Efficiency %
Measured writing speed ÷ ceiling.
Example: 86 ÷ 111.4 = 77%.
- Embedding
A list of numbers that captures a text’s meaning, used for search and matching.
Example: recomputing them for every document is a typical batch back-fill.
- Endpoint
A web address that answers requests for one model.
Example:
http://localhost:8081/v1for Day 8’s fine-tune.- Enforce
make the server block anything that breaks the rules.
Example:
json_schemamode enforces the schema.- Engine (inference server)
The program that loads a model and answers requests.
Example: llama.cpp, vLLM.
- Enum
A field that may hold only one value from a fixed list.
Example:
category: billing, outage, how-to or abuse.- Epoch
One full pass through all the training examples.
Example: 2 epochs over 800 tickets on Day 7, so every token is billed twice.
- Error file
Where a batch job puts the requests that failed, apart from the results.
Example: why
answeredcan be below 100.- Eval (eval harness)
A repeatable test of a model on a fixed set of prompts.
Example: 50 of the customer’s prompts in lesson 09 (Day 7).
- Expert
One specialist block inside a mixture-of-experts model; only the few the router picks are read for a token.
Example: lesson 12’s picture: 2 to 4 specialists out of dozens per token.
- Expert labels
Preference choices made by people who can check correctness.
Example:
toy_alignment.py --expert-labels: the reward model then agrees with them 78% of the time.- Expert parallelism
Placing different experts on different GPUs, so tokens travel to whichever GPU holds their chosen experts.
Example: the field guide’s 4 experts on 4 GPUs.
- f16 / FP16
A 16-bit number format, 2 bytes per number.
Example: the default format for the notes.
- Fabric (interconnect)
The links that join GPUs or whole machines so they can work on one model.
Example: two DGX Sparks joined by a 200-gigabit cable.
- felab
The kit’s small shared Python toolkit that every lesson script imports.
Example:
felab/targets.pylists every server’s address.- Fetch
Move data from memory to the chip that does the math.
Example: the whole 4.9 GB model is fetched once for every token written.
- Few-shot examples
A few solved examples placed in the prompt to show the model what a good answer looks like.
Example: tickets shown with their correct JSON before the real one.
- Fine-tune
Further training of an existing model on your own examples.
Example: a LoRA fine-tune on Day 7.
- firectl
Fireworks’ command-line tool for your account, quotas and deployments.
Example:
firectl deployment listshows what is billing by the hour.- Fireworks
A cloud service that runs AI models and charges per use; one platform used in this course.
Example: Day 1’s step 3 times its models for about 1 to 2 cents; one of the three, gpt-oss-20b, is expected to fail.
- Flag
An option typed after a command that changes how the program runs.
Example:
--parallel 4gives llama.cpp 4 slots.- Flash attention
A faster, memory-saving way to run attention; llama.cpp needs it to store the values in 8 bits.
Example:
EXTRA="-fa on".- Floating-point
A way of storing numbers with a sign, a size and a precise part, like scientific notation.
Example: FP16 and FP8.
- FLOPS (floating-point operations per second)
How many calculations on decimal numbers a chip can do each second; a teraflop is a trillion.
Example: an H100 does about 989 trillion (989 teraflops).
- Forgetting
Getting better at the trained task and worse at everything else.
Example: a drop in the 12-question general quiz, usually in the full row.
- FP4
A 4-bit number format, half a byte per number, with built-in math on the newest NVIDIA GPUs.
Example: FP4 deployment shapes cannot host LoRA adapters.
- FP8
An 8-bit number format, 1 byte per number, with built-in math on H100.
Example: a 70B model is 70 GB in FP8.
- fp32
A 32-bit number format, 4 bytes per number; training keeps the optimizer’s averages and the master copy in it.
Example: 12 of the 16 bytes per number in a full fine-tune.
- Fragmentation
Memory lost in gaps too small to use.
Example: part of the about 10% on top of weights and notes.
- Frontier model
One of the largest, most capable models, usually priced highest.
Example: $6.00 per million output tokens in the Platform Map example.
- Frozen
Left unchanged during training.
Example: QLoRA keeps the 4-bit model frozen.
- Full fine-tune
fine-tuning that changes every number in the model.
Example: about 121.6 GB of memory for a 7.6B model, by Day 8’s memory script.
- Fuse (fused model)
Baking the adapter into the base model’s numbers to make one self-contained model.
Example: the step after QLoRA training on Day 8.
- Gated model
A model you must request access to on Hugging Face before you can download it.
Example: why lesson 14 asks for
HF_TOKEN.- GB
1,000,000,000 bytes.
Example: 16,384 MiB = 17.18 GB.
- GB10
The DGX Spark’s chip: an NVIDIA Grace CPU and Blackwell GPU in one package, arm64, with 128 GB of shared memory.
Example: about 273 GB/s of memory speed.
- General quiz
Day 4’s 12 quick questions (math, facts, format), reused in lesson 17 to spot forgetting.
Example: about 9 of 12 for the untrained model.
- Generalise
Do well on new examples, not only the ones trained on.
Example: 89% on held-out examples for the clean run in the QLoRA On Your Own Box video.
- GGUF
llama.cpp’s model file format, one file per size.
Example: a
Q4_K_M.gguffile.- GiB
1,073,741,824 bytes (1,024 MiB), a little more than a GB.
Example: 59 GiB = 63.4 GB.
- Gigabit (Gb)
A billion bits; 8 gigabits make 1 GB.
Example: a 200-gigabit link moves about 25 GB a second.
- Goodhart’s law
When a measure becomes the target, it stops being a good measure.
Example: the toy’s average rating rises while the average true value falls.
- Goodput
The requests per second that met the speed promise; late ones do not count.
Example: 2.5 per second at 8 people on the practice server, 0.4 at 16.
- gpt-oss-20b, gpt-oss-120b
Two open mixture-of-experts models with about 21 and 117 billion parameters, of which only about 3.6 and 5.1 billion are used for each token. On Fireworks, gpt-oss-120b is billed per token (serverless); gpt-oss-20b no longer is (its model page, checked 27 September 2026: ‘Serverless: Not supported’), but batch jobs still take it.
Example: Day 6 runs gpt-oss-120b in step 1 and gpt-oss-20b for step 2’s batch job.
- GPU
The chip that runs the model.
Example: H100; on a Mac it shares memory with everything else.
- GPU cloud (raw GPU cloud)
A service that rents bare GPUs; the customer runs everything above them.
Example: the bottom layer of the Platform Map.
- GPU cores
The many small calculating units inside a GPU.
Example: on a Mac they mostly wait for memory while the model writes.
- GPU memory utilization (--gpu-memory-utilization)
The share of the GPU’s memory vLLM may claim for the model and its notes.
Example: 0.8 means 80% of the card.
- GPU-hour
One GPU rented for one hour.
Example: $8 for an H100; a month is about 730 of them.
- GQA (grouped-query attention)
Groups of attention heads share one set of notes.
Example: 8 KV heads instead of 32.
- Grader
Anything that scores a model’s reply: a code check, a judge model or a person.
Example: Day 7’s eval uses three, cheapest first.
- Gradient
The correction worked out for each trainable number at each training step.
Example: full fine-tuning keeps one for every number of the model.
- Gradient checkpointing (
--grad-checkpoint) Recomputing activations during training instead of storing them: less memory, somewhat slower.
Example: about 30% slower (the Full Fine-Tuning video).
- Grammar
A precise rulebook for which text is allowed, which a server can enforce one token at a time.
Example: llama.cpp turns a JSON schema into one.
- GRPO
A reinforcement method that scores a group of answers to the same prompt against each other, with no value model.
Example: used for many reasoning models (the RLHF video).
- H100
NVIDIA’s data-center GPU, 80 GB of memory at 3,350 GB/s.
Example: about $8 an hour on a dedicated Fireworks deployment (lab book).
- H200
The H100’s successor with faster memory.
Example: 4,800 GB/s against 3,350.
- Hand-labelled sample
Answers a person has marked right or wrong, used to check a judge.
Example: about 20 of the judge’s scores, redone by you.
- Harness
One test script you can point at any server, so the numbers compare fairly.
Example:
bench_ttft.py, built on Day 1 and reused on Day 2.- Head dim
How many numbers each note set holds per token.
Example: 128 in Llama 3.1 8B.
- Headroom
Spare capacity kept for surprises.
Example: plan 18 replicas where 15 are needed.
- Health endpoint
An address that answers once the server is ready to take requests.
Example:
curl -sf localhost:8000/health.- Held-out accuracy
How often the model is right on examples set aside before training and never trained on.
Example: 74% for the messy run, 89% for the clean run in the QLoRA On Your Own Box video.
- Held-out error
How wrong a trained model is on new examples it never saw; lower is better.
Example: 0.1321 for the toy’s rank-1 add-on at 200 examples.
- Held-out examples
Examples set aside before training and never trained on.
Example:
test.jsonl, 100 tickets; the toy tests on 500 new inputs.- Held-out prompts
Test prompts the model was never trained on.
Example: lesson 09’s
test.jsonl(Day 7).- HF_TOKEN (Hugging Face token)
A personal key that lets scripts download the models your Hugging Face account may use.
Example: set in
.env;remote.shpasses it to the Spark.- Homebrew (brew)
A Mac tool that installs programs from the command line.
Example:
brew install llama.cpp.- Hot-swap
Change which adapter a running server applies, without restarting it.
Example: what small adapter files make easy in multi-LoRA serving.
- Hugging Face
The public website where most open model files are shared.
Example:
serve_llamacpp.shdownloads its 4.9 GB file from there.
- Imitation model (SFT model)
The model after it copied examples of good answers.
Example: the toy’s step 2 model, which the preference pairs are drawn from.
- Inference
Running a trained model to get answers, as opposed to training it.
Example: an inference server answers requests.
- Input tokens
The tokens you send to the model; providers bill them per million.
Example: $0.30 per million, or $0.006 cached, in the Prefix Caching video.
- Instruct model
A model already trained to follow instructions and chat.
Example: Qwen2.5 7B Instruct.
- Integer
A whole number.
Example: q8_0 stores notes as 8-bit integers plus a shared scale.
- Interactive traffic
Requests where a person is waiting for the answer.
Example: a chat window; it stays on serverless or dedicated, never batch.
- Iteration (training step)
One update: the model sees a small batch of examples and adjusts.
Example: 300 in the QLoRA On Your Own Box video’s Mac example.
- ITL (inter-token latency)
The gap between one token and the next while a reply is being written.
Example: 22.4 ms in Day 1’s sample Ollama row: 1,000 ÷ 22.4 = 44.6 tokens per second.
- JSON
A plain-text format for data that programs can read: named fields inside curly brackets.
Example:
{"category": "billing", "severity": 3}.- JSON schema (schema)
A rulebook for JSON: which fields must appear, what type each holds and which values are allowed.
Example:
data/triage_schema.json: 3 fields, 4 categories, severity 1 to 5.- json_object, json_schema
The two enforced modes:
json_objectguarantees valid JSON with any fields;json_schemaguarantees JSON that follows your schema.Example: both 100% valid in lesson 07’s practice run.
- JSONL
A file with one JSON record per line.
Example:
batch_input.jsonl, 100 lines.- Judge (LLM judge)
A second model that scores answers against a rubric.
Example: one column of lesson 09’s table (Day 7).
- k (draft length)
How many tokens the draft model guesses per round.
Example: up to 8 in lesson 13’s llama.cpp server, 4 on Fireworks.
- K (in 8K, 32K, 128K)
1,024 when counting tokens.
Example: 128K = 131,072 tokens.
- KB, KiB
1,024 bytes in this course (the videos write KB,
kv_calc.pyprints KiB).Example: 128 KB = 131,072 bytes.
- Kernel
A small program that runs one piece of the model’s math on the GPU.
Example: running the model with fast kernels is a serving engine’s fourth job.
- Keys and values (K and V)
The two kinds of notes the model stores per token, per layer.
Example: the 2 in 2 x layers x KV heads x head dim x bytes.
- KL (Kullback-Leibler divergence)
A measure of how far one model’s choices have drifted from another’s; 0 means identical.
Example: the KL leash charges the model for that drift (previewed on Day 11, run on Day 12).
- KL leash (KL penalty)
A charge for drifting from the starting model, added to the reward so training cannot run off after loopholes.
Example: lesson 16’s toy loosens it from β 0.50 to 0.02.
- Knee
The point where seats run out and waiting time shoots up.
Example: after 8 people on the practice server.
- KV cache
The model’s notes on the conversation so far (keys and values), so it does not reread everything for each new word.
Example: 17.18 GB for one 131,072-token conversation.
- KV cache dtype
The number format the notes are stored in; 8-bit instead of 16-bit halves their memory.
Example:
--kv-cache-dtype fp8in vLLM.- KV head (note set)
One set of notes kept in each layer.
Example: 8 per layer in Llama 3.1 8B.
- Label
The known right answer stored with each test example.
Example: each test ticket’s category and severity.
- Labeller
A person who picks the better of two answers to make training data.
Example: the toy’s fast labellers cannot check the ticket’s category.
- Latency
How long someone waits.
Example: time to the first token.
- Latency-bound
A workload where what matters is how fast each reply arrives, typical of one or a few people at a time.
Example: where speculative decoding helps most.
- Layer
One of the stacked processing stages each token passes through; each keeps its own notes.
Example: 32 in an 8B model, 80 in a 70B.
- Leak (data leakage)
Data used to judge a model has also shaped it, so the score looks better than it is.
Example: the validation tickets, once they set the stopping point.
- Learning rate
How big a nudge each training step gives each number.
Example: 1e-4 (0.0001) for LoRA, 1e-5 for full.
- Lever
A change that moves cost or speed without new hardware.
Example: the memo’s five: caching, constrained output, batch, a LoRA fine-tune, speculative decoding.
- Llama 3.1 8B, Llama 3 70B
The open models used in the lessons and videos, with 8 and 70 billion parameters.
Example: 4.9 GB at 4-bit for the 8B.
- Llama 3.1 405B
A 405-billion-parameter Llama model with 126 layers (field guide).
Example: 810 GB of weights at 16 bits.
- llama.cpp
The open-source engine used on the Mac;
llama-serveranswers requests,llama-benchtimes the engine alone.Example: Day 2’s server.
- Load generator
A tool that sends many requests at set rates to test a server.
Example:
vllm bench servein2b_vllm_bench.sh.- Localhost
Your own computer, used as an address.
Example:
http://localhost:8080/v1is llama.cpp on your Mac.- Log-probability (log-prob)
The natural log of how likely a model is to write a given answer; DPO and
pref_eval.py --pairwisework in these.Example: the chosen answer +0.8 against the reference in the DPO video, about 2.2 times as likely.
- LoRA
A cheap way to customise a model: train a small add-on instead of changing the whole model.
Example: on Fireworks it can only be served on a dedicated deployment.
- Loss (training loss)
A score of how wrong the model’s guesses are on the examples it trains on; lower is better.
Example: falling loss with flat test accuracy means memorising.
- M4 Max, M4 Pro
Apple chips whose memory moves 546 and 273 GB/s.
Example: 546 ÷ 4.9 = 111 tokens/s on an M4 Max; 273 ÷ 4.9 = about 56 on an M4 Pro.
- make
A tool that runs named shortcuts listed in a file called
Makefile.Example:
make smoke,make fw-check.- Managed platform
A service that runs everything from the GPUs up to the model: the engine, scaling, routing and the models.
Example: Fireworks; its serverless models charge per token instead of per GPU-hour.
- Managed service
A provider runs the hardware and software for you and bills for use.
Example: Fireworks training and serving Day 7’s fine-tune.
- Margin
In DPO, how much more the model favours the chosen answer than the rejected one, in log-probability.
Example: 1.2 in the DPO video (against the frozen copy);
pref_eval.py --pairwisereports the raw gap, about +1 before and +10 after.- Markdown
Plain text with light formatting marks, such as # for a heading; the files end in
.md.Example:
SIZING_MEMO.md.- Mask (masking)
Block a token so the model cannot choose it.
Example: constrained decoding masks every token that would break the schema.
- Master copy
A 32-bit copy of every weight kept during training, so many tiny nudges add up accurately.
Example: 4 of the 16 bytes per number in a full fine-tune.
- matplotlib
The Python library that draws the kit’s charts.
Example: without it,
plot_sweep.pyprints a text chart instead.- Matrix math
Multiplying large grids of numbers, the core work of a model.
Example: what tensor cores speed up.
- Max model length (--max-model-len)
vLLM’s limit on how many tokens one conversation may hold, which caps that conversation’s notes.
Example:
--max-model-len 8192.- Maximum concurrency (vLLM log)
vLLM’s startup estimate of how many full-length conversations fit in its memory for notes: cache tokens ÷ the length cap.
Example: 8,234 blocks × 16 tokens ÷ 8,192 = about 16 in the Serving With vLLM video.
- MB
1,000,000 bytes.
Example: about 1.6 MB of notes for a 12-token prompt on Llama 3.1 8B.
- Memory-bound (bandwidth-bound)
Slowed by how fast memory delivers data.
Example: decode.
- Metal
Apple’s way for programs to use the Mac’s GPU.
Example:
brew install llama.cppgets a Metal build, with no CUDA.- MiB
1,048,576 bytes (about a million), the unit llama.cpp prints; 1,000 MiB is about 1 GB.
Example: 8,704 MiB = 9.13 GB.
- Mixed precision
Training that does its math with 16-bit numbers while keeping 32-bit copies where accuracy matters.
Example: bf16 math with an fp32 master copy and fp32 optimizer averages: about 16 bytes per number.
- MLP (feed-forward block)
The other big part of each layer besides attention, made of three large weight grids (gate, up, down).
Example:
peft_params.py --mlpadds them to what LoRA adapts.- MLX
Apple’s own software for running and training models on Apple chips;
mlx_lm.serveris its model server.Example: port 8081; the same library fine-tunes a model on Day 8.
- mlx-lm-lora
A tool built on MLX that trains preference methods such as DPO on a Mac.
Example:
mlx_lm_lora.train --train-mode dpo.- Model id
The exact name a program uses to pick a model on Fireworks.
Example:
accounts/fireworks/models/gpt-oss-120b, the default in Day 5’s script.- Model retirement
The provider switching a model off; code that names it stops working.
Example: the lab book saw several models due to retire two days after it read the prices.
- MoE (mixture of experts)
A model split into many expert parts, of which only a few run for each token.
Example: gpt-oss:20b reads 3.6 of its 21 billion parameters per token.
- ms (millisecond)
A thousandth of a second.
Example: 500 ms is half a second.
- Multi-LoRA
One server keeps one base model and many adapters, applying the right adapter to each request.
Example: 12 business units on 1 GPU; up to 100 adapters by default on Fireworks.
- Multi-node (scale-out)
Running one model across more than one machine, to get more memory than one machine has.
Example: a 235-billion-parameter model on two Sparks at 11.7 tokens a second.
- MXFP4
A 4-bit number format for model weights, supported on Blackwell GPUs such as the Spark’s.
Example: gpt-oss 120B in MXFP4 is 59 GiB on disk.
- N-gram
A run of n tokens in a row.
Example: a 3-gram is 3 tokens.
- N-gram speculation
Guessing the next tokens by copying what followed the same words earlier in the prompt; no draft model.
Example:
--ngram-speculation-length=3on Fireworks.- nan (not a number)
What a column shows when it has no value.
Example:
judge_1to5when no judge was asked.- NCCL
NVIDIA’s library for moving data between GPUs, including all-reduce.
Example:
all_reduce_perftests the link between two Sparks in lesson 14’s TWO_BOX.md.- NGC
NVIDIA’s catalogue of ready-made containers for its GPUs.
Example: the vLLM image
1_vllm.shstarts on the Spark.- Notes
This course’s plain word for the KV cache.
Example: 128 KB of notes per token for Llama 3.1 8B.
- Number format
The way each of a model’s numbers is stored, including how many bits it takes.
Example: BF16 (16 bits), FP8 (8 bits), FP4 (4 bits).
- numpy
A Python library for fast math on lists of numbers.
Example: the lesson 16 toys need nothing else.
- NVFP4
NVIDIA’s 4-bit number format with built-in math on Blackwell.
Example: named in the second question of Day 4, step 1.
- NVIDIA
The company whose GPUs run most data-center AI.
Example: vLLM and SGLang need an NVIDIA GPU.
- NVLink
NVIDIA’s very fast direct link between GPUs inside one server.
Example: about 900 GB/s on H100 (field guide).
- Observability
Seeing what a live service is doing, such as speed, errors and memory, while it runs.
Example: llama.cpp’s
--metricsflag publishes live counters at/metrics.- Offload (-ngl)
Choosing how many of the model’s layers run on the GPU; the rest run on the slower CPU.
Example:
-ngl 99puts every layer on the GPU.- Ollama
A one-command model server built on llama.cpp that chooses most settings for you.
Example:
ollama serveanswers on port 11434.- On-demand deployment
Fireworks’ name for a dedicated deployment: a private copy of a model, billed by the second at an hourly rate.
Example: the only way Fireworks serves a trained LoRA.
- Open model (open weights)
A model whose weights anyone can download, run and fine-tune.
Example: gpt-oss-120b, which the Platform Map video calls a mid-sized open model.
- OpenAI-compatible API
The request format OpenAI made popular, which most model servers accept, so one client works with all of them.
Example: Ollama, llama.cpp, MLX, vLLM and SGLang all speak it.
- Optimizer state
The running averages the training method (Adam) keeps per trainable number to size each correction.
Example: part of the about 16 bytes per number that full fine-tuning needs.
- Order of magnitude
About 10 times.
Example: 740 ms down to 36 ms is more than one order of magnitude.
- ORPO (odds-ratio preference optimization)
A preference method that does example training and preference training in one run, with no reference model.
Example:
--loss-method ORPOon Fireworks.- Out of memory
The server cannot get the memory it needs. It fails at startup (weights, or a cache reserved up front) or under load, as many conversations’ notes grow.
Example: the 131,072 16-bit row on a 16 to 24 GB Mac.
- Output tokens
The tokens the model writes back; providers bill them per million, usually at a higher price than input.
Example: $0.60 against $0.15 per million for gpt-oss-120b.
- Overfitting
Learning the training examples by heart instead of the task, so new examples go worse.
Example: 97% on training examples but 74% on held-out ones in the QLoRA On Your Own Box video.
- p50 (median)
The middle value: half the requests are faster, half slower.
Example:
ttft_p50_msin the comparison table.- p95
The value 95 out of 100 requests stay under.
Example: 95 of 100 prompts under 3,000 tokens.
- p99
The value 99 out of 100 requests stay under.
Example:
P99 TTFTin vLLM’s own benchmark.- PagedAttention
vLLM’s way of handing out note memory in small pages as a conversation grows, instead of reserving the maximum.
Example: 2 to 4 times more users (The KV Cache video).
- Parameters (8B, 70B)
The model’s learned numbers (weights).
Example: 8B = 8 billion.
- Parse
Read text as data; parsing fails when the text is not valid JSON.
Example:
parses_%: 86% when only asked in the prompt, on the practice server.- PCIe
The standard slot connection between a computer’s main board and a plug-in card such as a GPU; far slower than the GPU’s own memory.
Example: why moving layers off the GPU is slow (lesson 12).
- Peak memory
The most memory a job uses at any moment.
Example: about 11 GB for the 8B fine-tune in the QLoRA On Your Own Box video.
- Peak traffic
The most people using the service at the same moment.
Example: 120.
- PEFT (parameter-efficient fine-tuning)
The family of methods that customise a model by training a tiny fraction of it.
Example: LoRA, QLoRA, DoRA, adapter layers, prompt tuning and IA3.
- Percentile
The value a given share of requests stay under: p50 is half of them, p95 is 95 of 100.
Example: p95 TTFT of 41.4 ms at 8 people on the practice server.
- perf_metrics
Timing and speculation figures Fireworks returns with a reply when the request asks for them (
perf_metrics_in_response).Example: the
fireworks perf_metrics sample:line in lesson 13.- pip, wheel
pip installs Python packages; a wheel is a ready-built package file.
Example: a prebuilt wheel can expect a different CUDA than the machine has.
- Pipeline (integration)
The customer’s code that sends requests and uses the replies automatically, with no person checking each one.
Example: a help desk that routes each ticket by its category.
- Pipeline bubble
Time a GPU in a pipeline sits idle, waiting for work from the stage before it.
Example: too few requests in flight leave gaps in the line.
- Policy
In reinforcement learning, the model being trained, seen as a rule for choosing answers.
Example: RLHF’s policy writes answers and the reward model scores them.
- Policy gradient (REINFORCE)
Learning by sampling answers and making the above-average ones more likely.
Example: the toy’s sampled RLHF run, 5b.
- Poll
Ask again at intervals whether a job has finished.
Example:
submit_batch.shpolls every 30 seconds.- Port
A numbered door on a computer where one server program listens for requests.
Example: Ollama 11434, llama.cpp 8080, MLX 8081.
- PP (pipeline parallelism)
Giving each GPU a block of consecutive layers, so each token passes from one GPU to the next; it tolerates slower links but adds waiting.
Example: PP=2 across two linked Sparks (TWO_BOX.md).
- pp512 / tg128
llama-bench’s two tests: read a 512-token prompt, write 128 tokens, each in tokens per second.
Example: 1,050 and 86 at 4-bit.
- PPO
The classic reinforcement algorithm in RLHF: sample answers, score them, make high scorers more likely, on a leash.
Example: it needs 4 models in memory.
- Practice server (mock)
The kit’s stand-in server with realistic behaviour and made-up numbers.
Example: it has 8 slots.
- Preamble
The long fixed opening sent with every request: instructions, tool list and rules.
Example: 3,028 tokens in Day 5’s script.
- Precision
How many bits each stored number gets.
Example: 16-bit, 8-bit, 4-bit.
- Predicted outputs
A Fireworks option where you send the expected answer, such as the original text of an edit, as the draft.
Example: sending the original document when asking for a small change.
- Preference pair
Two answers to the same prompt, one chosen and one rejected.
Example: 800 in lesson 18, one per training ticket.
- Preference tuning
Training a model on which answers people prefer, rather than on single correct answers.
Example: Day 12’s DPO step.
- Prefill
The reading phase; the model reads the whole prompt at once.
Example:
pp512.- Prefill/decode interference
Long prompts being read slow the replies being written for others on the same server.
Example: one of the things that moves the knee.
- Prefix
The start of a prompt, up to the first token that differs from an earlier one.
Example: the shared 3,028-token opening.
- Prefix caching (prompt caching)
The server keeps its notes on a prompt’s opening and reuses them when the next request starts the same way.
Example: 740 ms for call 1, 36 ms for the rest in Day 5’s sample.
- Projection (weight grid)
One grid of learned numbers that a layer multiplies its input by.
Example: 4,096 x 4,096 = 16.8 million numbers in an 8B model’s attention.
- Prompt
The text you send the model.
Example: a 3,000-token support question.
- Prompt cache file
A file holding the saved notes on a prompt, so a later run can skip reading it.
Example:
preamble.safetensorsin the MLX script.- Prompt masking (
--mask-prompt) Counting the training error only on the answer, not on the system message or the customer’s text.
Example:
run_variants.shpasses--mask-prompt.- Proof of concept (POC)
A small working version of the customer’s system, built to prove it meets their needs before they commit.
Example: slide 5: the next two weeks of a POC.
- Prose traffic
open-ended, creative requests whose next words are hard to guess.
Example: “Write a poem about latency in the voice of a tired sea captain.”
- Proxy
A stand-in measure used because the real goal is hard to measure directly.
Example: a reward model is a proxy for what the customer wants.
- PyTorch
The Python library for model math that vLLM and many other AI tools are built on.
Example: a PyTorch and CUDA version clash is why the lab book avoids
pip install vllmon the Spark.
- q8_0 (cache)
llama.cpp’s 8-bit format for the notes, 1.0625 bytes per number.
Example: 8,704 MiB instead of 16,384.
- Q8_0, Q4_K_M, Q3_K_M
llama.cpp’s weight formats at about 8.5, 4.8 and 3.9 bits per number.
Example: 8.5, 4.9 and 4.0 GB for an 8B model.
- QLoRA
LoRA on a base model stored in 4 bits, so training fits in far less memory.
Example: about 11 GB at the peak for an 8B model in the QLoRA On Your Own Box video.
- Quality cliff
The size below which answers suddenly get worse.
Example: the 3-bit version (Q3) slips on multi-step and format questions first.
- Quantization (quantize)
Storing numbers with fewer bits.
Example: 16 to 8 to 4 bits, like saving a photo at lower quality.
- Queueing
Waiting in line for a free seat on the server.
Example: at 16 people on the practice server, 8 wait.
- Quota
A limit Fireworks sets on your account, such as how many GPUs you may use.
Example:
firectl quota listshows them.- Qwen2.5 7B Instruct
An open model with 7.6 billion numbers, already trained to follow instructions; Days 7 and 8 fine-tune it.
Example:
qwen2p5-7b-instructon Fireworks.- Qwen2.5-0.5B-Instruct
A small open instruct model with about half a billion parameters, used in the training lessons.
Example:
mlx-community/Qwen2.5-0.5B-Instruct-bf16, about 1 GB.
- RadixAttention
SGLang’s prefix cache: it keeps saved notes in a tree of shared openings and orders requests to reuse them.
Example: the video’s trunk, branch and leaf picture.
- RAG (retrieval-augmented generation)
An app that first looks up relevant documents and puts them in the prompt, then asks the question.
Example: Day 1’s long test prompt puts the context first and the question last.
- RAM (memory)
The computer’s working memory, where the model must sit while it runs.
Example: 36 GB on the sample Mac in Day 1’s setup step.
- Rank (LoRA rank)
How wide the adapter’s small grids are; a higher rank means more capacity and a bigger file.
Example: 8 on Day 7, whose script suggests 16 to 32 for harder tasks.
- Reasoning model
A model that writes out its thinking before its final answer.
Example: lesson 07: put the schema in its prompt and validate afterwards.
- Reference model
A frozen copy of the starting model that DPO and RLHF measure drift against.
Example:
--reference-model-pathindpo_mlx.sh.- Region-pinned deployment
A dedicated deployment fixed to one region (US, EU or APAC), billed at 1.5 times the normal rate.
Example: $8,760 a month for one H100 instead of $5,840.
- Reinforcement learning
Training by trial and error: the model tries answers, gets a score, and makes high-scoring answers more likely.
Example: the PPO loop in RLHF.
- Replica
One complete running copy of the server and model.
Example: 18 replicas for 120 people.
- Request
One message sent to the model for an answer.
Example: 95 of 100 requests under 3,000 tokens.
- Reserved capacity
GPUs committed for a long term at a lower hourly price, billed whether used or not.
Example: the fourth capacity mode; the lab book warns never to suggest it casually.
- response_format
The request setting that asks the server to enforce JSON (
json_object) or a schema (json_schema).Example: left out in prompt-only mode.
- Reward hacking
A model raising its score by exploiting the scorer’s blind spots instead of getting better.
Example: in lesson 16’s toy the score rises to +1.96 while true value falls to +3.45.
- Reward model
A model trained on people’s comparisons to give any answer a single score.
Example: in lesson 16’s toy it agrees with its labellers 71% of the time.
- RFT (reinforcement fine-tuning)
Training against a grader that scores each answer from 0 to 1, for tasks where correctness can be checked automatically.
Example: the tests pass, the format validates, the number matches.
- RLAIF
RLHF with an AI judge in place of human labellers.
Example: cheaper labels (the RLHF video).
- RLHF (reinforcement learning from human feedback)
Training a model to chase a reward model’s score, learned from people’s comparisons, on a leash.
Example: Day 12’s first step runs it on a toy.
- Roofline
A chart of a chip’s top speed against arithmetic intensity: a sloped memory-bound part, then a flat compute-bound roof.
Example: the GPU Bandwidth video’s chart.
- Router (MoE)
The small part of a mixture-of-experts model that picks which experts handle each token.
Example: the receptionist in lesson 12’s picture.
- Routing
Sending each request to the right model, or to the right copy of it.
Example: one of the layers a managed platform runs for you.
- RPM (requests per minute)
How many requests an account may send each minute.
Example: 10 on a new Fireworks account until a payment method is added.
- Rubric
The scoring guide a judge follows.
Example: 5 = specific, right owner, safe; 3 = plausible but vague; 1 = wrong or unsafe.
- Sampling
Letting the model write answers at random, following its own probabilities.
Example: RLHF samples answers at every training step; DPO does not.
- Scale
One shared number stored with a small block of numbers so they can be squeezed into fewer bits.
Example: why q8_0 is 1.0625 bytes, not 1.
- Scale to zero
A deployment setting that gives the GPU back after a stretch with no requests, so billing stops.
Example: Fireworks on-demand deployments do it after 1 hour idle by default (the lab book).
- Scheduler
The part of a server that decides which requests share the chip at each step.
Example: vLLM’s scheduler swaps requests in and out at every step.
- Schema-valid rate
The share of replies that come back as correctly formatted data.
Example: 99% in the Quantization video’s sample.
- Semaphore
A counter in code that lets only a set number of tasks run at the same time.
Example:
asyncio.Semaphore(8)keeps at most 8 requests running at once.- Sequence length
How many tokens one training example holds.
Example: activations grow with batch size x sequence length.
- Server
A program that waits for requests and answers them.
Example: llama.cpp’s
llama-server.- Serverless
Fireworks’ shared, always-on models, billed per token you send and receive.
Example: Day 1’s step 3 measurements cost about 1 to 2 cents.
- Serving (serve)
Keeping a model running so it can answer requests.
Example: $5.07 for 38 minutes of a dedicated GPU in Day 7’s sample.
- Session
One conversation.
Example: an 8K session holds up to 8,192 tokens.
- Session affinity
Sending every request of one session to the same replica, so it finds its saved notes.
Example: the
uservalue in Day 5’s script,lesson-06-session.- Severity (severity_%)
How urgent a ticket is, from 1 to 5;
severity_%is the share rated correctly.Example: 70.0 before and 95.0 after the fine-tune in Day 7’s sample.
- SFT (supervised fine-tuning)
Training on examples of the right answer, to teach format, tone and domain patterns.
Example: 800 labelled tickets on Day 7.
- SGLang
A fast open-source server for NVIDIA GPUs that reuses shared prompt openings.
Example: it runs in Day 9’s optional Spark step.
- Sharding (ZeRO, FSDP)
Splitting a model’s weights, gradients and optimizer state across several GPUs so each holds a share.
Example: a 70B full fine-tune across 16 H100s.
- Sigmoid (σ)
A curve that turns any number into a value between 0 and 1; 0 gives 0.5.
Example: σ(2.5) = 0.92.
- Single-stream
One request at a time, with no one else sharing the server.
Example: 54.9 tokens/s for one person on the practice server.
- Sizing memo
A short document recommending how to run a customer’s workload: how many copies, which option, what it costs, with the maths shown.
Example: Day 10’s
results/SIZING_MEMO.mdquotes Day 5’s speed-up.- SLA (service level agreement)
A written promise to a customer about uptime or speed.
Example:
COMPARISON.mdsays the Mac has none.- SLO (service level objective)
The speed promise.
Example: 95 of 100 requests get their first token within 500 ms.
- Slot
One seat on the server, a conversation it can serve at the same time.
Example:
--parallel 8gives 8.- Slot cache reuse
llama.cpp keeps each seat’s notes from its last request and reuses the part that matches.
Example: why calls 2 to 12 are fast on llama.cpp.
- Smoke test
A quick run of everything to prove the setup works before real use.
Example:
make smoke, 16 lessons in about 2 minutes.- Speculative decoding
A small draft model guesses several tokens and the big model checks them all in one pass; the output does not change.
Example: Day 9 measures when it pays.
- Speed-up (speedup_x)
How many times faster: the slow time divided by the fast one.
Example: 740 ÷ 36 = about 20.
- Spend cap
A monthly limit on your Fireworks spending; at 100% Fireworks pauses the API.
Example: set with
firectl quota update monthly-spend-usd.- SSH
A secure way to log in to another computer over the network and run commands there.
Example:
remote.shruns each script on the Spark over SSH.- Stable first, variable last
Build the prompt with the parts that never change at the top and anything that changes at the bottom.
Example: rules and tools first, the time and the question last.
- Startup log
The lines a server prints as it starts, including how much memory it set aside.
Example: vLLM’s shows 8.4 GB of weights and 8,234 blocks.
- Static batching
A fixed group of requests runs together, and the next group starts only when the longest one finishes.
Example: 12 requests on 4 seats finish at step 24, against 17 with continuous batching.
- Structured traffic
Requests whose answers follow a predictable form, such as JSON data, lists, code or extraction.
Example: lesson 13’s 5 structured prompts.
- Synthetic data
Examples made by a program or a model instead of collected from real users.
Example: the kit’s tickets, from
data/make_tickets.py.- System prompt
The fixed instructions at the start of every request that set how the model behaves.
Example: “You are the support agent for Acme Cloud.”
- Tail (latency tail)
The slowest few requests.
Example: p95 describes the tail.
- Target model
The big model whose answer you want; it checks the draft model’s guesses and decides every token.
Example: Llama 3.1 8B in lesson 13.
- TB
1,000,000,000,000 bytes, or 1,000 GB.
Example: about 1.1 TB to fully fine-tune a 70B model.
- Teardown
Deleting everything that bills by the hour at the end of a session, then checking the list is empty.
Example:
5_teardown.sh, thenmake fw-check.- Tenant
One customer, or business unit, sharing a platform with others.
Example: per-tenant fine-tunes: one adapter each.
- Tensor cores
The parts of an NVIDIA GPU built for fast matrix math.
Example: the H100’s tensor cores do math directly on FP8 numbers.
- Terminal
The window where you type commands.
Example: Day 2 uses four at once: one per server, plus one to compare.
- Test set (held-out set)
Examples kept out of training and out of every decision, used only for the final score.
Example:
test.jsonl, 100 tickets; the eval uses 60.- Throughput
Total output across everyone at once.
Example: 275.1 tokens/s with 8 people.
- Ticket triage
Sorting incoming support tickets by type and urgency.
Example: the kit’s shared task for Days 6 to 8.
- TODO
A marker for work not done yet; the memo prints it where a number is missing.
Example:
TODO (run lesson 01).- Token
A chunk of text, about three quarters of a word.
Example: 1,000 tokens is roughly 750 words.
- Token ID
The number a tokenizer gives each token; models compare these numbers, not the text.
Example: the big model checks the draft’s token ids.
- Tokenizer
The part of a model that cuts text into tokens and numbers them.
Example: Llama 3.1 8B and Llama 3.2 1B share one.
- Tokens per second (tok/s)
How many tokens are written each second.
Example: 38.9 per person with 8 people.
- Tool (tool definition)
An action an agent may ask for, described in the prompt so the model knows how to call it.
Example: search_tickets, get_invoice, refund, status_page and escalate in Day 5’s script.
- Total parameters
Every parameter a model has; all of them must sit in memory, whether or not they are read for a token.
Example: 21 billion for gpt-oss:20b, a 13.8 GB file.
- Toy (toy model)
A training problem shrunk until every number fits on one screen, keeping the real algorithm.
Example: one 64 x 64 layer in
toy_finetune.py.- TP (tensor parallelism)
Splitting each layer of a model across several GPUs that work on every token together.
Example: TP=2 on two H100s joined by NVLink in one server; they sync inside every layer, so it needs very fast links.
- Traffic profile
The shape of the load a test uses: tokens in, tokens out, requests per second, and the speed promise with its percentile.
Example: 1,024 in, 256 out, 200 requests per crowd size.
- Training accuracy
How often the model is right on the examples it learned from.
Example: 97% for the messy run in the QLoRA On Your Own Box video.
- Training set
The examples the model learns from.
Example:
train.jsonl, 800 tickets.- Training token
One token read during training; managed training bills every token on every pass.
Example: 2,000 examples x 600 tokens x 2 passes = 2.4 million.
- Trap (exit trap)
A shell instruction that runs a clean-up command whenever a script ends, even on Ctrl+C.
Example:
fireworks_spec.shdeletes its deployment this way.- True value
In the toy, the hidden score of what people really want, which the fast labellers’ reward model cannot see.
Example: 4.775 for right-category JSON with a friendly line;
E[true]is its average over the model’s answers.- TTFT (time to first token)
How long until the first token (word piece) of the reply appears.
Example: 41.4 ms (p95) at 8 people.
- Turn
One call to the model within a session.
Example: 20 turns in the video’s agent session.
- Unified memory
One pool of memory shared by a Mac’s CPU and GPU.
Example: the model and your browser share the same 36 GB.
- Unit cost
The price of one unit of work, such as a million tokens written.
Example: $6.00 against $0.20 per million in the Platform Map example.
- Usage block
The token counts a server returns with an answer.
Example: prompt_tokens and cached_tokens on Fireworks.
- Utilisation
The share of time a GPU does useful work.
Example: a dedicated GPU costs the same whether it is busy or idle.
- Validate
Check a reply with code after it arrives.
Example: Day 6’s parse and schema checks.
- Validation loss
A score of how wrong the model still is on the validation set; lower is better.
Example: Day 8’s run prints it every 50 steps.
- Validation set
Examples kept out of training and checked during it, to decide when to stop.
Example:
valid.jsonl, 100 tickets.- Value model
In PPO, an extra model that predicts how good an answer in progress will turn out, to steady training.
Example: one of RLHF’s 4 models.
- Verification pass
One pass of the big model that checks all the draft’s guesses at once, the way it reads a prompt.
Example: 3.36 tokens per pass at α 0.8 with 4 guesses.
- Virtual environment (venv)
A private folder of Python packages for one project, turned on with
source .venv/bin/activate.Example:
mlx_lm.serverandfelablive inside it.- vLLM
An open-source engine for NVIDIA GPUs built for many users at once.
Example: it uses PagedAttention.
- Vocabulary
The full list of tokens a tokenizer knows, each with its own number.
Example: the draft and the big model must share one.
- VRAM
The memory built into a separate graphics card. A Mac has none: its shared memory plays that role.
Example: ‘RAM is your VRAM’, in Day 1’s setup step.
- Wall-clock time
The real time that passes from start to finish.
Example: total output divides all tokens by it.
- Warm-up request
One request sent first and never counted, so loading the model does not skew the timings.
Example: Day 1’s stopwatch sends ‘Say hi.’ first.
- Weights
The model’s learned numbers, fetched in full for every token written (a mixture-of-experts model fetches only the part it uses).
Example: 4.9 GB for Llama 3.1 8B at 4-bit.
- Win rate
How often people or a judge prefer the new model’s answer to the old one’s.
Example: the DPO video says to measure it together with task accuracy.
- Workload
One kind of job a customer runs.
Example: a 4K support chat (lesson 05).