Skip to content

Glossary

404 terms from the course, each with an example.

.env file

The settings file in the 03-labs folder, holding defaults and your Fireworks key; it never leaves your Mac.

Example: cp .env.example .env creates it in Day 1’s first step, with a placeholder key.

4-bit, 8-bit, 16-bit

How many bits store each number; fewer bits, smaller model.

Example: an 8B model is 8.5 GB at 8-bit, 4.9 GB at 4-bit.

β (beta)

The leash in RLHF and DPO: how far training may pull the model from its starting copy; smaller means a longer leash.

Example: lesson 16’s toy loosens it from 0.50 to 0.02.

Acceptance rate

The share of a draft model’s guessed tokens the big model keeps; speculative decoding only pays when it is high.

Example: α (alpha) = 0.78 on structured prompts and 0.41 on prose in lesson 13’s sample.

Accuracy (category accuracy)

The share of replies with the right answer, checked against each test example’s known label.

Example: 78% of 50 tickets on the practice server (category_acc_%).

Activations

The working numbers a model computes for each example during training, kept for the correction step; they grow with batch size and example length.

Example: they come on top of the 16 bytes per number.

Active parameters

The numbers a mixture-of-experts model actually uses for each token, as opposed to all the numbers it stores.

Example: 3.6 of gpt-oss:20b’s 21 billion (17%).

Adam

The usual training method for adjusting a model’s numbers; it keeps two running averages per trainable number.

Example: its two averages are part of the 16 bytes per number a full fine-tune needs.

Adapter (LoRA adapter)

The small set of extra numbers LoRA trains and saves; the base model stays unchanged.

Example: about 30 MB in the Training Techniques video.

Adaptive speculative decoding

Fireworks’ option that trains the draft model on a customer’s own traffic, so more guesses are kept.

Example: the Batching & Scheduling video: the drafter trained on each customer’s own traffic.

Agent

A program that calls a model many times, using tools, to finish a task.

Example: Day 5’s Acme Cloud support agent, with 5 tools.

All-reduce

A network step where every GPU combines its partial results with all the others, so each ends up with the total.

Example: tensor parallelism does this inside every layer.

All-to-all

A network step where every GPU sends a different piece of data to every other GPU.

Example: moving each token to the GPUs that hold its experts.

API

The agreed format one program uses to send requests to another and get answers back.

Example: /v1/chat/completions is where a chat request goes.

API key

A secret code that proves a request comes from your account, so it can be billed.

Example: FIREWORKS_API_KEY in .env.

Arithmetic intensity

How many calculations the chip does for each byte it fetches from memory.

Example: about 1 to 2 for one person’s decode; an H100 needs about 300 to keep its math busy.

arm64

The chip design used by Apple silicon and by the DGX Spark’s processor; software must be built for it.

Example: container images tagged for GB10/arm64.

Attention (the look-back step)

The step where each new token looks back at the notes on earlier tokens.

Example: each writing step reads every earlier token’s notes, which is why a long prompt slows it a little.

Attention head (query head)

One of the parallel parts of a layer that looks back over the conversation.

Example: Llama 3.1 8B has 32 per layer; with GQA they share 8 note sets.

Auto Reload

Fireworks’ setting that tops up your prepaid credit automatically when it runs low.

Example: leave it off while you are learning.

Automatic prefix caching

vLLM’s built-in prefix cache: it reuses saved notes for identical blocks at the start of a prompt.

Example: vllm serve $MODEL --enable-prefix-caching.

Autoscaling

Adding or removing copies of a model as traffic rises and falls.

Example: one of the layers a managed platform runs for you.

B200

NVIDIA’s Blackwell data-center GPU, 192 GB of memory at 8,000 GB/s.

Example: 8,000 ÷ 70 = 114 tokens/s limit for a 70B model at 8 bits.

Back-fill

Running a job over all your existing records at once, for example after adding a new field.

Example: re-sorting last year’s tickets into categories.

Bandwidth (memory bandwidth)

How many GB per second memory delivers to the chip.

Example: M4 Max 546 GB/s, H100 3,350 GB/s.

Base model

The model a fine-tune starts from and leaves unchanged.

Example: Qwen2.5 7B Instruct on Day 7.

Batch API

Fireworks’ way to send many requests as one file and collect the answers later, at about half the serverless price.

Example: Day 6’s 100 tickets in batch_input.jsonl.

Batch job (batch-inference job)

One run of the Batch API: your uploaded file of requests, processed when GPUs are free.

Example: triage-batch-…, checked every 30 seconds until COMPLETED.

Batch size

How many requests share one pass over the model.

Example: 20.5 tokens a second at batch 1 and 368 in total at batch 32 on a Spark (lab book).

Batch size (training)

How many examples the model learns from in one training step.

Example: 2 in 1_train_mlx.sh; lower it to 1 if memory runs out.

Batching

Serving several people with one pass over the model.

Example: 8 people share each fetch of the weights.

Benchmark (public benchmark)

A standard test that model makers report scores on, used to compare models at launch.

Example: it says little about one customer’s tickets.

bf16 (bfloat16)

A 16-bit number format used for training, 2 bytes per number.

Example: multi-LoRA on Fireworks needs a BF16 deployment shape.

Bias (judge bias)

A judge’s habit of favouring something that is not correctness.

Example: longer, more polished answers.

Bit

A single 0 or 1; 8 bits make a byte.

Example: a 16-bit number is 2 bytes.

Blackwell

The NVIDIA GPU generation after H100, with built-in 4-bit math.

Example: B200, 8,000 GB/s.

Block (KV block)

A fixed-size piece of note memory holding the notes for a few tokens; vLLM hands blocks out as conversations grow.

Example: 16 tokens per block; 8,234 blocks in the Serving With vLLM startup log.

Block hash

A fingerprint of one fixed-size block of a prompt together with everything before it; a matching fingerprint means the saved notes can be reused.

Example: 64-token blocks in the practice server.

Bottleneck

The slowest part of a process, the one that holds everything up.

Example: memory speed while the model writes.

Bradley-Terry model

The rule that turns two scores into the chance a person prefers one: the sigmoid of the gap between them.

Example: scores 2.1 and -0.4 give about 92%.

Break-even

The point where two options cost the same.

Example: about 800 million tokens a day for serverless against one dedicated H100, in the Platform Map example.

Build versus buy

Doing a job with your own hardware and people, or paying a service to do it.

Example: COMPARISON.md: the same fine-tune on your Mac and on Fireworks.

Bursty traffic

Traffic that comes in sudden peaks with quiet spells between.

Example: the Platform Map video matches it to serverless.

Business case

The argument, in money and results, for making a change.

Example: 88% off the input bill for an agent whose input is 90% repeats, at the video’s prices.

BYOC (bring your own cloud)

A managed service’s software running inside the customer’s own cloud account, so the data stays there.

Example: the lab book’s answer to “Can we keep data in-region?”

Byte

8 bits.

Example: one FP16 number takes 2 bytes.

Cache hit (hit rate)

A request that finds its opening already saved; the hit rate is the share of requests that do.

Example: the Prefix Caching video: measure the hit rate, because an unused cache only takes memory.

Cached input

Input tokens the server reused from an earlier request, billed at a discount.

Example: $0.015 instead of $0.15 per million for gpt-oss 120B.

Cached tokens (cached_tokens)

Input tokens the server reused instead of reading again; Fireworks reports them and bills them at a discount.

Example: usage: prompt_tokens=… cached_tokens=… from Day 5’s --usage run.

Capacity (of one replica)

The most people one copy of the server serves at once while keeping the speed promise.

Example: 8 on the practice server within 500 ms.

Capacity modes

The ways a platform sells computing: serverless, dedicated, batch and reserved.

Example: the four in the Platform Map video.

Ceiling (speed limit)

The fastest a model can write: memory speed ÷ model size (for a mixture-of-experts model, only the part it reads per token).

Example: 546 ÷ 4.9 = 111 tokens/s.

Changelog

A provider’s dated list of changes to its product.

Example: where the lab book found adaptive rate limits described.

Chatty answer (chatty_%)

A reply that wraps the JSON in chat or writes prose, so a program cannot read it.

Example: about 20% before DPO, about 1% after, as the lesson expects.

Checkpoint (saved model file)

The file a training run saves: the whole changed model for a full fine-tune, only the adapter for LoRA.

Example: about 990 MB against about 9 MB in lesson 17’s expected table.

Chip

The processor that does the model’s math.

Example: an Apple M4 Max, or an NVIDIA H100 (a GPU).

Chosen and rejected

The two answers in a preference pair: the one a person approved, and the one they edited away or turned down.

Example: chosen: the clean JSON; rejected: the same JSON wrapped in chat.

Chunked prefill

Reading a long prompt in pieces, between other people’s writing steps, so one long prompt does not freeze everyone else.

Example: --enable-chunked-prefill in vLLM.

Closed model

A model you can only rent through its maker’s service; you never get its weights.

Example: where the Platform Map video’s customer journey starts.

Cluster

Several GPUs, often across machines, working as one.

Example: what a full fine-tune of an 8B model needs, at about 128 GB.

Code completion

An editor feature that suggests the next few lines of code.

Example: 200 tokens in, 30 out, in the Inference 101 video.

Cold and warm

Cold: nothing is saved yet, so the whole prompt is read; warm: saved notes are reused. Not the same as a cold start (Day 1), when a server is still loading the model.

Example: call 1 is cold (740 ms) and calls 2 to 12 are warm (36 ms) in Day 5’s sample.

Cold start

A slow request while a server that was idle or newly started loads the model.

Example: why Day 1’s stopwatch never counts its first call.

Completion

The text the model writes back.

Example: the JSON reply to one ticket.

Compute-bound

Slowed by calculation speed rather than memory.

Example: prefill.

Concurrency (people served at once)

How many conversations the server runs at the same moment.

Example: 8 on the practice server.

Concurrency sweep

The same test run at 1, 2, 4, 8, 16 and 32 people at once, then drawn as a curve.

Example: lesson 03’s sweep.py.

Config (settings file)

The file shipped with a model that lists its shape.

Example: 32 layers, 8 KV heads, head dim 128.

Constrained decoding

At each step the server blocks every token that would break the schema or grammar, so the reply cannot come out malformed.

Example: json_schema mode: 50 of 50 valid in lesson 07’s practice run. A reply cut off by the length limit (max_tokens) can still break.

Container

A sealed package holding a program and the exact software versions it needs, so it runs the same on any machine.

Example: Serving With vLLM starts the vllm/vllm-openai:latest container with docker run (Docker is the usual container tool).

Container image (tag)

The packaged software a container starts from; the tag names its version.

Example: nvcr.io/nvidia/vllm:25.09-py3 in 1_vllm.sh.

Context length (context, ctx)

The most tokens one conversation may hold.

Example: 8K = 8,192 tokens.

Continuous batching

Letting requests join and leave the shared group at every writing step, instead of waiting for the whole group to finish.

Example: a finished reply’s seat is refilled at the next step, like a taxi rank rather than a tour bus.

CPU

The computer’s general-purpose processor, slower than the GPU for model math.

Example: model layers not placed on the GPU run here.

Crossover

The monthly volume at which two ways of serving cost the same; the memo’s name for the break-even.

Example: about 744 million tickets a month for the memo’s example customer, if its current 18 copies could carry that much.

CSV file

A plain table file: one row per line, commas between the columns.

Example: results/01-latency.csv.

CUDA

NVIDIA’s software layer that lets programs run on its GPUs.

Example: a Mac has none; llama.cpp uses Metal instead.

CUDA graphs

A GPU speed-up that records a sequence of work once and replays it, instead of launching each piece separately.

Example: the documented no-container workaround gives it up and loses 20 to 30% of throughput.

custom_id

The label on each request in a batch file, used to match every answer back to its question.

Example: t-000 to t-099.

Customer journey (maturity curve)

The usual path: rent a closed model, improve the prompt and context, move to open models, then train on your own data.

Example: the Platform Map video’s four stages.

Dataset (Fireworks dataset)

A file stored in your Fireworks account for jobs to read or write.

Example: firectl dataset create uploads the batch file; the job writes its results to another.

Decode

The writing phase, one token at a time, fetching the whole model for each (a mixture-of-experts model fetches only the part it uses).

Example: 86 tokens/s at 4-bit.

Dedicated deployment

A GPU reserved for you on Fireworks, billed by the second at an hourly rate from the moment it is created, used or not.

Example: about $8 an hour for one H100 (lab book); by default it stops only after a full hour with no requests.

Dense model

A model that uses all of its numbers for every token.

Example: Llama 3.1 8B.

Deployment (production)

A model running as a live service for real users.

Example: running out of note memory is the most common production failure.

Deployment shape

A preset for a dedicated deployment on Fireworks: the GPU and the number format it runs.

Example: --deployment-shape default in 3_deploy.sh.

DGX Spark

NVIDIA’s desktop AI computer, with 128 GB of memory shared by its CPU and GPU.

Example: about 273 GB/s of memory speed; Day 9’s optional step runs vLLM and SGLang on it.

Discovery call

The first meeting where you learn a customer’s task, volume, speed target, quality bar and data limits.

Example: the memo is what you send after it.

Docker

The common tool for building and running containers.

Example: Serving With vLLM starts vLLM with docker run.

Document classification

Sorting documents into a fixed set of categories.

Example: ticket triage is one kind.

DoRA

A LoRA variant that splits each weight into a size (magnitude) and a direction, and trains LoRA on the direction.

Example: often 1 to 2 points above LoRA, a little slower (the Parameter-Efficient Fine-Tuning video).

DPO (direct preference optimization)

Training from pairs of answers where one is marked better, to teach taste or judgement.

Example: Day 12 runs it on your Mac.

Draft model

A small, fast model that guesses the next few tokens for a big model to check.

Example: Llama 3.2 1B drafting for Llama 3.1 8B in lesson 13.

Driver

The software that lets the operating system use a piece of hardware such as a GPU.

Example: the driver, CUDA and the Python packages must agree on an NVIDIA machine.

Efficiency %

Measured writing speed ÷ ceiling.

Example: 86 ÷ 111.4 = 77%.

Embedding

A list of numbers that captures a text’s meaning, used for search and matching.

Example: recomputing them for every document is a typical batch back-fill.

Endpoint

A web address that answers requests for one model.

Example: http://localhost:8081/v1 for Day 8’s fine-tune.

Enforce

make the server block anything that breaks the rules.

Example: json_schema mode enforces the schema.

Engine (inference server)

The program that loads a model and answers requests.

Example: llama.cpp, vLLM.

Enum

A field that may hold only one value from a fixed list.

Example: category: billing, outage, how-to or abuse.

Epoch

One full pass through all the training examples.

Example: 2 epochs over 800 tickets on Day 7, so every token is billed twice.

Error file

Where a batch job puts the requests that failed, apart from the results.

Example: why answered can be below 100.

Eval (eval harness)

A repeatable test of a model on a fixed set of prompts.

Example: 50 of the customer’s prompts in lesson 09 (Day 7).

Expert

One specialist block inside a mixture-of-experts model; only the few the router picks are read for a token.

Example: lesson 12’s picture: 2 to 4 specialists out of dozens per token.

Expert labels

Preference choices made by people who can check correctness.

Example: toy_alignment.py --expert-labels: the reward model then agrees with them 78% of the time.

Expert parallelism

Placing different experts on different GPUs, so tokens travel to whichever GPU holds their chosen experts.

Example: the field guide’s 4 experts on 4 GPUs.

f16 / FP16

A 16-bit number format, 2 bytes per number.

Example: the default format for the notes.

Fabric (interconnect)

The links that join GPUs or whole machines so they can work on one model.

Example: two DGX Sparks joined by a 200-gigabit cable.

felab

The kit’s small shared Python toolkit that every lesson script imports.

Example: felab/targets.py lists every server’s address.

Fetch

Move data from memory to the chip that does the math.

Example: the whole 4.9 GB model is fetched once for every token written.

Few-shot examples

A few solved examples placed in the prompt to show the model what a good answer looks like.

Example: tickets shown with their correct JSON before the real one.

Fine-tune

Further training of an existing model on your own examples.

Example: a LoRA fine-tune on Day 7.

firectl

Fireworks’ command-line tool for your account, quotas and deployments.

Example: firectl deployment list shows what is billing by the hour.

Fireworks

A cloud service that runs AI models and charges per use; one platform used in this course.

Example: Day 1’s step 3 times its models for about 1 to 2 cents; one of the three, gpt-oss-20b, is expected to fail.

Flag

An option typed after a command that changes how the program runs.

Example: --parallel 4 gives llama.cpp 4 slots.

Flash attention

A faster, memory-saving way to run attention; llama.cpp needs it to store the values in 8 bits.

Example: EXTRA="-fa on".

Floating-point

A way of storing numbers with a sign, a size and a precise part, like scientific notation.

Example: FP16 and FP8.

FLOPS (floating-point operations per second)

How many calculations on decimal numbers a chip can do each second; a teraflop is a trillion.

Example: an H100 does about 989 trillion (989 teraflops).

Forgetting

Getting better at the trained task and worse at everything else.

Example: a drop in the 12-question general quiz, usually in the full row.

FP4

A 4-bit number format, half a byte per number, with built-in math on the newest NVIDIA GPUs.

Example: FP4 deployment shapes cannot host LoRA adapters.

FP8

An 8-bit number format, 1 byte per number, with built-in math on H100.

Example: a 70B model is 70 GB in FP8.

fp32

A 32-bit number format, 4 bytes per number; training keeps the optimizer’s averages and the master copy in it.

Example: 12 of the 16 bytes per number in a full fine-tune.

Fragmentation

Memory lost in gaps too small to use.

Example: part of the about 10% on top of weights and notes.

Frontier model

One of the largest, most capable models, usually priced highest.

Example: $6.00 per million output tokens in the Platform Map example.

Frozen

Left unchanged during training.

Example: QLoRA keeps the 4-bit model frozen.

Full fine-tune

fine-tuning that changes every number in the model.

Example: about 121.6 GB of memory for a 7.6B model, by Day 8’s memory script.

Fuse (fused model)

Baking the adapter into the base model’s numbers to make one self-contained model.

Example: the step after QLoRA training on Day 8.

Gated model

A model you must request access to on Hugging Face before you can download it.

Example: why lesson 14 asks for HF_TOKEN.

GB

1,000,000,000 bytes.

Example: 16,384 MiB = 17.18 GB.

GB10

The DGX Spark’s chip: an NVIDIA Grace CPU and Blackwell GPU in one package, arm64, with 128 GB of shared memory.

Example: about 273 GB/s of memory speed.

General quiz

Day 4’s 12 quick questions (math, facts, format), reused in lesson 17 to spot forgetting.

Example: about 9 of 12 for the untrained model.

Generalise

Do well on new examples, not only the ones trained on.

Example: 89% on held-out examples for the clean run in the QLoRA On Your Own Box video.

GGUF

llama.cpp’s model file format, one file per size.

Example: a Q4_K_M.gguf file.

GiB

1,073,741,824 bytes (1,024 MiB), a little more than a GB.

Example: 59 GiB = 63.4 GB.

Gigabit (Gb)

A billion bits; 8 gigabits make 1 GB.

Example: a 200-gigabit link moves about 25 GB a second.

Goodhart’s law

When a measure becomes the target, it stops being a good measure.

Example: the toy’s average rating rises while the average true value falls.

Goodput

The requests per second that met the speed promise; late ones do not count.

Example: 2.5 per second at 8 people on the practice server, 0.4 at 16.

gpt-oss-20b, gpt-oss-120b

Two open mixture-of-experts models with about 21 and 117 billion parameters, of which only about 3.6 and 5.1 billion are used for each token. On Fireworks, gpt-oss-120b is billed per token (serverless); gpt-oss-20b no longer is (its model page, checked 27 September 2026: ‘Serverless: Not supported’), but batch jobs still take it.

Example: Day 6 runs gpt-oss-120b in step 1 and gpt-oss-20b for step 2’s batch job.

GPU

The chip that runs the model.

Example: H100; on a Mac it shares memory with everything else.

GPU cloud (raw GPU cloud)

A service that rents bare GPUs; the customer runs everything above them.

Example: the bottom layer of the Platform Map.

GPU cores

The many small calculating units inside a GPU.

Example: on a Mac they mostly wait for memory while the model writes.

GPU memory utilization (--gpu-memory-utilization)

The share of the GPU’s memory vLLM may claim for the model and its notes.

Example: 0.8 means 80% of the card.

GPU-hour

One GPU rented for one hour.

Example: $8 for an H100; a month is about 730 of them.

GQA (grouped-query attention)

Groups of attention heads share one set of notes.

Example: 8 KV heads instead of 32.

Grader

Anything that scores a model’s reply: a code check, a judge model or a person.

Example: Day 7’s eval uses three, cheapest first.

Gradient

The correction worked out for each trainable number at each training step.

Example: full fine-tuning keeps one for every number of the model.

Gradient checkpointing (--grad-checkpoint)

Recomputing activations during training instead of storing them: less memory, somewhat slower.

Example: about 30% slower (the Full Fine-Tuning video).

Grammar

A precise rulebook for which text is allowed, which a server can enforce one token at a time.

Example: llama.cpp turns a JSON schema into one.

GRPO

A reinforcement method that scores a group of answers to the same prompt against each other, with no value model.

Example: used for many reasoning models (the RLHF video).

H100

NVIDIA’s data-center GPU, 80 GB of memory at 3,350 GB/s.

Example: about $8 an hour on a dedicated Fireworks deployment (lab book).

H200

The H100’s successor with faster memory.

Example: 4,800 GB/s against 3,350.

Hand-labelled sample

Answers a person has marked right or wrong, used to check a judge.

Example: about 20 of the judge’s scores, redone by you.

Harness

One test script you can point at any server, so the numbers compare fairly.

Example: bench_ttft.py, built on Day 1 and reused on Day 2.

Head dim

How many numbers each note set holds per token.

Example: 128 in Llama 3.1 8B.

Headroom

Spare capacity kept for surprises.

Example: plan 18 replicas where 15 are needed.

Health endpoint

An address that answers once the server is ready to take requests.

Example: curl -sf localhost:8000/health.

Held-out accuracy

How often the model is right on examples set aside before training and never trained on.

Example: 74% for the messy run, 89% for the clean run in the QLoRA On Your Own Box video.

Held-out error

How wrong a trained model is on new examples it never saw; lower is better.

Example: 0.1321 for the toy’s rank-1 add-on at 200 examples.

Held-out examples

Examples set aside before training and never trained on.

Example: test.jsonl, 100 tickets; the toy tests on 500 new inputs.

Held-out prompts

Test prompts the model was never trained on.

Example: lesson 09’s test.jsonl (Day 7).

HF_TOKEN (Hugging Face token)

A personal key that lets scripts download the models your Hugging Face account may use.

Example: set in .env; remote.sh passes it to the Spark.

Homebrew (brew)

A Mac tool that installs programs from the command line.

Example: brew install llama.cpp.

Hot-swap

Change which adapter a running server applies, without restarting it.

Example: what small adapter files make easy in multi-LoRA serving.

Hugging Face

The public website where most open model files are shared.

Example: serve_llamacpp.sh downloads its 4.9 GB file from there.

Imitation model (SFT model)

The model after it copied examples of good answers.

Example: the toy’s step 2 model, which the preference pairs are drawn from.

Inference

Running a trained model to get answers, as opposed to training it.

Example: an inference server answers requests.

Input tokens

The tokens you send to the model; providers bill them per million.

Example: $0.30 per million, or $0.006 cached, in the Prefix Caching video.

Instruct model

A model already trained to follow instructions and chat.

Example: Qwen2.5 7B Instruct.

Integer

A whole number.

Example: q8_0 stores notes as 8-bit integers plus a shared scale.

Interactive traffic

Requests where a person is waiting for the answer.

Example: a chat window; it stays on serverless or dedicated, never batch.

Iteration (training step)

One update: the model sees a small batch of examples and adjusts.

Example: 300 in the QLoRA On Your Own Box video’s Mac example.

ITL (inter-token latency)

The gap between one token and the next while a reply is being written.

Example: 22.4 ms in Day 1’s sample Ollama row: 1,000 ÷ 22.4 = 44.6 tokens per second.

JSON

A plain-text format for data that programs can read: named fields inside curly brackets.

Example: {"category": "billing", "severity": 3}.

JSON schema (schema)

A rulebook for JSON: which fields must appear, what type each holds and which values are allowed.

Example: data/triage_schema.json: 3 fields, 4 categories, severity 1 to 5.

json_object, json_schema

The two enforced modes: json_object guarantees valid JSON with any fields; json_schema guarantees JSON that follows your schema.

Example: both 100% valid in lesson 07’s practice run.

JSONL

A file with one JSON record per line.

Example: batch_input.jsonl, 100 lines.

Judge (LLM judge)

A second model that scores answers against a rubric.

Example: one column of lesson 09’s table (Day 7).

k (draft length)

How many tokens the draft model guesses per round.

Example: up to 8 in lesson 13’s llama.cpp server, 4 on Fireworks.

K (in 8K, 32K, 128K)

1,024 when counting tokens.

Example: 128K = 131,072 tokens.

KB, KiB

1,024 bytes in this course (the videos write KB, kv_calc.py prints KiB).

Example: 128 KB = 131,072 bytes.

Kernel

A small program that runs one piece of the model’s math on the GPU.

Example: running the model with fast kernels is a serving engine’s fourth job.

Keys and values (K and V)

The two kinds of notes the model stores per token, per layer.

Example: the 2 in 2 x layers x KV heads x head dim x bytes.

KL (Kullback-Leibler divergence)

A measure of how far one model’s choices have drifted from another’s; 0 means identical.

Example: the KL leash charges the model for that drift (previewed on Day 11, run on Day 12).

KL leash (KL penalty)

A charge for drifting from the starting model, added to the reward so training cannot run off after loopholes.

Example: lesson 16’s toy loosens it from β 0.50 to 0.02.

Knee

The point where seats run out and waiting time shoots up.

Example: after 8 people on the practice server.

KV cache

The model’s notes on the conversation so far (keys and values), so it does not reread everything for each new word.

Example: 17.18 GB for one 131,072-token conversation.

KV cache dtype

The number format the notes are stored in; 8-bit instead of 16-bit halves their memory.

Example: --kv-cache-dtype fp8 in vLLM.

KV head (note set)

One set of notes kept in each layer.

Example: 8 per layer in Llama 3.1 8B.

Label

The known right answer stored with each test example.

Example: each test ticket’s category and severity.

Labeller

A person who picks the better of two answers to make training data.

Example: the toy’s fast labellers cannot check the ticket’s category.

Latency

How long someone waits.

Example: time to the first token.

Latency-bound

A workload where what matters is how fast each reply arrives, typical of one or a few people at a time.

Example: where speculative decoding helps most.

Layer

One of the stacked processing stages each token passes through; each keeps its own notes.

Example: 32 in an 8B model, 80 in a 70B.

Leak (data leakage)

Data used to judge a model has also shaped it, so the score looks better than it is.

Example: the validation tickets, once they set the stopping point.

Learning rate

How big a nudge each training step gives each number.

Example: 1e-4 (0.0001) for LoRA, 1e-5 for full.

Lever

A change that moves cost or speed without new hardware.

Example: the memo’s five: caching, constrained output, batch, a LoRA fine-tune, speculative decoding.

Llama 3.1 8B, Llama 3 70B

The open models used in the lessons and videos, with 8 and 70 billion parameters.

Example: 4.9 GB at 4-bit for the 8B.

Llama 3.1 405B

A 405-billion-parameter Llama model with 126 layers (field guide).

Example: 810 GB of weights at 16 bits.

llama.cpp

The open-source engine used on the Mac; llama-server answers requests, llama-bench times the engine alone.

Example: Day 2’s server.

Load generator

A tool that sends many requests at set rates to test a server.

Example: vllm bench serve in 2b_vllm_bench.sh.

Localhost

Your own computer, used as an address.

Example: http://localhost:8080/v1 is llama.cpp on your Mac.

Log-probability (log-prob)

The natural log of how likely a model is to write a given answer; DPO and pref_eval.py --pairwise work in these.

Example: the chosen answer +0.8 against the reference in the DPO video, about 2.2 times as likely.

LoRA

A cheap way to customise a model: train a small add-on instead of changing the whole model.

Example: on Fireworks it can only be served on a dedicated deployment.

Loss (training loss)

A score of how wrong the model’s guesses are on the examples it trains on; lower is better.

Example: falling loss with flat test accuracy means memorising.

M4 Max, M4 Pro

Apple chips whose memory moves 546 and 273 GB/s.

Example: 546 ÷ 4.9 = 111 tokens/s on an M4 Max; 273 ÷ 4.9 = about 56 on an M4 Pro.

make

A tool that runs named shortcuts listed in a file called Makefile.

Example: make smoke, make fw-check.

Managed platform

A service that runs everything from the GPUs up to the model: the engine, scaling, routing and the models.

Example: Fireworks; its serverless models charge per token instead of per GPU-hour.

Managed service

A provider runs the hardware and software for you and bills for use.

Example: Fireworks training and serving Day 7’s fine-tune.

Margin

In DPO, how much more the model favours the chosen answer than the rejected one, in log-probability.

Example: 1.2 in the DPO video (against the frozen copy); pref_eval.py --pairwise reports the raw gap, about +1 before and +10 after.

Markdown

Plain text with light formatting marks, such as # for a heading; the files end in .md.

Example: SIZING_MEMO.md.

Mask (masking)

Block a token so the model cannot choose it.

Example: constrained decoding masks every token that would break the schema.

Master copy

A 32-bit copy of every weight kept during training, so many tiny nudges add up accurately.

Example: 4 of the 16 bytes per number in a full fine-tune.

matplotlib

The Python library that draws the kit’s charts.

Example: without it, plot_sweep.py prints a text chart instead.

Matrix math

Multiplying large grids of numbers, the core work of a model.

Example: what tensor cores speed up.

Max model length (--max-model-len)

vLLM’s limit on how many tokens one conversation may hold, which caps that conversation’s notes.

Example: --max-model-len 8192.

Maximum concurrency (vLLM log)

vLLM’s startup estimate of how many full-length conversations fit in its memory for notes: cache tokens ÷ the length cap.

Example: 8,234 blocks × 16 tokens ÷ 8,192 = about 16 in the Serving With vLLM video.

MB

1,000,000 bytes.

Example: about 1.6 MB of notes for a 12-token prompt on Llama 3.1 8B.

Memory-bound (bandwidth-bound)

Slowed by how fast memory delivers data.

Example: decode.

Metal

Apple’s way for programs to use the Mac’s GPU.

Example: brew install llama.cpp gets a Metal build, with no CUDA.

MiB

1,048,576 bytes (about a million), the unit llama.cpp prints; 1,000 MiB is about 1 GB.

Example: 8,704 MiB = 9.13 GB.

Mixed precision

Training that does its math with 16-bit numbers while keeping 32-bit copies where accuracy matters.

Example: bf16 math with an fp32 master copy and fp32 optimizer averages: about 16 bytes per number.

MLP (feed-forward block)

The other big part of each layer besides attention, made of three large weight grids (gate, up, down).

Example: peft_params.py --mlp adds them to what LoRA adapts.

MLX

Apple’s own software for running and training models on Apple chips; mlx_lm.server is its model server.

Example: port 8081; the same library fine-tunes a model on Day 8.

mlx-lm-lora

A tool built on MLX that trains preference methods such as DPO on a Mac.

Example: mlx_lm_lora.train --train-mode dpo.

Model id

The exact name a program uses to pick a model on Fireworks.

Example: accounts/fireworks/models/gpt-oss-120b, the default in Day 5’s script.

Model retirement

The provider switching a model off; code that names it stops working.

Example: the lab book saw several models due to retire two days after it read the prices.

MoE (mixture of experts)

A model split into many expert parts, of which only a few run for each token.

Example: gpt-oss:20b reads 3.6 of its 21 billion parameters per token.

ms (millisecond)

A thousandth of a second.

Example: 500 ms is half a second.

Multi-LoRA

One server keeps one base model and many adapters, applying the right adapter to each request.

Example: 12 business units on 1 GPU; up to 100 adapters by default on Fireworks.

Multi-node (scale-out)

Running one model across more than one machine, to get more memory than one machine has.

Example: a 235-billion-parameter model on two Sparks at 11.7 tokens a second.

MXFP4

A 4-bit number format for model weights, supported on Blackwell GPUs such as the Spark’s.

Example: gpt-oss 120B in MXFP4 is 59 GiB on disk.

N-gram

A run of n tokens in a row.

Example: a 3-gram is 3 tokens.

N-gram speculation

Guessing the next tokens by copying what followed the same words earlier in the prompt; no draft model.

Example: --ngram-speculation-length=3 on Fireworks.

nan (not a number)

What a column shows when it has no value.

Example: judge_1to5 when no judge was asked.

NCCL

NVIDIA’s library for moving data between GPUs, including all-reduce.

Example: all_reduce_perf tests the link between two Sparks in lesson 14’s TWO_BOX.md.

NGC

NVIDIA’s catalogue of ready-made containers for its GPUs.

Example: the vLLM image 1_vllm.sh starts on the Spark.

Notes

This course’s plain word for the KV cache.

Example: 128 KB of notes per token for Llama 3.1 8B.

Number format

The way each of a model’s numbers is stored, including how many bits it takes.

Example: BF16 (16 bits), FP8 (8 bits), FP4 (4 bits).

numpy

A Python library for fast math on lists of numbers.

Example: the lesson 16 toys need nothing else.

NVFP4

NVIDIA’s 4-bit number format with built-in math on Blackwell.

Example: named in the second question of Day 4, step 1.

NVIDIA

The company whose GPUs run most data-center AI.

Example: vLLM and SGLang need an NVIDIA GPU.

Observability

Seeing what a live service is doing, such as speed, errors and memory, while it runs.

Example: llama.cpp’s --metrics flag publishes live counters at /metrics.

Offload (-ngl)

Choosing how many of the model’s layers run on the GPU; the rest run on the slower CPU.

Example: -ngl 99 puts every layer on the GPU.

Ollama

A one-command model server built on llama.cpp that chooses most settings for you.

Example: ollama serve answers on port 11434.

On-demand deployment

Fireworks’ name for a dedicated deployment: a private copy of a model, billed by the second at an hourly rate.

Example: the only way Fireworks serves a trained LoRA.

Open model (open weights)

A model whose weights anyone can download, run and fine-tune.

Example: gpt-oss-120b, which the Platform Map video calls a mid-sized open model.

OpenAI-compatible API

The request format OpenAI made popular, which most model servers accept, so one client works with all of them.

Example: Ollama, llama.cpp, MLX, vLLM and SGLang all speak it.

Optimizer state

The running averages the training method (Adam) keeps per trainable number to size each correction.

Example: part of the about 16 bytes per number that full fine-tuning needs.

Order of magnitude

About 10 times.

Example: 740 ms down to 36 ms is more than one order of magnitude.

ORPO (odds-ratio preference optimization)

A preference method that does example training and preference training in one run, with no reference model.

Example: --loss-method ORPO on Fireworks.

Out of memory

The server cannot get the memory it needs. It fails at startup (weights, or a cache reserved up front) or under load, as many conversations’ notes grow.

Example: the 131,072 16-bit row on a 16 to 24 GB Mac.

Output tokens

The tokens the model writes back; providers bill them per million, usually at a higher price than input.

Example: $0.60 against $0.15 per million for gpt-oss-120b.

Overfitting

Learning the training examples by heart instead of the task, so new examples go worse.

Example: 97% on training examples but 74% on held-out ones in the QLoRA On Your Own Box video.

p50 (median)

The middle value: half the requests are faster, half slower.

Example: ttft_p50_ms in the comparison table.

p95

The value 95 out of 100 requests stay under.

Example: 95 of 100 prompts under 3,000 tokens.

p99

The value 99 out of 100 requests stay under.

Example: P99 TTFT in vLLM’s own benchmark.

PagedAttention

vLLM’s way of handing out note memory in small pages as a conversation grows, instead of reserving the maximum.

Example: 2 to 4 times more users (The KV Cache video).

Parameters (8B, 70B)

The model’s learned numbers (weights).

Example: 8B = 8 billion.

Parse

Read text as data; parsing fails when the text is not valid JSON.

Example: parses_%: 86% when only asked in the prompt, on the practice server.

PCIe

The standard slot connection between a computer’s main board and a plug-in card such as a GPU; far slower than the GPU’s own memory.

Example: why moving layers off the GPU is slow (lesson 12).

Peak memory

The most memory a job uses at any moment.

Example: about 11 GB for the 8B fine-tune in the QLoRA On Your Own Box video.

Peak traffic

The most people using the service at the same moment.

Example: 120.

PEFT (parameter-efficient fine-tuning)

The family of methods that customise a model by training a tiny fraction of it.

Example: LoRA, QLoRA, DoRA, adapter layers, prompt tuning and IA3.

Percentile

The value a given share of requests stay under: p50 is half of them, p95 is 95 of 100.

Example: p95 TTFT of 41.4 ms at 8 people on the practice server.

perf_metrics

Timing and speculation figures Fireworks returns with a reply when the request asks for them (perf_metrics_in_response).

Example: the fireworks perf_metrics sample: line in lesson 13.

pip, wheel

pip installs Python packages; a wheel is a ready-built package file.

Example: a prebuilt wheel can expect a different CUDA than the machine has.

Pipeline (integration)

The customer’s code that sends requests and uses the replies automatically, with no person checking each one.

Example: a help desk that routes each ticket by its category.

Pipeline bubble

Time a GPU in a pipeline sits idle, waiting for work from the stage before it.

Example: too few requests in flight leave gaps in the line.

Policy

In reinforcement learning, the model being trained, seen as a rule for choosing answers.

Example: RLHF’s policy writes answers and the reward model scores them.

Policy gradient (REINFORCE)

Learning by sampling answers and making the above-average ones more likely.

Example: the toy’s sampled RLHF run, 5b.

Poll

Ask again at intervals whether a job has finished.

Example: submit_batch.sh polls every 30 seconds.

Port

A numbered door on a computer where one server program listens for requests.

Example: Ollama 11434, llama.cpp 8080, MLX 8081.

PP (pipeline parallelism)

Giving each GPU a block of consecutive layers, so each token passes from one GPU to the next; it tolerates slower links but adds waiting.

Example: PP=2 across two linked Sparks (TWO_BOX.md).

pp512 / tg128

llama-bench’s two tests: read a 512-token prompt, write 128 tokens, each in tokens per second.

Example: 1,050 and 86 at 4-bit.

PPO

The classic reinforcement algorithm in RLHF: sample answers, score them, make high scorers more likely, on a leash.

Example: it needs 4 models in memory.

Practice server (mock)

The kit’s stand-in server with realistic behaviour and made-up numbers.

Example: it has 8 slots.

Preamble

The long fixed opening sent with every request: instructions, tool list and rules.

Example: 3,028 tokens in Day 5’s script.

Precision

How many bits each stored number gets.

Example: 16-bit, 8-bit, 4-bit.

Predicted outputs

A Fireworks option where you send the expected answer, such as the original text of an edit, as the draft.

Example: sending the original document when asking for a small change.

Preference pair

Two answers to the same prompt, one chosen and one rejected.

Example: 800 in lesson 18, one per training ticket.

Preference tuning

Training a model on which answers people prefer, rather than on single correct answers.

Example: Day 12’s DPO step.

Prefill

The reading phase; the model reads the whole prompt at once.

Example: pp512.

Prefill/decode interference

Long prompts being read slow the replies being written for others on the same server.

Example: one of the things that moves the knee.

Prefix

The start of a prompt, up to the first token that differs from an earlier one.

Example: the shared 3,028-token opening.

Prefix caching (prompt caching)

The server keeps its notes on a prompt’s opening and reuses them when the next request starts the same way.

Example: 740 ms for call 1, 36 ms for the rest in Day 5’s sample.

Projection (weight grid)

One grid of learned numbers that a layer multiplies its input by.

Example: 4,096 x 4,096 = 16.8 million numbers in an 8B model’s attention.

Prompt

The text you send the model.

Example: a 3,000-token support question.

Prompt cache file

A file holding the saved notes on a prompt, so a later run can skip reading it.

Example: preamble.safetensors in the MLX script.

Prompt masking (--mask-prompt)

Counting the training error only on the answer, not on the system message or the customer’s text.

Example: run_variants.sh passes --mask-prompt.

Proof of concept (POC)

A small working version of the customer’s system, built to prove it meets their needs before they commit.

Example: slide 5: the next two weeks of a POC.

Prose traffic

open-ended, creative requests whose next words are hard to guess.

Example: “Write a poem about latency in the voice of a tired sea captain.”

Proxy

A stand-in measure used because the real goal is hard to measure directly.

Example: a reward model is a proxy for what the customer wants.

PyTorch

The Python library for model math that vLLM and many other AI tools are built on.

Example: a PyTorch and CUDA version clash is why the lab book avoids pip install vllm on the Spark.

q8_0 (cache)

llama.cpp’s 8-bit format for the notes, 1.0625 bytes per number.

Example: 8,704 MiB instead of 16,384.

Q8_0, Q4_K_M, Q3_K_M

llama.cpp’s weight formats at about 8.5, 4.8 and 3.9 bits per number.

Example: 8.5, 4.9 and 4.0 GB for an 8B model.

QLoRA

LoRA on a base model stored in 4 bits, so training fits in far less memory.

Example: about 11 GB at the peak for an 8B model in the QLoRA On Your Own Box video.

Quality cliff

The size below which answers suddenly get worse.

Example: the 3-bit version (Q3) slips on multi-step and format questions first.

Quantization (quantize)

Storing numbers with fewer bits.

Example: 16 to 8 to 4 bits, like saving a photo at lower quality.

Queueing

Waiting in line for a free seat on the server.

Example: at 16 people on the practice server, 8 wait.

Quota

A limit Fireworks sets on your account, such as how many GPUs you may use.

Example: firectl quota list shows them.

Qwen2.5 7B Instruct

An open model with 7.6 billion numbers, already trained to follow instructions; Days 7 and 8 fine-tune it.

Example: qwen2p5-7b-instruct on Fireworks.

Qwen2.5-0.5B-Instruct

A small open instruct model with about half a billion parameters, used in the training lessons.

Example: mlx-community/Qwen2.5-0.5B-Instruct-bf16, about 1 GB.

RadixAttention

SGLang’s prefix cache: it keeps saved notes in a tree of shared openings and orders requests to reuse them.

Example: the video’s trunk, branch and leaf picture.

RAG (retrieval-augmented generation)

An app that first looks up relevant documents and puts them in the prompt, then asks the question.

Example: Day 1’s long test prompt puts the context first and the question last.

RAM (memory)

The computer’s working memory, where the model must sit while it runs.

Example: 36 GB on the sample Mac in Day 1’s setup step.

Rank (LoRA rank)

How wide the adapter’s small grids are; a higher rank means more capacity and a bigger file.

Example: 8 on Day 7, whose script suggests 16 to 32 for harder tasks.

Reasoning model

A model that writes out its thinking before its final answer.

Example: lesson 07: put the schema in its prompt and validate afterwards.

Reference model

A frozen copy of the starting model that DPO and RLHF measure drift against.

Example: --reference-model-path in dpo_mlx.sh.

Region-pinned deployment

A dedicated deployment fixed to one region (US, EU or APAC), billed at 1.5 times the normal rate.

Example: $8,760 a month for one H100 instead of $5,840.

Reinforcement learning

Training by trial and error: the model tries answers, gets a score, and makes high-scoring answers more likely.

Example: the PPO loop in RLHF.

Replica

One complete running copy of the server and model.

Example: 18 replicas for 120 people.

Request

One message sent to the model for an answer.

Example: 95 of 100 requests under 3,000 tokens.

Reserved capacity

GPUs committed for a long term at a lower hourly price, billed whether used or not.

Example: the fourth capacity mode; the lab book warns never to suggest it casually.

response_format

The request setting that asks the server to enforce JSON (json_object) or a schema (json_schema).

Example: left out in prompt-only mode.

Reward hacking

A model raising its score by exploiting the scorer’s blind spots instead of getting better.

Example: in lesson 16’s toy the score rises to +1.96 while true value falls to +3.45.

Reward model

A model trained on people’s comparisons to give any answer a single score.

Example: in lesson 16’s toy it agrees with its labellers 71% of the time.

RFT (reinforcement fine-tuning)

Training against a grader that scores each answer from 0 to 1, for tasks where correctness can be checked automatically.

Example: the tests pass, the format validates, the number matches.

RLAIF

RLHF with an AI judge in place of human labellers.

Example: cheaper labels (the RLHF video).

RLHF (reinforcement learning from human feedback)

Training a model to chase a reward model’s score, learned from people’s comparisons, on a leash.

Example: Day 12’s first step runs it on a toy.

Roofline

A chart of a chip’s top speed against arithmetic intensity: a sloped memory-bound part, then a flat compute-bound roof.

Example: the GPU Bandwidth video’s chart.

Router (MoE)

The small part of a mixture-of-experts model that picks which experts handle each token.

Example: the receptionist in lesson 12’s picture.

Routing

Sending each request to the right model, or to the right copy of it.

Example: one of the layers a managed platform runs for you.

RPM (requests per minute)

How many requests an account may send each minute.

Example: 10 on a new Fireworks account until a payment method is added.

Rubric

The scoring guide a judge follows.

Example: 5 = specific, right owner, safe; 3 = plausible but vague; 1 = wrong or unsafe.

Sampling

Letting the model write answers at random, following its own probabilities.

Example: RLHF samples answers at every training step; DPO does not.

Scale

One shared number stored with a small block of numbers so they can be squeezed into fewer bits.

Example: why q8_0 is 1.0625 bytes, not 1.

Scale to zero

A deployment setting that gives the GPU back after a stretch with no requests, so billing stops.

Example: Fireworks on-demand deployments do it after 1 hour idle by default (the lab book).

Scheduler

The part of a server that decides which requests share the chip at each step.

Example: vLLM’s scheduler swaps requests in and out at every step.

Schema-valid rate

The share of replies that come back as correctly formatted data.

Example: 99% in the Quantization video’s sample.

Semaphore

A counter in code that lets only a set number of tasks run at the same time.

Example: asyncio.Semaphore(8) keeps at most 8 requests running at once.

Sequence length

How many tokens one training example holds.

Example: activations grow with batch size x sequence length.

Server

A program that waits for requests and answers them.

Example: llama.cpp’s llama-server.

Serverless

Fireworks’ shared, always-on models, billed per token you send and receive.

Example: Day 1’s step 3 measurements cost about 1 to 2 cents.

Serving (serve)

Keeping a model running so it can answer requests.

Example: $5.07 for 38 minutes of a dedicated GPU in Day 7’s sample.

Session

One conversation.

Example: an 8K session holds up to 8,192 tokens.

Session affinity

Sending every request of one session to the same replica, so it finds its saved notes.

Example: the user value in Day 5’s script, lesson-06-session.

Severity (severity_%)

How urgent a ticket is, from 1 to 5; severity_% is the share rated correctly.

Example: 70.0 before and 95.0 after the fine-tune in Day 7’s sample.

SFT (supervised fine-tuning)

Training on examples of the right answer, to teach format, tone and domain patterns.

Example: 800 labelled tickets on Day 7.

SGLang

A fast open-source server for NVIDIA GPUs that reuses shared prompt openings.

Example: it runs in Day 9’s optional Spark step.

Sharding (ZeRO, FSDP)

Splitting a model’s weights, gradients and optimizer state across several GPUs so each holds a share.

Example: a 70B full fine-tune across 16 H100s.

Sigmoid (σ)

A curve that turns any number into a value between 0 and 1; 0 gives 0.5.

Example: σ(2.5) = 0.92.

Single-stream

One request at a time, with no one else sharing the server.

Example: 54.9 tokens/s for one person on the practice server.

Sizing memo

A short document recommending how to run a customer’s workload: how many copies, which option, what it costs, with the maths shown.

Example: Day 10’s results/SIZING_MEMO.md quotes Day 5’s speed-up.

SLA (service level agreement)

A written promise to a customer about uptime or speed.

Example: COMPARISON.md says the Mac has none.

SLO (service level objective)

The speed promise.

Example: 95 of 100 requests get their first token within 500 ms.

Slot

One seat on the server, a conversation it can serve at the same time.

Example: --parallel 8 gives 8.

Slot cache reuse

llama.cpp keeps each seat’s notes from its last request and reuses the part that matches.

Example: why calls 2 to 12 are fast on llama.cpp.

Smoke test

A quick run of everything to prove the setup works before real use.

Example: make smoke, 16 lessons in about 2 minutes.

Speculative decoding

A small draft model guesses several tokens and the big model checks them all in one pass; the output does not change.

Example: Day 9 measures when it pays.

Speed-up (speedup_x)

How many times faster: the slow time divided by the fast one.

Example: 740 ÷ 36 = about 20.

Spend cap

A monthly limit on your Fireworks spending; at 100% Fireworks pauses the API.

Example: set with firectl quota update monthly-spend-usd.

SSH

A secure way to log in to another computer over the network and run commands there.

Example: remote.sh runs each script on the Spark over SSH.

Stable first, variable last

Build the prompt with the parts that never change at the top and anything that changes at the bottom.

Example: rules and tools first, the time and the question last.

Startup log

The lines a server prints as it starts, including how much memory it set aside.

Example: vLLM’s shows 8.4 GB of weights and 8,234 blocks.

Static batching

A fixed group of requests runs together, and the next group starts only when the longest one finishes.

Example: 12 requests on 4 seats finish at step 24, against 17 with continuous batching.

Structured traffic

Requests whose answers follow a predictable form, such as JSON data, lists, code or extraction.

Example: lesson 13’s 5 structured prompts.

Synthetic data

Examples made by a program or a model instead of collected from real users.

Example: the kit’s tickets, from data/make_tickets.py.

System prompt

The fixed instructions at the start of every request that set how the model behaves.

Example: “You are the support agent for Acme Cloud.”

Tail (latency tail)

The slowest few requests.

Example: p95 describes the tail.

Target model

The big model whose answer you want; it checks the draft model’s guesses and decides every token.

Example: Llama 3.1 8B in lesson 13.

TB

1,000,000,000,000 bytes, or 1,000 GB.

Example: about 1.1 TB to fully fine-tune a 70B model.

Teardown

Deleting everything that bills by the hour at the end of a session, then checking the list is empty.

Example: 5_teardown.sh, then make fw-check.

Tenant

One customer, or business unit, sharing a platform with others.

Example: per-tenant fine-tunes: one adapter each.

Tensor cores

The parts of an NVIDIA GPU built for fast matrix math.

Example: the H100’s tensor cores do math directly on FP8 numbers.

Terminal

The window where you type commands.

Example: Day 2 uses four at once: one per server, plus one to compare.

Test set (held-out set)

Examples kept out of training and out of every decision, used only for the final score.

Example: test.jsonl, 100 tickets; the eval uses 60.

Throughput

Total output across everyone at once.

Example: 275.1 tokens/s with 8 people.

Ticket triage

Sorting incoming support tickets by type and urgency.

Example: the kit’s shared task for Days 6 to 8.

TODO

A marker for work not done yet; the memo prints it where a number is missing.

Example: TODO (run lesson 01).

Token

A chunk of text, about three quarters of a word.

Example: 1,000 tokens is roughly 750 words.

Token ID

The number a tokenizer gives each token; models compare these numbers, not the text.

Example: the big model checks the draft’s token ids.

Tokenizer

The part of a model that cuts text into tokens and numbers them.

Example: Llama 3.1 8B and Llama 3.2 1B share one.

Tokens per second (tok/s)

How many tokens are written each second.

Example: 38.9 per person with 8 people.

Tool (tool definition)

An action an agent may ask for, described in the prompt so the model knows how to call it.

Example: search_tickets, get_invoice, refund, status_page and escalate in Day 5’s script.

Total parameters

Every parameter a model has; all of them must sit in memory, whether or not they are read for a token.

Example: 21 billion for gpt-oss:20b, a 13.8 GB file.

Toy (toy model)

A training problem shrunk until every number fits on one screen, keeping the real algorithm.

Example: one 64 x 64 layer in toy_finetune.py.

TP (tensor parallelism)

Splitting each layer of a model across several GPUs that work on every token together.

Example: TP=2 on two H100s joined by NVLink in one server; they sync inside every layer, so it needs very fast links.

Traffic profile

The shape of the load a test uses: tokens in, tokens out, requests per second, and the speed promise with its percentile.

Example: 1,024 in, 256 out, 200 requests per crowd size.

Training accuracy

How often the model is right on the examples it learned from.

Example: 97% for the messy run in the QLoRA On Your Own Box video.

Training set

The examples the model learns from.

Example: train.jsonl, 800 tickets.

Training token

One token read during training; managed training bills every token on every pass.

Example: 2,000 examples x 600 tokens x 2 passes = 2.4 million.

Trap (exit trap)

A shell instruction that runs a clean-up command whenever a script ends, even on Ctrl+C.

Example: fireworks_spec.sh deletes its deployment this way.

True value

In the toy, the hidden score of what people really want, which the fast labellers’ reward model cannot see.

Example: 4.775 for right-category JSON with a friendly line; E[true] is its average over the model’s answers.

TTFT (time to first token)

How long until the first token (word piece) of the reply appears.

Example: 41.4 ms (p95) at 8 people.

Turn

One call to the model within a session.

Example: 20 turns in the video’s agent session.

Unified memory

One pool of memory shared by a Mac’s CPU and GPU.

Example: the model and your browser share the same 36 GB.

Unit cost

The price of one unit of work, such as a million tokens written.

Example: $6.00 against $0.20 per million in the Platform Map example.

Usage block

The token counts a server returns with an answer.

Example: prompt_tokens and cached_tokens on Fireworks.

Utilisation

The share of time a GPU does useful work.

Example: a dedicated GPU costs the same whether it is busy or idle.

Validate

Check a reply with code after it arrives.

Example: Day 6’s parse and schema checks.

Validation loss

A score of how wrong the model still is on the validation set; lower is better.

Example: Day 8’s run prints it every 50 steps.

Validation set

Examples kept out of training and checked during it, to decide when to stop.

Example: valid.jsonl, 100 tickets.

Value model

In PPO, an extra model that predicts how good an answer in progress will turn out, to steady training.

Example: one of RLHF’s 4 models.

Verification pass

One pass of the big model that checks all the draft’s guesses at once, the way it reads a prompt.

Example: 3.36 tokens per pass at α 0.8 with 4 guesses.

Virtual environment (venv)

A private folder of Python packages for one project, turned on with source .venv/bin/activate.

Example: mlx_lm.server and felab live inside it.

vLLM

An open-source engine for NVIDIA GPUs built for many users at once.

Example: it uses PagedAttention.

Vocabulary

The full list of tokens a tokenizer knows, each with its own number.

Example: the draft and the big model must share one.

VRAM

The memory built into a separate graphics card. A Mac has none: its shared memory plays that role.

Example: ‘RAM is your VRAM’, in Day 1’s setup step.

Wall-clock time

The real time that passes from start to finish.

Example: total output divides all tokens by it.

Warm-up request

One request sent first and never counted, so loading the model does not skew the timings.

Example: Day 1’s stopwatch sends ‘Say hi.’ first.

Weights

The model’s learned numbers, fetched in full for every token written (a mixture-of-experts model fetches only the part it uses).

Example: 4.9 GB for Llama 3.1 8B at 4-bit.

Win rate

How often people or a judge prefer the new model’s answer to the old one’s.

Example: the DPO video says to measure it together with task accuracy.

Workload

One kind of job a customer runs.

Example: a 4K support chat (lesson 05).