Batch API
course 10 of 22
lesson 08 · lessons/08-batch-api
25 min Fractions of a cent Fireworks
What you will do and why
Jobs where nobody waits for the answer can go to Fireworks’ Batch API: one file in, answers later, at about half the usual price. You rehearse the round trip offline, then run it for real on 100 test tickets.
Why it matters: The Platform Map video’s example app spends $4.80 a day on 20 million tokens (word pieces). If none of it needed an instant answer, batch would cut that to about $2.40.
You are done when: The practice batch shows answered 100 of 100, and your Fireworks batch results are scored the same way. If the job is still running when you stop, you have its name written down and can finish the last two commands later.
Start this first
Batch jobs take minutes to hours, so start the real one now. It may finish during the video or take many hours (the lab book allows up to 72); both are normal. Run it in a second terminal: from 03-labs, source .venv/bin/activate, then the command. It builds the file of 100 requests, uploads it, starts the job and checks on it every 30 seconds. The script estimates fractions of a cent, so it is safe to start before the offline rehearsal below. Write down the job name from the create job: line: you need it to check on the job later. Ctrl+C is then safe: it stops the checking, and the job keeps running. The model, gpt-oss-20b, is no longer offered per token (serverless), but batch still takes it: batch runs any model that Fireworks can put on a GPU reserved for one customer (an on-demand deployment). If the job is still pending, with no error, when you finish today’s other work, see Stuck? below.
cd lessons/08-batch-api && bash submit_batch.sh accounts/fireworks/models/gpt-oss-20bThe Platform Map · 2:51
Download mp4 (11.0 MB)Chapters
In this video What a managed AI platform runs for its customers, the four ways it sells computing, and the sums that decide between paying per token and renting a whole GPU.
The narrator’s “next video” is the file order. Your next step is below.
3 key points
A managed platform runs everything from the GPUs up to the model, so the customer builds only the app. Its shared (serverless) models charge per token, not per hour of GPU time.
The 5 layers, bottom to top: GPUs and data centres; the serving engine (the program that runs the model); scaling and routing (adding copies as traffic grows, and sending each request to one); the model; then the app. A raw GPU cloud rents only the bottom layer and leaves the rest to you.
Four ways to buy computing: shared models billed per token, your own rented GPUs, batch for work that can wait, and GPUs reserved on a long contract.
Shared per-token models (serverless) suit traffic that comes in sudden peaks. Your own GPUs (dedicated) suit steady volume or strict speed. Batch, this step, costs about half of serverless because answers come later. Reserved GPUs cost less per hour but bill for the whole contract.
Do the sums before renting a GPU: most products start far below the point where a dedicated one pays.
The video’s app uses 20 million tokens a day: about $146 a month on serverless. One dedicated H100 (a top data-centre GPU) running all month costs $5,840, 40 times more.
In plain words
Section titled “In plain words”Some jobs have nobody waiting for the answer: testing a model on 100 tickets, or re-sorting last year’s documents. The Batch API takes all the requests as one file, runs them when Fireworks has spare GPUs, and charges about half the usual per-token price. You collect the answers later and match each one to its question by a label.
Picture it
Same-day dry cleaning costs extra; leave it for next week and it is cheaper, because the shop fits it in between rush jobs. Each item gets a numbered tag, because the clothes do not come back in the order they went in.
With real numberslesson 08’s scripts, the lab book and the Platform Map video’s example prices
- The file: the 100 test tickets the model never trained on, one request per line, labelled
t-000tot-099. - Price: about half of serverless (the lab book says 50%).
- The script’s own estimate: 100 short tickets on a small model cost well under a cent even at full price; in batch, about half that. It was written when gpt-oss-20b had a per-token price; it has none now, so there is no listed price to halve for it (see the run note).
- Scaled up, with the Platform Map video’s example: 20 million tokens a day costs $4.80 on serverless. If none of it needed an instant answer, batch would make it about $2.40 (4.80 ÷ 2).
- The trade: results come back in minutes to hours. Work where someone waits needs an answer in about 300 ms (0.3 seconds), the lesson’s example.
Words to know
- Batch API
- Fireworks’ way to send many requests as one file and collect the answers later, at about half the serverless price. Example: the lesson’s 100 tickets.
- JSONL
- A file with one JSON record per line. Example:
batch_input.jsonl, 100 lines. - custom_id
- The label on each request in a batch file, used to match every answer back to its question. Example:
t-000tot-099. - Serverless
- Fireworks’ shared, always-on models, billed per token you send and receive. Example: step 1’s Fireworks run.
From the lesson
lessons/08-batch-api/README.md
What and why
Section titled “What and why”Not every request needs an answer in 300 ms. Evals, back-fills, nightly classification and synthetic data can all wait an hour. The Batch API takes a file of requests, runs them when capacity is free and charges about half the serverless price. Prompt caching discounts stack on top.
It is the difference between a courier and the regular post: same letter, different urgency, different price.
interactive → serverless / dedicated pay for latencybulk → batch pay ~50%, get results in minutes to hoursRead the code first
Section titled “Read the code first”make_batch.py: one JSON line per request,{"custom_id", "body"}. The body is the normal chat-completions request.submit_batch.sh: upload the dataset, create the job, poll, then download.score_batch.py → find_content()joins results back to labels bycustom_id, because results may come back in any order.
cd lessons/08-batch-apipython make_batch.pypython score_batch.py --simulate --target mock # rehearse the whole loop offline
bash submit_batch.sh accounts/fireworks/models/gpt-oss-20bfirectl dataset download <OUTPUT_DATASET_ID>python score_batch.py <downloaded>.jsonlWhat each command does
python make_batch.pyTurns the 100 test tickets into one file,
batch_input.jsonl: one line per request, each with its own label (custom_id,t-000tot-099) and the usual chat request with the schema enforced. If you started the Fireworks job first, it already made this file; running it again writes the same file. Look forwrote …batch_input.jsonl (100 requests).python score_batch.py --simulate --target mockRehearses the whole round trip offline: sends the 100 requests to the practice server one by one, saves the answers in the shape a batch job returns, then scores them. It takes under a minute: 100 requests, each under about half a second (in step 1, 95 of 100 enforced replies from the practice server came back within 447 ms). Look for answered 100 of 100, schema_valid_% 100.0 and category_acc_% 78.0.
bash submit_batch.sh accounts/fireworks/models/gpt-oss-20bThe real thing (paid, fractions of a cent by the script’s estimate). It uploads the file to your Fireworks account and starts the batch job on a small model (gpt-oss-20b). Then it checks every 30 seconds until the job says COMPLETED, FAILED or EXPIRED. If you started it before the video, do not run it again: one job is enough. Look for the job name on the
create job:line and, at the end, the id of the output dataset (the results file the job saved in your account). A plain note on the model: gpt-oss-20b is no longer offered per token. On 27 September 2026 its model page said “Serverless: Not supported”. Batch still takes it: Fireworks’ batch guide (checked the same day) accepts “Any model that supports On-Demand Deployments”: models Fireworks can run on a GPU reserved for one customer, and gpt-oss-20b is one. The same guide prices batch at “50% off Serverless per-token prices”, and gpt-oss-20b no longer has one, so its batch price is not listed. If the job never leaves pending, see Stuck? below.firectl dataset download <OUTPUT_DATASET_ID>Downloads the results. Replace
<OUTPUT_DATASET_ID>with the output dataset id from the job details the script printed. Look for the name of the.jsonlfile it saves: the next command needs it.python score_batch.py <downloaded>.jsonlScores the real answers exactly like the rehearsal, matching each one to its ticket by
custom_id. Replace<downloaded>.jsonlwith that file’s path. Look for answered at or near 100 of 100, and your model’s own schema_valid_% and category_acc_%: these are real numbers, not the practice server’s. gpt-oss-20b also writes out its thinking before it answers (step 1’s check 3). If schema_valid_% is low or replies are empty, that thinking may have used up the batch file’s 120-token reply limit: write that down as a finding.
What you should see
Section titled “What you should see”How to read it
Each results file gives one row. answered counts results matched to a ticket by custom_id, out of the 100 sent (of). The next three are shares of the answered results: those that follow the schema, have the right category, and have the right severity (urgency, 1 to 5). The practice run’s 100 of 100, 100%, 78% and 79% are made-up; your Fireworks row is a real model’s score on the same tickets.
file answered of schema_valid_% category_acc_% severity_acc_%simulated_results.jsonl 100 100 100.0 78.0 79.0That is the practice server’s made-up score on the 100 test tickets, which no model trains on. Your Fireworks run gives a real model’s score on the same tickets. The fine-tunes on Days 7 and 8 are each compared with their own base model.
Check yourself
Section titled “Check yourself”2 questions. Say your answer out loud, then tap to check it.
Which of a customer’s jobs (workloads) would you move to the Batch API, where requests go in as one file and the answers come back later at about half price?Show answerHide
In plain words
Any job where nobody is waiting for the answer: testing a model on a fixed set of examples, reprocessing old records, sorting documents into categories, or generating practice data.
Picture it
A restaurant cooks for the diners at its tables first. Tomorrow’s catering order waits for the quiet hours, when the kitchen is idle and the work costs less.
With real numberslesson 08 and the Platform Map video’s example prices
- Work where someone waits, such as a chat window, pays for speed: the lesson’s example is an answer in 300 ms (0.3 seconds). It stays on serverless or dedicated.
- A test of the model on 100 tickets can wait minutes to hours: batch, at about half price.
- The Platform Map video’s app, 20 million tokens a day: $4.80 on serverless, about $2.40 if it could all go to batch.
- Discounts for a repeated prompt opening (prompt caching, Day 5) apply on top of the batch price.
Words to know
- Back-fill
- Running a job over all your existing records at once, for example after adding a new field.
- Embedding
- A list of numbers that captures a text’s meaning, used for search and matching. Recomputing them for every document is a typical back-fill.
- Synthetic data
- Examples made by a program or a model instead of collected from real users. Example: the kit’s tickets.
- Eval
- A repeatable test of a model on a fixed set of prompts. Example: the 100 test tickets.
Go deeper: the engineer version
The kit's question
Which customer workloads would you move to batch?
The kit's answer
Anything where no user is waiting: evals, embeddings back-fills, document classification, synthetic data.
More detail: The test is whether anyone waits, not the size of the job. Batch jobs run when the provider has idle GPUs (submit_batch.sh), so they can take minutes to hours; the lab book lists 12 to 72 hour windows. Prompt-caching discounts stack on top of the batch price. Anything user-facing (chat, agents, code completion) stays on serverless or a dedicated deployment.
In the batch file, every request carries its own label, custom_id (such as t-000). Why does it need one?Show answerHide
In plain words
So each answer can be matched to its question. Results can come back in any order, and failed requests go to a separate error file, so answer number 5 is not always for question number 5.
Picture it
A cloakroom gives each coat a numbered tag. If staff instead handed coats back in line order and one coat had gone missing, everyone behind it would get the wrong coat. The tag makes the order irrelevant.
With real numberslesson 08’s scripts
- The lesson’s file holds 100 requests, labelled
t-000tot-099. - Suppose request
t-042fails. The results then hold 99 answers. - Even if the other 99 came back in order, matching by line order would pair every answer after the gap with the wrong ticket: 57 wrong pairings (
t-043tot-099). - Matched by
custom_id, the 99 are scored correctly and the table says answered 99 of 100.
Words to know
- custom_id
- The label on each request in a batch file. Example:
t-042. - Error file
- Where a batch job puts the requests that failed, apart from the results. Example: why answered can be below 100.
- JSONL
- A file with one JSON record per line. Example: one line per request, one line per result.
Go deeper: the engineer version
The kit's question
Why custom_id?
The kit's answer
Results aren’t ordered, and some requests may fail and land in the error file.
More detail: Batch output is a set of result records, not an ordered response stream: order is not guaranteed and failures are written separately. score_batch.py builds a table of labels keyed t-000 to t-099 and looks each result up by custom_id. Its find_content() then searches the record for the reply text, because the exact output shape can differ between providers and versions.
Explain what you learned
Section titled “Explain what you learned”Question Testing every change on our own data sounds expensive. How do we afford it?
One clear answer
Evaluation is batch work. At half price you can afford to evaluate every prompt or model change, not just the big ones.
What this means
- “Evaluation is batch work”: Testing a model on a fixed set of examples has nobody waiting for the answers, so it can go in one batch file and come back later.
- “At half price”: The Batch API charges about half the normal per-token price (serverless). The Platform Map video’s $4.80 a day would be about $2.40.
- “you can afford to evaluate every prompt or model change”: The same budget buys twice as many test runs, so every change to the instructions or the model gets tested. Today’s 100 tickets cost well under a cent, by the script’s own estimate.
- “not just the big ones”: Small tweaks cause surprises too. Testing only the big releases lets those slip through.
Your numbersSaved on this device and collected in the Day 6 wrap-up.
Hint: From python score_batch.py --simulate --target mock: the answered and category_acc_% columns. Expect 100 of 100 and 78.
Hint: submit_batch.sh prints it on the create job: line: triage-batch- followed by the date and time you started it.
Hint: From python score_batch.py <downloaded>.jsonl: the answered, schema_valid_% and category_acc_% columns. Write all three, for example “100 / 100 / 78”.
Run make fw-check now: it should list nothing. It lists anything on Fireworks that is still billing, so an empty list means nothing is costing you money.
Done when
Section titled “Done when”The practice batch shows answered 100 of 100, and your Fireworks batch results are scored the same way. If the job is still running when you stop, you have its name written down and can finish the last two commands later.
Stuck?
Section titled “Stuck?”Connection refusedon the practice run.- The practice server is not running. From the
03-labsfolder runmake mock &, then run the command again. firectl: command not found.- Install it as on Day 1:
brew tap fw-ai/firectl && brew install firectl, thenfirectl signin. - A firectl command rejects its options.
- Commands change. Run it with
--help, for examplefirectl batch-inference-job create --help, and follow what it says; the script’s own note says the same. - The job is still running when you have to stop.
- Nothing is lost: Ctrl+C stops the checking, not the job. Later, from
lessons/08-batch-api, runfirectl batch-inference-job get <job name>with the name from thecreate job:line. When it says COMPLETED, run the last two commands. - The job stays pending and shows no error.
- Fireworks’ batch guide (checked 27 September 2026) warns that a job on a model batch does not support can stay pending with no error. Batch should still take gpt-oss-20b, but if yours has not started by the time you finish today’s other work, start a second job on gpt-oss-120b, which is offered per token: from
lessons/08-batch-api, runbash submit_batch.sh accounts/fireworks/models/gpt-oss-120b. Write down its new job name and follow that one. Your score is then gpt-oss-120b’s, so label it that way. - The job ends FAILED or EXPIRED.
- Read the job details the script printed at the end for the reason. Model ids change: if the model is no longer offered, copy a current one from the Fireworks model library and run
bash submit_batch.sh <that id>. answeredis below 100.- Some requests failed and went to the job’s error file instead of the results. The score covers the ones answered, each matched to its ticket by
custom_id: that is exactly why the label exists. Pass a results file, or --simulate.- The path after
score_batch.pydoes not point to a file. Use the exact name of the downloaded.jsonlfile;lslists what is in the folder.
Go deeper: the lab book's Fireworks lab for this step All fixes
Code in this step
Section titled “Code in this step”make_batch.py Python · 42 lines
"""Lesson 08 · step 1 — turn the eval set into a Batch API input file.
One line per request: {"custom_id": ..., "body": {<normal chat-completions body>}}.custom_id is how you join results back to labels — results may come back in any order.
python make_batch.py # → batch_input.jsonl (100 tickets) python make_batch.py --model accounts/fireworks/models/gpt-oss-20b"""import argparseimport jsonfrom pathlib import Path
from felab.tickets import SCHEMA, SYSTEM, load_test
HERE = Path(__file__).parent
def main() -> None: ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) ap.add_argument("--model", default=None, help="optional: some setups want the model per line") a = ap.parse_args()
rows = load_test() out = HERE / "batch_input.jsonl" with out.open("w") as f: for i, r in enumerate(rows): body = { "messages": [{"role": "system", "content": SYSTEM + "\nSchema:\n" + json.dumps(SCHEMA)}, {"role": "user", "content": r["ticket"]}], "max_tokens": 120, "temperature": 0, "response_format": {"type": "json_schema", "json_schema": {"name": "triage", "schema": SCHEMA}}, } if a.model: body["model"] = a.model f.write(json.dumps({"custom_id": f"t-{i:03d}", "body": body}) + "\n") print(f"wrote {out} ({len(rows)} requests)")
if __name__ == "__main__": main()score_batch.py Python · 82 lines
"""Lesson 08 · step 3 — join batch results back to labels and score them.
python score_batch.py results.jsonl # a downloaded Fireworks output file python score_batch.py --simulate --target mock # no Fireworks needed: runs batch_input.jsonl through any target synchronously # and writes simulated_results.jsonl in the same shape, then scores it
The output file format can differ slightly between providers and versions, so wedon't hard-code a path to the answer: we find `custom_id`, then the first assistant`content` string anywhere inside that line."""import argparseimport jsonfrom pathlib import Path
from felab import add_target_args, client, record, resolve, tablefrom felab.tickets import grade, load_test
HERE = Path(__file__).parent
def find_content(obj): """Depth-first search for choices[..].message.content (or any 'content' string).""" if isinstance(obj, dict): msg = obj.get("message") if isinstance(msg, dict) and isinstance(msg.get("content"), str): return msg["content"] for v in obj.values(): c = find_content(v) if c is not None: return c elif isinstance(obj, list): for v in obj: c = find_content(v) if c is not None: return c return None
def simulate(args) -> Path: t = resolve(args) cli = client(t) out = HERE / "simulated_results.jsonl" with out.open("w") as f: for line in (HERE / "batch_input.jsonl").read_text().splitlines(): req = json.loads(line) body = {k: v for k, v in req["body"].items() if k != "model"} resp = cli.chat.completions.create(model=t.model, **body) f.write(json.dumps({"custom_id": req["custom_id"], "response": resp.model_dump()}) + "\n") print(f"simulated {out}") return out
def main() -> None: ap = add_target_args(argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)) ap.add_argument("results", nargs="?", help="downloaded results .jsonl") ap.add_argument("--simulate", action="store_true") a = ap.parse_args() path = simulate(a) if a.simulate else Path(a.results or "") if not path.is_file(): raise SystemExit("Pass a results file, or --simulate.")
labels = {f"t-{i:03d}": r for i, r in enumerate(load_test())} grades = [] for line in path.read_text().splitlines(): rec = json.loads(line) cid = rec.get("custom_id") if cid in labels: grades.append(grade(find_content(rec), labels[cid])) n = len(grades) row = {"file": path.name, "answered": n, "of": len(labels), "schema_valid_%": 100 * sum(g["valid"] for g in grades) / max(1, n), "category_acc_%": 100 * sum(g["category_ok"] for g in grades) / max(1, n), "severity_acc_%": 100 * sum(g["severity_ok"] for g in grades) / max(1, n)} print(table([row])) record("08-batch", row)
if __name__ == "__main__": main()submit_batch.sh Bash · 41 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 08 · step 2 — submit to the Fireworks Batch API (50% off serverless prices).## Batch = "I don't need the answer in 300 ms; I need 100k answers by tomorrow".# The provider fills idle GPU time with your work, so it is cheaper.# Evals, back-fills, synthetic data, nightly classification → batch.## 100 short tickets on a small model costs well under a cent even at full price.# Flags per the docs at time of writing; `firectl batch-inference-job create --help` wins.# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"MODEL=${1:-accounts/fireworks/models/gpt-oss-20b}STAMP=$(date +%m%d-%H%M)IN="triage-batch-in-$STAMP"JOB="triage-batch-$STAMP"
[[ -f batch_input.jsonl ]] || python make_batch.py
echo "▸ upload input dataset: $IN"firectl dataset create "$IN" ./batch_input.jsonl
echo "▸ create job: $JOB (model $MODEL)"firectl batch-inference-job create --job-id "$JOB" --model "$MODEL" --input-dataset-id "$IN"
echo "▸ poll every 30 s (Ctrl-C is safe — the job keeps running; re-check with the get command)"while true; do state=$(firectl batch-inference-job get "$JOB" | grep -iE "^\s*state" | head -1 || true) echo " $(date +%H:%M:%S) $state" [[ "$state" =~ COMPLETED|FAILED|EXPIRED ]] && break sleep 30done
firectl batch-inference-job get "$JOB"cat <<EOF
▸ Next: find the output dataset id in the output above, then firectl dataset download <OUTPUT_DATASET_ID> python score_batch.py <downloaded results .jsonl>EOF