AI Field Engineer · Enterprise · Fireworks AI

Learn it by running it.

A hands-on guide to AI field engineering: learn how inference platforms work, then run nineteen labs from your Mac — seven against Fireworks for a few dollars, nine locally for nothing, three on how models are trained — and eighteen narrated videos that explain every concept underneath them.

7 Fireworks labs · ~$6 total 9 local labs · Mac-first · $0 18 narrated videos · rendered Prices & docs checked 23 Sept 2026
The job

What "AI Field Engineer" actually means here

Field engineers build proofs of concept, benchmark systems, debug production issues, and explain tradeoffs to customers. Each skill below maps to a lab you can run.

Field engineering taskWhat to learnWhere to practise
Build end-to-end POCs inside customer codebasesCan you ship working code under someone else's constraints, fast?Capstone
Architect inference foundations and size deploymentsMemory maths, concurrency, GPU shapes, cost per million tokensL3 · L5 · F1
Run load tests; set latency, throughput and cost baselinesDo you benchmark properly — traffic profile, p95, goodput?L3 · F1
Deploy new model families on vLLM and SGLang; pick shapes, quantization, serving patternsHands-on-keyboard with the open serving stackL2 · L4 · L9
Guide model selection, fine-tuning strategy (SFT, DPO, RFT) and evaluationDo you know which method fits which problem — and the cost of each?F5 · F6 · L8 · T1–T3
Build and run fine-tuning pipelines with customersDataset format, job launch, deploy, measure, iterateF5 · L8
Design evaluation frameworks measuring production qualityNot MMLU: task accuracy, tool-call success, latency, costF6
Lead discovery; earn trust with ML engineers and VPs in one meetingTwo registers, one story. Practise both out loudTalk tracks
Turn recurring pain into product proposalsHave you got an opinion about what's missing in the platform?Capstone memo

Build a portfolio of evidence

  • A benchmark you ran on Fireworks and on your own hardware, with the traffic profile written down.
  • A fine-tune you shipped: dataset, job, deployment, eval delta, cost.
  • A sizing memo with two or three options and the maths visible.
  • One product criticism, specific and constructive. Explain the tradeoff and suggest a fix.

Their own writing, worth citing

  • Frontier RL Is Cheaper Than You Think — over 98% of BF16 weights are bit-identical between consecutive RL checkpoints, so shipping ~2% deltas cut cross-region transfer by roughly 94%.
  • The Fine-Tuning Bottleneck Isn't the Algorithm — the blocker is integration friction and iteration speed; the fastest teams run 100+ jobs, one customer ships a checkpoint about every five hours.
  • Open-source agents with frontier advisors — an open model doing the work with a frontier model as a sparse advisor beat frontier-only on a legal benchmark at roughly a third of the cost.

Read all three before the loop. Quoting one accurately, then asking a good question about it, is the cheapest credibility you will ever buy.

The product

Everything Fireworks sells, in one page

The platform splits into four jobs: serve a model, train a model, evaluate it, and route between models. Underneath all of it sits their own inference engine — the reason the pitch is "specialized intelligence you own", not "cheap GPUs".

Who runs what

The stack, and where the customer stops

Your app
prompts, RAG, agents, product logic — the customer
Model + adapters
library models, your uploaded weights, LoRA adapters — shared
Serving engine
FireAttention kernels, FP8/FP4 quantization, prompt caching — Fireworks
Optimisation
FireOptimizer, adaptive speculative decoding, multi-LoRA batching — Fireworks
Capacity
serverless · on-demand · batch · reserved · Virtual Cloud / BYOC — Fireworks
Hardware
H100 · H200 · B200 · B300 · GB300, multi-region — Fireworks

Contrast to hold in your head: a raw GPU cloud sells you the bottom row and hands you the rest. Fireworks sells everything except your application — and increasingly, the training loop too.

Serverless

Per-token API, OpenAI-compatible, Standard / Priority / Fast tiers. Billed on input, cached input and output separately.

Default starting point. No provisioning, no idle cost.

On-demand deployments

Private replicas billed per GPU-second, autoscaling, scale-to-zero (1 hour idle by default — shorten it).

Needed for LoRA serving, custom shapes, speculative-decoding flags.

Batch API

JSONL in, JSONL out, 50% of serverless price, caching discounts stack. 12–72h windows.

Where your eval runs and bulk jobs belong.

Reserved capacity

Committed GPU capacity, lower hourly price, sold through sales.

Bills for the whole term whether used or not — never suggest it casually.

Managed Training

SFT, DPO, ORPO and RFT as managed jobs via UI, firectl or REST. Priced per 1M training tokens.

The "own your weights" story, and the cheapest real fine-tune you can run.

Training API

Custom training loops in Python — custom loss, distillation, inference in the loop. Serverless (per token) or dedicated (per GPU-hour) compute.

GA in Aug 2026. This is what "Fireworks Training" means now.

Eval Protocol

Their open-source, trace-first eval framework (pip install eval-protocol). Rule-based, LLM-judge or hybrid rewards; feeds RFT.

The bridge from "we think it's better" to a reward function.

Multi-LoRA

One base deployment, many adapters (default quota 100). Needs a BF16 shape with --enable-addons; FP8/FP4 shapes can't host adapters.

How per-customer fine-tunes stay affordable.

FireAttention

Their own kernels. V4 targets B200 with NVFP4 and claims 250+ tokens/sec.

The answer to "why not just run vLLM ourselves?"

Speculative decoding

Default model-based speculation, custom --draft-model, n-gram, or predicted outputs. Measure with perf_metrics_in_response.

Docs warn a bad drafter makes things slower — good nuance to voice.

Structured output

JSON mode and JSON-schema mode (2020-12, $ref/$defs, no external refs), plus grammar mode.

Two docs rules: put the schema in the prompt too; for reasoning models omit response_format.

Virtual Cloud / BYOC

The engine runs inside the customer's own VPC, across many clouds and regions.

The answer to data-residency objections in regulated APAC accounts.

Routing (Nexus · FireRouter)

Task-level, cache-aware routing across open and closed models, with centralized cost controls.

Reframes "which model?" as a portfolio question. Great discovery hook.

Serverless, $ per 1M tokensinputcachedoutput
Nemotron 3.5 Lightning 30B A3B0.050.010.20
GLM 5.3 Flash0.150.030.50
gpt-oss 120B0.150.0150.60
DeepSeek V4.1 Flash0.300.0061.20
GLM 5.31.400.264.40
Qwen 3.8 Max2.000.256.00
Kimi K33.000.3015.00
Size tier <4B / 4–16B / >16B0.10 · 0.20 · 0.90 flat
On-demand GPUs$ / hour
H100 80GB · H200 141GB8.00
B200 180GB13.00
B300 288GB15.00
GB300 288GB20.00
Billed per GPU-second, no start-up charge. Region-pinned (US/EU/APAC) costs 1.5×.
Managed training, $ / 1M training tokensLoRA SFTLoRA DPOFull SFT
≤16B0.501.001.00
16–80B3.006.006.00
80–300B6.0012.0012.00
>300B10.0020.0020.00
Prices move weekly. Everything above was read from Fireworks' pricing and docs pages on 23 September 2026, and several models were scheduled for retirement two days later. Re-check before using a price in a sizing decision, and date your estimate.
Before you spend a cent

Spend control, and the one trap that costs real money

You can do every Fireworks lab here for roughly the price of lunch — as long as you understand one asymmetry: per-token endpoints cost cents; a GPU deployment costs dollars per hour, whether or not you use it.

The trap: a LoRA fine-tune cannot be served on serverless. Fireworks' docs are explicit that trained LoRA models deploy only to on-demand (dedicated) deployments. So the moment you finish a $2 fine-tune, serving it has an $8/hour floor. Deploy it, run your eval, screenshot the numbers, and firectl deployment delete it in the same sitting.
Checklist · do this first

Guardrails, 15 minutes

0 / 7
firectl signin
firectl whoami
firectl quota list                          # see GPU + LoRA + spend quotas
firectl quota update monthly-spend-usd      # hard pause at 100%, warning at 80%
firectl deployment list                     # should be empty when you finish for the day
Calculator

What will this workload cost?

–
Cost / request
–
Per day
–
Per month
–
Saved by caching
–
Dedicated break-even
–

Break-even compares the monthly serverless bill against one H100 on-demand replica at $8/hour running 24×7 ($5,840/month). Below that line, per-token is almost always the right answer — and it is the honest answer to give a customer.

Calculator

What will a fine-tune cost?

Training tokens
–
Training cost
–
Serving (1× H100)
–
Total for this experiment
–

A realistic first fine-tune — 2,000 examples, 600 tokens each, 2 epochs of LoRA SFT on an 8B model — is about 2.4M training tokens, which is roughly $1.20. The serving window is what you actually have to watch.

Hands-on · the companion code

Page on the left, terminal on the right

Every lab below has a folder in the field-engineer-labs repo. The page explains what to do and why; the folder holds commented code that you run in order, and every script saves its numbers to results/. The capstone turns those numbers into a customer sizing memo.

1 · Watch

The matching 3-minute video, for the idea in plain words.

2 · Read

The lab card here: the point of the lab, the steps, the line to say.

3 · Run

cd lessons/NN-…, read its README, run the scripts. Try --target mock first, then the real server.

4 · Keep

The numbers land in results/*.csv. Lesson 15 builds the memo from them.

Start in ten minutes, with no model download

cd field-engineer-labs
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt && pip install -e .
cp .env.example .env

make mock &      # offline fake LLM that behaves like a real one
make smoke       # every lesson, end to end, ~2 min
make setup       # then: Ollama, llama.cpp, MLX on your Mac

Every script takes the same switch, so one harness measures every backend:
--target mock | ollama | llamacpp | mlx | fireworks | spark

What is in the repo

felab/            shared toolkit (targets, timing, results, graders)
  mock_server.py  offline server: prefill, batching, prefix cache, JSON
data/             the triage dataset used by lessons 07–11
lessons/NN-*/     README + commented scripts, in course order
results/          your measurements → the capstone memo
scripts/          smoke_test.sh (also runs in CI)
Makefile          make help

The nineteen lessons, in order

#Lesson and folderLab cardVideoCost

Only lessons 10, 13 and (optionally) 18 create billed Fireworks deployments. Lesson 10 ends in a teardown script, and lesson 13 deletes its deployment automatically on exit. Run make fw-check at the end of every session.

Hands-on · cloud

Seven Fireworks labs, about six dollars

Each lab is scoped to one practical field engineering skill, and ends with something you can show or say. Run them in order, after the guardrails checklist in the previous section.

Your hardware

Your two machines: the Mac and the Spark

Use the Mac for iteration and the Spark for the server stack. Both are unified-memory machines, so the same sizing logic applies to each — only the bandwidth changes. The Spark is compute-rich and bandwidth-poor: 128 GB of unified memory at 273 GB/s, with prefill performance far beyond its price class and decode speed that is firmly mid-range. That combination makes it an unusually good teaching machine — the bottleneck is visible in every measurement.

Unified memory
128 GB
~115 GB usable for weights + KV
Memory bandwidth
273 GB/s
the number that sets decode speed
vs H100
12× less
H100 = 3.35 TB/s
Peak
~1 PFLOP
FP4, with sparsity · 5th-gen tensor cores
Precisions
FP8 · FP4
NVFP4 and MXFP4 both supported
Two boxes
200 Gb QSFP
ConnectX-7, RoCE, up to 405B in FP4
Calculator

Decode ceiling: predict before you measure

–
Bytes read / token
–
Theoretical ceiling
–
Realistic (×0.65)
–
Published measurement
–

Pick your own Mac in the device list — the same arithmetic governs Apple silicon, which is also unified memory. For dense models, measured single-stream decode lands at roughly 60–85% of the ceiling. For mixture-of-experts models the file is much bigger than the bytes actually read per token — which is exactly why gpt-oss-120b (59 GiB on disk) decodes faster on a Spark than a dense 14B model.

The demo that lands

Invert the formula on an MoE model and you can discover its active footprint in front of someone: gpt-oss-120b measures about 55 tokens/sec on a Spark, so 273 ÷ 55 ≈ 5 GB read per token — out of a 63 GB file. Under 10% of the model is touched per token. That single sum explains MoE better than any diagram.

Published DGX Spark measurementsBackendPrefill tok/sDecode tok/s
Llama 3.1 8B · NVFP4TensorRT-LLM10,25738.7
Qwen3 14B · NVFP4TensorRT-LLM5,92922.7
gpt-oss 20B · MXFP4llama.cpp3,67082.7
gpt-oss 120B · MXFP4llama.cpp1,72555.4
Llama 3.1 8B · FP8 · batch 1 → 32SGLang7,99120.5 → 368 aggregate
Llama 3.1 70B · FP8 · batch 1SGLang8032.7
Qwen3 235B · NVFP4 · two SparksTensorRT-LLM23,47711.7
Stack hygiene

Five rules that save you a weekend

Containers, not host pip

NVIDIA's own guidance: the host OS, driver and CUDA move together on a fixed cadence; get new features from NGC containers instead.

The vLLM pip path on arm64 has bitten people (torch/CUDA mismatches, --enforce-eager costing 20–30%).

nvidia-smi memory is blank

Unified memory means the usual memory fields report N/A. Use free -h, htop, or the DGX Dashboard on localhost:11000.

Don't debug an OOM with the wrong instrument.

Drop caches between big runs

sudo sh -c 'sync; echo 3 > /proc/sys/vm/drop_caches' appears in four of NVIDIA's own playbooks.

Page-cache pressure in a unified pool shows up as mysterious slowdowns.

bitsandbytes needs a hint

export BNB_CUDA_VERSION=130 before QLoRA runs.

Straight from NVIDIA's fine-tuning guide: CUDA 13.1 isn't supported yet.

Build llama.cpp for sm_121

-DCMAKE_CUDA_ARCHITECTURES=121a-real

GB10 is compute capability 12.1; generic builds miss it.

Start from the playbooks

NVIDIA ships ~65 DGX Spark playbooks (vLLM, SGLang, TensorRT-LLM, llama.cpp, Unsloth, PyTorch fine-tune, NVFP4 quantization, two-Spark clustering).

They pin working container tags — which is most of the battle on arm64.

Hands-on · local · macOS

Nine local labs, Mac-first, zero dollars

Everything here runs on an Apple-silicon Mac with Ollama, llama.cpp and MLX. The last lab is the Spark track over SSH, for the parts that need real server software: vLLM, SGLang, NVFP4 and multi-node. Same benchmark script throughout, pointed at different base URLs.

Hands-on · how models are trained

Three training labs, from toy to real

These pair with the training deep dive, videos 14 to 18. T1 shrinks each algorithm until you can see it working, in numpy, in seconds. T2 runs full fine-tuning, LoRA and DoRA on a real model on your Mac. T3 turns realistic failures into preference pairs and runs DPO locally, and optionally on Fireworks.

The artifact

Capstone: a POC you can walk someone through

One workload, taken end to end, is worth more than fifteen disconnected labs. Pick something enterprise-shaped — support-ticket triage, contract clause extraction, a coding agent for an internal SDK — and run the whole arc.

  1. DiscoveryWrite the brief: task, volume, latency target, quality bar, data constraints
  2. Eval set50–100 real examples with graders, before you touch a model
  3. BaselineA closed frontier model: quality, latency, cost per 1K tasks
  4. Open modelSame eval on 2–3 open models via Fireworks serverless
  5. SpecialiseLoRA SFT on your data; re-run the eval; measure the delta
  6. Local optionSame model on the Spark for the data-residency variant
  7. Sizing memoTwo or three deployment options with the maths shown
  8. ReadoutFive slides: problem, evidence, recommendation, risks, next step

Sizing memo template

  • Traffic profile: requests/day, peak concurrency, input/output token distribution.
  • Latency targets: p95 TTFT, p95 ITL, and what the product does when they are missed.
  • Option A — serverless: $/1M in and out, cache hit rate, monthly total, limits.
  • Option B — dedicated: GPU type and count, replicas, utilisation, monthly total, break-even volume.
  • Option C — specialised model: smaller fine-tuned model, training cost, quality delta, new unit cost.
  • Risks: cold starts, long-context KV pressure, model retirement, eval drift.
  • Recommendation in one sentence, with the number that drives it.

Eval harness, minimum viable

  • Inputs, expected outputs, and a grader per example (exact match, JSON-schema validity, tool-call success, or an LLM judge with a rubric).
  • Report quality, p50/p95 latency and $ per 1,000 tasks in the same table. Always the three together.
  • Run it through the Batch API at half price when latency isn't what you're measuring.
  • Keep every run's raw traces. Fireworks' Eval Protocol is trace-first for exactly this reason, and traces are what turn an eval into an RFT reward later.
Watch first, then run

Eighteen narrated explainers

Every concept in this lab book has a 3–4 minute video: what it is, why it matters, how you will run it, with worked examples. Videos 14–18 are a training deep dive: full fine-tuning, SFT, LoRA and its family, RLHF and DPO. Rendered with Remotion, narrated with Kokoro, subtitled, and delivered as MP4 files in this conversation. Each one opens with what you will learn and closes with a three-point recap and the lab to run next.

#VideoWhat it explainsThen run
1inference-101Prefill vs decode, TTFT, ITL, throughput, goodputF1
2kv-cacheKV maths, PagedAttention, FP8 KV cacheL5
3gpu-bandwidthRoofline, decode ceiling, device comparisonL3 · L4
4quantizationFP16 → FP8 → 4-bit, NVFP4, KV quantizationL4
5batchingContinuous batching, chunked prefill, speculationL3
6trainingSFT, LoRA, DPO, RFT and how to chooseF5 · L8
7offloadingSizing formula, offloading, MoE, TP vs PPL7
8platformServing layers, capacity modes, maturity curveCapstone
9serving-stackWhat an inference server does; why containers; the five flagsL2 · L9
10benchmarkingTraffic profiles, concurrency sweeps, percentilesL3
11prefix-cachingWhy agent traffic is mostly repeats, and how to exploit itL6 · F2
12qloraFrozen 4-bit base + adapter; data and evals decide itL8 · F5
13scale-outSpeculative decoding, acceptance rate, what a second box buysL9
14full-finetuningThe training loop; every weight moves; ≈16 bytes/param; ZeRO, FSDP, checkpointing; when it is worth itT1 · T2
15sftData format, loss masking, how much data, the knobs, reading loss curves, what SFT can't doT2 · F5 · L8
16peftLoRA's B·A, adapter arithmetic, QLoRA, DoRA, adapters, prefix tuning, IA³, multi-LoRA economicsT2
17rlhfWhy preferences, reward models (Bradley-Terry), PPO, the KL leash, reward hacking, GRPO and RFTT1
18dpoThe DPO insight and loss, one step with numbers, pairs from your own product, IPO/KTO/ORPO/SimPOT3
Regenerate on your Mac

Kokoro through MLX, then render

The delivered MP4s were rendered on Linux with the ONNX build of Kokoro. On your Mac, mlx-audio runs the same Kokoro-82M weights on the Apple GPU — faster, and it is the path to use if you want to change the script, the voice or the pacing.

cd fireworks-explainers
npm install
npx remotion browser ensure

# narration on Apple silicon (MLX)
pip install mlx-audio soundfile numpy
python scripts/tts_mlx.py --voice af_heart --speed 0.96      # all 18
python scripts/tts_mlx.py --video kv-cache                   # just one

npm run dev                      # Remotion Studio, preview with audio
npm run render                   # out/<id>.mp4, subtitles burned in

Edit the script

src/narration.json is the single source of truth: narration text, scene order and the visual each scene uses.

Re-run TTS, re-render. Subtitles regenerate from the same text.

Pacing and pauses

Each scene gets a 0.35s lead-in and a 0.7s tail (1.1s on titles and recaps), and speed is 0.96.

Change them in scripts/tts_*.py if you want it brisker.

Subtitles

Sentence-level cues in a fixed 200px band below the content area, timed by character count across each scene's audio.

The band is reserved, so captions can never cover a diagram.

Add a video

Add an entry to narration.json; a composition appears automatically in the Studio.

New visuals go in src/components/Visuals.tsx and are referenced by name.

Licence note. Remotion is free for individuals and companies of up to three people, including commercial use; beyond that it needs a licence, and rendering in CI counts as an automation. Kokoro-82M is Apache-2.0.
Coverage

Concept → video → lab → what you say

ConceptVideoLabThe sentence that shows you've run it
Prefill vs decodeinference-101F1"TTFT is prefill plus queue; ITL is decode. Which one is the customer complaining about?"
KV cache sizingkv-cacheL5"Weights fit. A single 128K-token request needs ~40 GB of KV on top — that's the real limit."
PagedAttentionkv-cacheL5 · L9"Fixed-size blocks and a block table, like OS paging — 2–4× more concurrent requests."
Memory bandwidthgpu-bandwidthL4"Decode ceiling is bandwidth over bytes per token. On a Spark that's 273 over the model size."
Roofline / batchinggpu-bandwidthL3"Batch 1 is memory-bound; I sweep concurrency and report the curve, not one number."
QuantizationquantizationL4"FP8 is close to free on Hopper and Blackwell; 4-bit needs an eval before I'd ship it."
Prefix / prompt cachingprefix-cachingF2 · L6"Their system prompt is 6K tokens on every call — caching removes most of the prefill cost."
Speculative decodingbatchingF7"Same output distribution, fewer big-model passes — and a mismatched drafter makes it slower."
SFT / LoRAtrainingF5 · L8"One base deployment, many adapters — a fine-tune per customer at base-model economics."
Full fine-tuningfull-finetuningT2"Sixteen bytes a parameter to train against two to serve. I right-size and try LoRA before I go full."
PEFT familypeftT1 · T2"QLoRA when memory is the limit, DoRA when there's a quality gap, multi-LoRA when there are many tenants."
RLHF & reward hackingrlhfT1"A reward model is a proxy; the KL leash stops the policy gaming it. Checkable tasks get a grader instead."
DPOdpoT3"Same pairs as RLHF, two models instead of four, no RL loop. SFT first, then DPO, and I watch accuracy."
DPO vs RFTtrainingF6"Can you write the answer, compare two, or score one automatically? That picks the method."
MoE & active paramsoffloadingL7"Memory is sized on total parameters; speed on active ones. That's why 120B beats dense 30B here."
OffloadingoffloadingL7"Offload is the last resort — you're trading a 273 GB/s pipe for a far narrower one."
Tensor vs pipeline paralleloffloadingL9"Tensor parallel syncs every layer, so it stays inside a box. Pipeline tolerates slower links."
Serving engine internalsserving-stackL2 · L9"API, scheduler, KV manager, kernels — and containers keep the CUDA stack honest."
Benchmark methodbenchmarkingL3"Traffic profile first, then a concurrency curve, then where p95 crosses the target."
Agent prefix reuseprefix-cachingL6 · F2"Stable content first, variable last — otherwise there is nothing to cache."
QLoRA mechanicsqloraL8"Frozen four-bit base, small adapter. The dataset and the eval set decide the result."
Multi-node realityscale-outL9"Two boxes add capacity, not bandwidth. 235B across two decodes at about 12 tok/s."
Platform economicsplatformF3 · capstone"Serverless until utilisation justifies dedicated; batch for evals; specialise to cut unit cost."
Schedule

Twelve days, about two hours a day

Days 1–10 cover serving, benchmarking and the capstone. Days 11–12 are the training track (videos 14–18, labs T1–T3), for when you want depth on how models are trained.

Say it well

Talk tracks and questions

Two registers, same content

The VP answer and the engineer answer

QuestionTo a VPTo an ML engineer
"Why are we slow?""Two different problems wear the same word. The wait before the answer starts is one fix; the speed it types at is another. Yours is the first, and it's mostly repeated work we can cache.""p95 TTFT is 2.8s with a 6K static prefix and no prefix caching. Enable it, add chunked prefill, and re-run the same traffic replay."
"Should we fine-tune?""Only if you can measure what 'better' means. Give me 100 examples you'd defend in front of a customer and I'll tell you in a week — with the cost.""SFT for format and domain, DPO if you only have preferences, RFT if the grader is programmatic. LoRA first; full-parameter only if three tests say so."
"Why not self-host vLLM?""You can, and some teams should. The question is whether serving is a differentiator for you or a tax.""vLLM is excellent. You'd own kernel upgrades, quantization tuning, autoscaling, multi-LoRA batching and the on-call. Here's what that cost us in hours last quarter."
"Can we keep data in-region?""Yes — the engine can run inside your own cloud account, so data never leaves your environment.""BYOC/Virtual Cloud deployment in your VPC; region-pinned dedicated deployments are also possible at a 1.5× rate."
"How do we cut cost 3×?""Three levers, in order: stop paying for repeated context, move routine traffic to a smaller model, and specialise that model on your data.""Prompt caching, task-level routing, then a LoRA on a 4–16B base. Measure each with the same eval set and report $/1K tasks."

Ask them

  • Where does adaptive speculative decoding pay off most, and where does the drafter mismatch bite?
  • How do field engineers get fixes into the platform — what's the path from a customer pattern to a shipped feature?
  • With Virtual Cloud and the Foundry partnership, how do you decide when an account belongs in BYOC?
  • What constraints matter most when a customer chooses a deployment model?
  • How do you keep evaluation honest when customers want a benchmark number instead of a task metric?

Prepare for these

  • "Walk me through what happens between the API call and the first token."
  • "A customer's 70B deployment OOMs under load. Debug it out loud."
  • "Size a deployment for 20 requests/second, 1K in and 300 out, p95 TTFT under one second."
  • "They want to move off a closed model. What's your first week?"
  • "Their fine-tune scored better offline and worse in production. What happened?"
Check the work

Sources & caveats

Everything here was read on 23 September 2026. Fireworks' catalogue and prices change weekly; NVIDIA's playbooks pin container tags that move. Verify anything you plan to quote.

Known ambiguities, worth knowing before you assert them: the pricing page says fine-tuned models serve "at the same price as base models", while the LoRA docs say trained LoRAs deploy only to on-demand deployments — reconcile those before promising a customer serverless economics for a fine-tune. Published rate limits (a flat 6,000 RPM) also conflict with a recent changelog entry describing adaptive limits. And DGX Spark bandwidth is quoted as 273 GB/s on the datasheet, though NVIDIA's Hot Chips material implies the silicon supports more.
  1. Fireworks docs: quickstart
  2. Fireworks docs: serverless pricing
  3. Fireworks: pricing (GPU-hour, training)
  4. Fireworks docs: on-demand deployments
  5. Fireworks docs: Batch API
  6. Fireworks docs: supervised fine-tuning & dataset format
  7. Fireworks docs: DPO fine-tuning
  8. Fireworks docs: reinforcement fine-tuning
  9. Fireworks docs: deploying LoRAs (serverless not supported)
  10. Fireworks docs: prompt caching
  11. Fireworks docs: speculative decoding
  12. Fireworks docs: structured responses
  13. Fireworks docs: quotas & spend limits
  14. Fireworks docs: firectl
  15. Fireworks blog: Frontier RL is cheaper than you think
  16. Fireworks blog: the fine-tuning bottleneck isn't the algorithm
  17. Fireworks blog: open-source agents with frontier advisors
  18. Fireworks blog: Training API GA
  19. Fireworks blog: Eval Protocol
  20. Fireworks blog: FireAttention V4 (NVFP4 on B200)
  21. Fireworks blog: Virtual Cloud / BYOC
  22. NVIDIA: DGX Spark hardware specs
  23. NVIDIA: DGX Spark performance measurements
  24. NVIDIA: DGX Spark playbooks (vLLM, SGLang, TRT-LLM, Unsloth…)
  25. LMSYS: SGLang on DGX Spark
  26. llama.cpp: DGX Spark benchmark thread
  27. Unsloth: fine-tuning on DGX Spark
  28. vLLM recipes: DGX Spark GB10
  29. Remotion docs
  30. Kokoro-82M (Apache-2.0)
  31. mlx-audio — Kokoro on Apple silicon
  32. mlx-lm — local serving and LoRA on Apple silicon
  33. Use the linked vendor documentation for current specifications and prices.