Serverless
Per-token API, OpenAI-compatible, Standard / Priority / Fast tiers. Billed on input, cached input and output separately.
Default starting point. No provisioning, no idle cost.
On-demand deployments
Private replicas billed per GPU-second, autoscaling, scale-to-zero (1 hour idle by default — shorten it).
Needed for LoRA serving, custom shapes, speculative-decoding flags.
Batch API
JSONL in, JSONL out, 50% of serverless price, caching discounts stack. 12–72h windows.
Where your eval runs and bulk jobs belong.
Reserved capacity
Committed GPU capacity, lower hourly price, sold through sales.
Bills for the whole term whether used or not — never suggest it casually.
Managed Training
SFT, DPO, ORPO and RFT as managed jobs via UI, firectl or REST. Priced per 1M training tokens.
The "own your weights" story, and the cheapest real fine-tune you can run.
Training API
Custom training loops in Python — custom loss, distillation, inference in the loop. Serverless (per token) or dedicated (per GPU-hour) compute.
GA in Aug 2026. This is what "Fireworks Training" means now.
Eval Protocol
Their open-source, trace-first eval framework (pip install eval-protocol). Rule-based, LLM-judge or hybrid rewards; feeds RFT.
The bridge from "we think it's better" to a reward function.
Multi-LoRA
One base deployment, many adapters (default quota 100). Needs a BF16 shape with --enable-addons; FP8/FP4 shapes can't host adapters.
How per-customer fine-tunes stay affordable.
FireAttention
Their own kernels. V4 targets B200 with NVFP4 and claims 250+ tokens/sec.
The answer to "why not just run vLLM ourselves?"
Speculative decoding
Default model-based speculation, custom --draft-model, n-gram, or predicted outputs. Measure with perf_metrics_in_response.
Docs warn a bad drafter makes things slower — good nuance to voice.
Structured output
JSON mode and JSON-schema mode (2020-12, $ref/$defs, no external refs), plus grammar mode.
Two docs rules: put the schema in the prompt too; for reasoning models omit response_format.
Virtual Cloud / BYOC
The engine runs inside the customer's own VPC, across many clouds and regions.
The answer to data-residency objections in regulated APAC accounts.
Routing (Nexus · FireRouter)
Task-level, cache-aware routing across open and closed models, with centralized cost controls.
Reframes "which model?" as a portfolio question. Great discovery hook.