Fireworks AI
A managed inference and fine-tuning cloud for open and custom models. You call an API, pay per token, and Fireworks runs the engine, the GPUs and the scaling.
- Founded
- 2022, Redwood City. CEO Lin Qiao, previously led PyTorch at Meta, with six co-founders.
- Funding
- Series D of $1.5B at a $17.5B valuation (July 2026). Series C of $250M at $4B (Oct 2025).
- Scale
- Reported >$1B annualised revenue; tens of trillions of tokens served per day.
- Customers
- Cursor, Notion, Perplexity, Sourcegraph, Uber, DoorDash, Shopify and others.
- Serverless & on-demand: OpenAI-compatible API per token, or dedicated GPUs by the hour.
- FireAttention: their own CUDA attention kernels and quantization (FP8/FP4).
- FireOptimizer: adaptive speculative decoding tuned to each customer's traffic.
- Post-training: SFT, LoRA, reinforcement fine-tuning, and serving many LoRA adapters on one base model.
- Compound AI: function calling, JSON/grammar-constrained output, agent workflows.
The pitch: "Own your specialized model. Fast, cheap, and your data stays yours." They say most tokens they serve come from customer-specialized models.