Solutions Architect · Forward Deployed Engineer

From prompt to token, visually.

A study guide for building and explaining AI inference systems. It follows the LLM inference roadmap step by step, explains each idea in plain words, and ends with labs, customer scenarios and flashcards.

Fireworks AI · managed inference Runpod · GPU cloud Researched Sept 2026
Know the room

Two companies, two layers of the same stack

Both sell "run AI models on GPUs", but at different heights. Fireworks hands you an API and runs everything underneath. Runpod hands you the GPU and lets you run whatever you like on it. You should be able to say that in one sentence, in a customer conversation.

Fireworks AI

A managed inference and fine-tuning cloud for open and custom models. You call an API, pay per token, and Fireworks runs the engine, the GPUs and the scaling.

Founded
2022, Redwood City. CEO Lin Qiao, previously led PyTorch at Meta, with six co-founders.
Funding
Series D of $1.5B at a $17.5B valuation (July 2026). Series C of $250M at $4B (Oct 2025).
Scale
Reported >$1B annualised revenue; tens of trillions of tokens served per day.
Customers
Cursor, Notion, Perplexity, Sourcegraph, Uber, DoorDash, Shopify and others.
  • Serverless & on-demand: OpenAI-compatible API per token, or dedicated GPUs by the hour.
  • FireAttention: their own CUDA attention kernels and quantization (FP8/FP4).
  • FireOptimizer: adaptive speculative decoding tuned to each customer's traffic.
  • Post-training: SFT, LoRA, reinforcement fine-tuning, and serving many LoRA adapters on one base model.
  • Compound AI: function calling, JSON/grammar-constrained output, agent workflows.

The pitch: "Own your specialized model. Fast, cheap, and your data stays yours." They say most tokens they serve come from customer-specialized models.

Runpod

A developer-first GPU cloud. Rent a GPU by the second, deploy a serverless endpoint that scales to zero, or spin up a multi-node cluster.

Founded
2022. CEO Zhen Lu and CTO Pardeep Singh. Remote-first, with people in APAC.
Funding
$100M led by Summit Partners at a $1B valuation (June 2026). Reportedly turned down $500M+ buyout offers.
Scale
Reported ~$240M annualised revenue and 1M+ developers.
Customers
Indie developers up to enterprises spending millions a year.
  • Pods: dedicated GPU containers (H100, H200, B200, A100, RTX 4090 and more). SSH, Jupyter, your own stack.
  • Secure vs Community Cloud: vetted data centres vs a cheaper marketplace of hosts.
  • Serverless: autoscaling workers, per-second billing, scale to zero, FlashBoot for faster cold starts, an official vLLM worker.
  • Instant Clusters: multi-node GPUs with fast interconnect for models too big for one box.
  • Network volumes & templates: persistent storage and one-click vLLM, SGLang and ComfyUI images.

The pitch: "Cheapest, fastest path to a GPU. No sales calls, no lock-in, pay by the second, free egress."

Who runs what

The inference stack, layer by layer

Layer
Fireworks serverless
Runpod Serverless
Runpod Pod
Your appprompts, RAG, agents
you
you
you
Model choice & tuningwhich model, LoRA, evals
shared*
you
you
Serving enginevLLM / SGLang / FireAttention
Fireworks
you (template)
you
Batching & kernelsquantization, spec decoding
Fireworks
you (engine config)
you
Autoscalingworkers up and down
Fireworks
Runpod
you
GPUs & data centrehardware, network, power
Fireworks
Runpod
Runpod
Fireworks runs itRunpod runs itThe customer runs it* pick from their library or bring your own weights; fine-tuning is a managed service
Plain words

Fireworks is a restaurant. You order from the menu (or bring a family recipe), and the kitchen, chefs and ovens are theirs. You pay per plate.

Runpod rents you a professional kitchen by the second. You bring the recipe and the chef (vLLM, SGLang), and you get the ovens cheaper. Serverless is a kitchen that switches its lights off when nobody is ordering.

CompetitorWhat it isHow to position against it
Together AIOpen-model inference and training on its own clustersClosest like-for-like for Fireworks: compare latency, fine-tuning depth, enterprise controls
BasetenManaged inference with strong developer toolingFireworks rival on dedicated deployments; Runpod is cheaper if the customer can run their own engine
ModalServerless GPUs for any Python functionDirect Runpod Serverless rival: compare cold starts, price per second, GPU choice
Groq · CerebrasCustom silicon, very high tokens per secondFastest raw decode, but narrower model catalogues and limited fine-tuning
CoreWeave · Lambda · Vast.aiGPU clouds and marketplacesRunpod's direct rivals: CoreWeave for big enterprise HPC, Vast for bargain marketplace
AWS · GCP · AzureHyperscaler managed AI (Bedrock, Vertex, SageMaker)Both: faster access to new open models, better price-performance, less lock-in
Roadmap step 1 Foundations

What happens when you make an LLM call

Every request has two phases. Prefill reads your whole prompt at once. Decode writes the answer one token at a time. Almost every optimization in this guide targets one of the two.

Walkthrough · tap any box

The life of one request

Interactive

Play a request

Prompt tokens (each square ≈ many tokens)
Output tokens
prefill
decode
0 s–
TTFT
–
ITL / TPOT
–
Total time
–
Share in decode
–

Illustrative numbers for a 70B model on one modern GPU: prefill ≈ 8,000 tokens/s, decode ≈ 22 ms per token. Try a long prompt with a short answer (RAG, classification), then the reverse (writing, code).

Interactive · toy tokenizer

What the model actually sees

Characters
–
Words
–
Tokens
–
Chars per token
–
At $0.90 / 1M tokens
–

A toy splitter for illustration. Real tokenizers (BPE) learn their splits from data, but the idea is the same: long or rare words break into pieces. A · marks a leading space, and the small numbers stand in for token IDs.

Plain words

Prefill is a chef reading the whole order ticket in one glance. Decode is plating dishes one at a time. The KV cache is the chopped ingredients kept on the counter, so the chef never re-chops for the next dish. TTFT is how long until the first bite arrives, and ITL is the rhythm of the rest of the meal.

Tokenization

Text is cut into sub-word pieces with numeric IDs. About ¾ of an English word per token.

Snapping a sentence into LEGO bricks.

Customers pay per token, so tokens are both the cost and the latency unit.

Forward pass

One trip of the input through every layer of the model, producing next-token probabilities.

One run down an assembly line.

In decode, every output token costs one forward pass.

Autoregressive

Generate a token, append it, feed everything back in, repeat.

Writing one word at a time, re-reading the sentence before each word.

Why long answers are slow: the steps can't run in parallel.

KV cache

Stored key and value vectors for every past token, so they aren't recomputed each step.

Mise en place on the counter.

Grows with context × users. After weights, it's the biggest user of GPU memory.

TTFT

Time to first token. Roughly queue time plus prefill time.

How long until the waiter brings the first bite.

Drives how responsive chat and agents feel.

ITL / TPOT

Inter-token latency, or time per output token, in decode.

How smoothly words stream onto the screen.

Voice agents and live coding need low ITL; 20–40 ms feels fluid.

Throughput vs latency

Total tokens/s for all users vs speed for one user.

Cars per hour on the highway vs your own trip time.

Bigger batches raise throughput but can slow each user. This is the core serving trade-off.

Goodput

Throughput that still meets the latency target (e.g. p95 TTFT < 500 ms).

Deliveries that arrive on time, not just deliveries made.

The number to agree with a customer before a benchmark.

Roadmap step 2 Transformer basics

Attention: every token decides who to listen to

A transformer is a stack of identical blocks. In each block, self-attention lets every token look back at the earlier tokens and pull in what's relevant. Tap a word below to make it the one asking.

Interactive

Who does "it" refer to?

query: it

The copper bar under each word is its attention weight. Later words are greyed out: a decoder only looks backwards (a "causal mask").

QQuery: "what am I looking for?" The current word's question.
KKey: "what do I contain?" A label on every earlier word.
VValue: the actual information passed along once Q matches K.
Plain words

A library. Your query is what you type into the search box. The keys are the titles on every spine. The values are what's inside the books. Attention compares your query with every title and blends the contents of the best matches.

The KV cache exists because the keys and values of earlier words never change, so there's no need to re-shelve the library every time you add a word.

Embeddings

Each token becomes a vector: a long list of numbers encoding meaning.

GPS coordinates for words; similar words live nearby.

Also sold as embedding endpoints for search and RAG.

Transformer block

Attention, then a feed-forward network, with normalization. Stacked 32–126 times.

One floor of an office tower; the model is the whole tower.

Layer count sets KV-cache size and compute per token.

Multi-head & GQA

Many attention "heads" run in parallel. Grouped-query attention shares K/V across heads.

Several readers skimming for different things, sharing one card catalogue.

GQA shrinks the KV cache 4–8×. It's why Llama 3 70B needs only ~0.3 MB per token.

Context window

The maximum number of tokens the model can attend over (8K to 1M+).

How big a desk the reader has.

Long context = large KV cache = fewer users per GPU.

Roadmap step 3 GPU & hardware

Inference is mostly a memory problem

A GPU has huge compute (FLOPS) but it can only use it if data arrives fast enough from memory (bandwidth). In decode, every token needs all the weights read once. Watch one token go through the model, then switch to 32 tokens.

Walkthrough · one pass through the modeldrag to tilt

Step 1

Fetching weights
0.0 s
Computing
0.0 s
Cores busy –Tokens finished 0 / 1Slowed down, not to scale: on a real H100 at batch 1, fetching takes over 100× longer than computing.
Interactive · roofline

Where does your workload sit?

memory-bound
Arithmetic intensity
–
Achievable
–
GPU busy
–
70B tokens/s (all users)
–

Simplified: counts weight reads only (≈2 FLOPs per parameter per token) and ignores KV-cache reads, which grow with context. Peak numbers are dense vendor specs. The lesson holds: batching is how you buy back idle compute.

Plain words

The cooks (SMs) chop incredibly fast, but ingredients come from a fridge down the hall (HBM) through one pipe (bandwidth). The cutting board in front of them (SRAM) is tiny but instant. Cooking one plate at a time, the cooks mostly wait for the pipe. Cook 200 plates from the same trip to the fridge, and now the cooks are the bottleneck. That's batching.

GPUMemoryBandwidthBF16 denseFP8 denseWhy a customer picks it
RTX 409024 GB1.0 TB/s~165 TF~330 TFCheap dev and small models (Runpod Community Cloud)
L40S48 GB0.86 TB/s~362 TF~733 TFImage/video models, mid-size LLMs on a budget
A100 80GB80 GB2.0 TB/s312 TFno FP8Proven workhorse, cheaper per hour; use INT8/AWQ instead of FP8
H100 SXM80 GB3.35 TB/s989 TF1,979 TFDefault for production LLM serving; FP8 tensor cores
H200141 GB4.8 TB/s989 TF1,979 TFSame compute as H100, more and faster memory: 70B on one GPU, long context
B200192 GB8.0 TB/s~2,250 TF~4,500 TFNewest NVIDIA; adds FP4. Biggest models, highest throughput
MI300X192 GB5.3 TB/s~1,307 TF~2,615 TFAMD alternative with lots of memory; ROCm software
Explain the tradeoff this way: "H100 and H200 have the same compute. The H200 only wins on memory: 141 GB at 4.8 TB/s instead of 80 GB at 3.35. So for decode-heavy traffic, or a 70B that won't fit in 80 GB, it's the better buy, even though its FLOPS are identical."

SM

Streaming multiprocessor: one of 100+ parallel compute units on the GPU, each with tensor cores.

One cook in a very large kitchen.

Why GPUs beat CPUs at the matrix multiplies inside every layer.

HBM vs SRAM

HBM: large, fast off-chip memory (the "VRAM"). SRAM: tiny, much faster on-chip memory.

The fridge vs the cutting board.

FlashAttention's whole trick is fewer fridge trips.

Arithmetic intensity

FLOPs done per byte moved. Below the "ridge point" you're memory-bound; above it, compute-bound.

Plates cooked per trip to the fridge.

Decode at batch 1 ≈ 1–2 FLOP/byte; H100's ridge ≈ 300. Huge headroom.

NVLink / InfiniBand

NVLink joins GPUs inside a server (~900 GB/s on H100). InfiniBand joins servers.

Express conveyors between kitchens.

Needed when a model is split across GPUs (see step 6).

Roadmap step 4 Optimization techniques

How modern systems get faster and cheaper

Four ideas carry most of the weight: manage KV memory like an operating system, never let the GPU idle, guess ahead cheaply, and shrink the numbers. Each has a live demo.

Simulation · PagedAttention

Same GPU memory, same requests, two allocators

Reserve max length up frontnaive

Allocate small pages on demandpaged

Each colour is one request's KV cache; each square is one block of 16 tokens. Hatched squares are reserved but never used. vLLM's PagedAttention paper reports 2–4× more throughput from this alone.

Plain words

The naive way is a restaurant that reserves a table for ten for every party, just in case. PagedAttention seats people at small tables as they arrive and adds a table when the party grows. A seating chart (the "block table") keeps track of who sits where.

Simulation · continuous batching

Four GPU slots, twelve requests of different lengths

Static · all done at
–
Static · GPU busy
–
Continuous · all done at
–
Continuous · GPU busy
–
Plain words

Static batching is a tour bus: nobody leaves until the whole group has finished, and nobody new boards mid-tour. Continuous batching is a taxi rank: the moment one passenger gets out, the next one gets in. Same cars, more trips.

Step-through · speculative decoding

A small model drafts, the big model checks

This round: draft proposes → target verifies in one pass
Press "Next round".
Output so far
Big-model passes
0
Tokens produced
0
Tokens per pass (this run)
–
Expected per pass
–
draft accepteddraft rejectedtoken from the big model

Expected tokens per big-model pass = (1 − αk+1) / (1 − α). Output quality is identical: the big model checks every token. It pays off when the draft matches your traffic, which is why Fireworks trains drafters on each customer's data (Cursor reported up to 2× lower latency).

Plain words

A junior writer drafts the next four words; the senior editor reads all four at once and ticks the ones that are right. Checking four words takes the editor about as long as writing one, so every correct guess is time saved. When the junior guesses wrong, the editor writes that word and the junior starts again from there.

Interactive · quantization

One weight in 16, 8 and 4 bits

π stored as
–
70B model weights
–
Fits one H100 (80 GB)?
–
H100 · 80 GB2 × H100

FlashAttention

A fused attention kernel that works in tiles inside SRAM and never writes the full attention matrix to HBM.

Chopping on the board without walking to the fridge between cuts.

Fireworks' FireAttention is their own take on this idea.

Chunked prefill

Split a long prompt's prefill into chunks and interleave them with other users' decode steps.

Prepping a banquet order in batches so regular diners still get served.

Stops one 100K-token prompt from freezing everyone's streaming.

Prompt / prefix caching

Keep the KV cache for a shared prefix (system prompt, documents, chat history) and reuse it.

Keeping a pot of base sauce ready instead of starting from scratch.

Agents and RAG repeat huge prefixes. Fireworks bills cached input tokens at a discount.

Quantization

Store weights (and sometimes KV) in fewer bits: FP8, INT8, INT4 (AWQ, GPTQ), FP4.

Vacuum-packing food: same meal, half the freezer space.

FP8 on H100+ halves memory and roughly doubles decode speed with minimal quality loss. Always run evals.

KV-cache compression

FP8 KV cache, GQA/MLA attention, or evicting old tokens.

Smaller containers double how much fits in the fridge.

Halving KV size roughly doubles concurrent users.

Disaggregated serving

Run prefill and decode on separate GPU pools and ship the KV cache between them.

A prep kitchen and a plating line that never block each other.

Tunes TTFT and ITL independently at large scale.

Roadmap step 5 Engines & frameworks

The frameworks, and what each platform runs

A serving framework (or "engine") is the software that loads model weights onto GPUs and turns requests into tokens. The key difference between the two companies: on Runpod you choose the engine (anything that runs in a Docker container), and on Fireworks the engine is Fireworks' own. You bring weights; they serve them.

FrameworkWhat it isOn RunpodOn Fireworks
vLLMThe default open-source LLM server. PagedAttention, continuous batching, OpenAI-compatible API.Official worker worker-vllm in the Serverless Hub; also Pod templates.Not used Fireworks runs its own engine. Know vLLM because customers compare you against self-hosting it.
SGLangFast server that reuses shared prompt prefixes (RadixAttention); strong at agents and JSON output.Official worker worker-sglang.Not used Prefix reuse is offered as prompt caching instead.
OllamaOne-command local runner built on llama.cpp; pulls quantized models from a library.Pod tutorial PyTorch template, expose port 11434, set OLLAMA_HOST=0.0.0.0. Community serverless workers exist.Not applicable Local/dev tool. A common starting point before a customer moves to a hosted API.
llama.cppC/C++ engine for GGUF quantized models on CPU, Apple Metal, CUDA.Bring your own container on a Pod or custom worker.Not applicable
TensorRT-LLM / TritonNVIDIA's compiled engine plus inference server; top throughput for a fixed model.Bring your own container NVIDIA images on Pods or a custom Serverless worker.Not used
TGIHugging Face's server, now in maintenance mode.Archived worker worker-tgi. Steer new builds to vLLM or SGLang.Not used
ComfyUI · SDXLImage-generation pipelines.Official workers worker-comfyui, worker-sdxl.Managed API Image models served via API.
faster-whisperSpeech-to-text.Official worker worker-faster_whisper.Managed API Audio transcription models via API.
Infinity embeddingsEmbedding server for search and RAG.Official worker worker-infinity-embedding.Managed API Embedding models via API.
Axolotl (fine-tuning)Open-source fine-tuning toolkit.Official worker llm-fine-tuning; or any Pod.Managed service SFT, LoRA and reinforcement fine-tuning without running a trainer.
Your own weightsA Hugging Face model or fine-tune you own.Yes Point the vLLM worker's MODEL_NAME at it, or bake it into your image.Yes firectl model create uploads safetensors + config; LoRA adapters deploy to on-demand deployments.
green first-party / officialamber supported with setup, or managed APIgrey archiveddash not how that platform works
Map

Where each framework fits

Positions are qualitative. Read it as: the more users and GPUs, the further up and right you go, and the more batching and memory management matter. Tap a bubble to open that framework below.

Plain words

Ollama is a home espresso machine: one button, great coffee, one cup at a time. vLLM and SGLang are café machines built to pour hundreds of cups an hour. TensorRT-LLM is a factory line tuned for one drink. Fireworks is a coffee chain: you never touch the machine, you just order.

Explorer · tap any box

How each framework works inside

Best for

On Runpod

On Fireworks

Run it
Interactive · SGLang RadixAttention / prompt caching

Shared prefixes are computed once

Prompt tokens sent
0
Actually computed
0
Reused from cache
0
Prefill saved
–

Send requests in any order. A segment turns copper when it's computed for the first time and green when a later request reuses it. Agents repeat the same system prompt and tool definitions on every call, which is why this matters so much for them. Fireworks exposes the same idea as prompt caching (cheaper cached input tokens); on Runpod you get it from SGLang or from vLLM's prefix caching.

Deploy flows

From zero to an API call on each platform

    
            

    Interactive

    Engine chooser

    Where does it run?
    What's the traffic like?

    Roadmap step 6 For mid/senior: scale out

    When one GPU isn't enough

    A 405B model needs ~810 GB just for FP16 weights, more than eight H100s hold. There are four ways to split work across GPUs, and real deployments combine them. Pick a strategy: each walkthrough shows the model being split, one request travelling through, and the trade-off at the end.

    Data paralleldrag to tilt

    Step 1

    MoE models

    Mixture of experts: only a few "expert" sub-networks run per token (e.g. Mixtral, DeepSeek, Llama 4).

    A hospital that sends each patient to two specialists, not every doctor.

    Memory is sized by total parameters; speed by active parameters.

    LoRA & multi-LoRA

    Tiny trainable adapters on a frozen base model. Many adapters can share one base on one GPU.

    One base recipe, a different spice packet per customer.

    Fireworks serves hundreds of fine-tunes cheaply this way.

    Autoscaling & cold starts

    Add workers as traffic rises; the first request on a new worker waits for the model to load.

    Opening more checkout lanes, but staff take time to arrive.

    Runpod FlashBoot snapshots help; a warm minimum worker is the reliable fix.

    Cost per million tokens

    All-in serving cost divided by tokens actually served.

    Cost per plate, including rent and the idle hours.

    The number the customer's CFO cares about. See the lab below.

    Practice

    Sizing & cost lab

    The most common whiteboard question for this role: "How many GPUs does this need, and what will it cost?" Change the inputs and watch the method, not just the answer.

    Calculator

    Will it fit? GPU memory sizing

    fits
    weights
    KV
    Weights
    –
    KV cache
    –
    Total w/ overhead
    –
    GPUs needed
    –
    Max requests at this context
    –
    Single-user decode ceiling
    –

    Overhead of 10% covers activations, CUDA context and fragmentation; engines typically use ~90% of GPU memory. GPU count rounds up to 1, 2, 4 or 8 for tensor parallelism. The decode ceiling is bandwidth ÷ bytes read per token, an upper bound for one user.

    Calculator

    Serverless per-token vs dedicated GPUs

    serverless cheaper
    Serverless / month
    –
    Dedicated / month
    –
    Break-even volume
    –
    Dedicated utilisation
    –

    Example defaults: a 70B-class model at ~$0.90 per 1M tokens serverless vs two H100s at a Runpod-like $2.89/hour serving ~2,500 tok/s with batching. Prices move often: check live pricing before quoting. Runpod's own Serverless-vs-Pod rule of thumb is the same idea: break-even utilisation ≈ Pod $/hr ÷ Serverless $/hr (≈ 60% on H100).

    Practice

    Customer scenarios, worked through

    These show up in role-plays and technical rounds. Use the same shape every time: diagnose with questions and metrics, fix in order of cheapest first, measure before and after.

    "Our chatbot takes 3 seconds before anything appears."

    +
    Diagnose
    • That's TTFT, so look at prefill and queueing, not decode.
    • How long are prompts? Same system prompt every call?
    • Is it cold starts (first request after idle) or every request?
    • p50 vs p95: spiky means queueing or noisy neighbours.
    Fix
    • Prefix caching for the repeated system prompt and history.
    • Trim RAG context; retrieve fewer, better chunks.
    • Chunked prefill so long prompts don't block others.
    • Keep a minimum warm worker; pick a priority/fast tier.
    Measure
    • p50/p95 TTFT before and after, same traffic replay.
    • Cache hit rate.
    • Cost per 1K conversations.
    One-liner"TTFT is almost all prefill and queue time. Your system prompt is 6K tokens and identical on every call, so prompt caching should remove most of it. Let's replay yesterday's traffic and compare p95."

    "We want to move off OpenAI to open models to cut cost."

    +
    Diagnose
    • What tasks? Chat, extraction, code, agents with tools?
    • Current volume, latency targets, monthly bill.
    • What does "good enough" mean? Is there an eval set?
    Fix
    • Both platforms are OpenAI-compatible: swap base URL and key.
    • Shortlist 2–3 open models; run the customer's eval.
    • Close quality gaps with LoRA / fine-tuning on their data.
    • Use JSON / grammar mode for reliable structured output.
    • Send offline jobs to batch pricing.
    Measure
    • Eval score vs the incumbent model.
    • $ per 1M tokens and total cost of ownership.
    • p95 latency; tool-call success rate.
    One-liner"The API swap takes an afternoon; the real work is the eval. Give me 200 real examples and I'll show quality, latency and cost side by side for three models by Friday."

    "Size GPUs for a 70B model at 20 requests/second."

    +
    Diagnose
    • Average input and output tokens per request?
    • Latency target (p95 TTFT, ITL)?
    • Peak vs average traffic; can it queue?
    Fix
    • Say 1K in / 300 out: 20 rps × 300 = 6K output tok/s.
    • FP8 weights 70 GB; add KV for concurrency (lab above).
    • Per-replica throughput from a benchmark, e.g. ~2.5K tok/s on 2×H100.
    • 6K ÷ 2.5K → 3 replicas, +1 for headroom = 4 replicas (8 H100s), or fewer H200s.
    Measure
    • Load test at 1×, 1.5×, 2× peak.
    • Goodput at the latency target.
    • $ per 1M tokens per option.
    One-liner"Decode is memory-bound, so I size on bandwidth and KV headroom, not FLOPS. Here are three options with monthly cost. I'd validate the per-replica number with a one-hour benchmark before you commit."

    "Serverless or dedicated?"

    +
    Diagnose
    • Traffic shape: bursty, business hours, or flat 24/7?
    • Strictness of the latency SLA.
    • Monthly token volume and growth.
    Fix
    • Bursty or low volume → serverless / scale to zero.
    • Steady high volume or strict SLA → dedicated or reserved capacity.
    • Often both: dedicated baseline, serverless for spikes.
    Measure
    • Utilisation of dedicated capacity.
    • Cold-start rate on serverless.
    • Blended $ per 1M tokens.
    One-liner"It's a utilisation question. Above roughly 60% steady utilisation, dedicated wins. Below it, you're paying for idle GPUs."

    "Cold starts on Runpod Serverless are killing our p95."

    +
    Diagnose
    • Worker logs: time spent pulling image vs loading weights vs compiling.
    • Is the model downloaded at request time?
    • How often do workers scale to zero?
    Fix
    • Load the model at worker start, outside the handler.
    • Bake weights into the image or use a network volume.
    • Cache compiled kernels / CUDA graphs.
    • Enable FlashBoot; set min workers ≥ 1 for SLA traffic.
    Measure
    • Cold-start count and duration per day.
    • p95 end-to-end latency.
    • Extra cost of warm workers vs SLA value.
    One-liner"FlashBoot helps when a snapshot is warm, but a miss means a full reload. Moving the model load out of the handler and keeping one warm worker costs about $X a month and removes the tail."

    "The pod crashed with CUDA out of memory."

    +
    Diagnose
    • nvidia-smi: what's using memory?
    • Engine settings: max model length, max sequences, memory utilisation.
    • Did it fail at load (weights) or under load (KV cache)?
    Fix
    • At load: quantize (FP8/AWQ) or use a bigger GPU / tensor parallelism.
    • Under load: lower max context or max concurrent sequences; FP8 KV cache.
    • Leave headroom for other processes.
    Measure
    • Peak memory under a load test.
    • Max concurrency before latency degrades.
    One-liner"Weights fit fine. It's the KV cache: your max context is 128K, and a single request at that length needs about 40 GB. Cap it at 32K or move to FP8 KV and you'll get four times the users."
    Recall

    Flashcards

    Tap a card to flip it. Say the answer out loud before you flip; then check what you missed.

    Schedule

    A 7-day plan

    About two focused hours a day. Tick days off as you go; your progress stays in this browser.

    Last look before the call

    Cheat sheet

    Fireworks · Series D
    $1.5B @ $17.5B

    July 2026. Reported >$1B annualised revenue.

    Runpod · growth round
    $100M @ $1B

    June 2026, led by Summit Partners. ~$240M annualised revenue, 1M+ developers.

    Weights, 70B model
    140 / 70 / 35 GB

    FP16 / FP8 / INT4. Bytes per param × params.

    KV cache, Llama 3 70B
    ≈0.33 MB / token

    2 × 80 layers × 8 KV heads × 128 dim × 2 bytes. 128K context ≈ 43 GB.

    VRAM rule
    W + KV + 10%

    Then round GPUs up to 1, 2, 4 or 8.

    Decode ceiling
    BW ÷ bytes

    H100 3.35 TB/s ÷ 70 GB ≈ 48 tok/s for one user.

    H100 → H200
    same FLOPS

    80 → 141 GB, 3.35 → 4.8 TB/s. Memory is the upgrade.

    H100 ridge point
    ≈300 FLOP/B

    Batch-1 decode ≈ 1–2. That's why batching matters.

    Spec decoding
    (1−αᵏ⁺¹)/(1−α)

    Tokens per big-model pass. α 0.75, k 4 ≈ 3.1.

    Serverless vs dedicated
    ≈60% utilisation

    Break-even on Runpod H100: Pod $/hr ÷ Serverless $/hr.

    Engines
    vLLM default

    SGLang for shared prefixes/JSON, TensorRT-LLM for max NVIDIA perf, llama.cpp for edge. TGI in maintenance.

    Latency vocabulary
    TTFT · ITL · goodput

    Agree the target percentile (p95) before any benchmark.

    Check the work

    Sources

    Company figures are as reported in September 2026; performance claims are vendor benchmarks; prices change often. Re-check pricing pages before making a cost estimate.