Serving With vLLM
file 09 · 2:56 · subtitles burned in
Chapters
In this video What a server program for AI models does and the five settings that matter most, shown on vLLM, a server built for NVIDIA chips.
What it explains: what an inference server does, containers, key flags
3 key points
A server program for AI models does 4 jobs. It takes requests, groups them to share the chip, keeps each conversation’s notes (the model’s memory of it so far), and runs the math fast.
Take one away and you have a demo, not a service. vLLM became the default because of the middle two jobs. People join and leave the group at every step (continuous batching), and notes are handed out a page at a time (PagedAttention).
On NVIDIA AI machines such as the DGX Spark (a desktop AI computer), run the server inside a container: a sealed package with the exact software versions it needs.
The chip’s driver, NVIDIA’s CUDA software and the Python packages must all match, and often do not. Installing without a container often needs a documented workaround, which costs 20 to 30% of total output.
Five settings decide most of a server’s performance: how fast it answers and how many people it serves. Three of them took the same machine from 16 conversations at once to 50.
The three: reuse a shared prompt opening (prefix caching), read long prompts in pieces (chunked prefill), and store the notes in 8 bits (FP8 KV cache). Only after that, the video says, argue about hardware.
Used in the course
- Day 2 · Step 1 · Serve one model three ways
- Day 9 · Step 3 · The Spark track (rewatch, 0:48 to 1:48)
The title card in the video says “Video 9 of 18”: that is the file order. The course plays the videos in the order of its days.