Skip to content

Serving With vLLM

file 09 · 2:56 · subtitles burned in

Chapters

In this video What a server program for AI models does and the five settings that matter most, shown on vLLM, a server built for NVIDIA chips.

What it explains: what an inference server does, containers, key flags

3 key points

  1. A server program for AI models does 4 jobs. It takes requests, groups them to share the chip, keeps each conversation’s notes (the model’s memory of it so far), and runs the math fast.

    Take one away and you have a demo, not a service. vLLM became the default because of the middle two jobs. People join and leave the group at every step (continuous batching), and notes are handed out a page at a time (PagedAttention).

  2. On NVIDIA AI machines such as the DGX Spark (a desktop AI computer), run the server inside a container: a sealed package with the exact software versions it needs.

    The chip’s driver, NVIDIA’s CUDA software and the Python packages must all match, and often do not. Installing without a container often needs a documented workaround, which costs 20 to 30% of total output.

  3. Five settings decide most of a server’s performance: how fast it answers and how many people it serves. Three of them took the same machine from 16 conversations at once to 50.

    The three: reuse a shared prompt opening (prefix caching), read long prompts in pieces (chunked prefill), and store the notes in 8 bits (FP8 KV cache). Only after that, the video says, argue about hardware.

The title card in the video says “Video 9 of 18”: that is the file order. The course plays the videos in the order of its days.

Download mp4 (11.5 MB)