The Platform Map
file 08 · 2:51 · subtitles burned in
Chapters
In this video What a managed AI platform runs for its customers, the four ways it sells computing, and the sums that decide between paying per token and renting a whole GPU.
What it explains: serving layers, capacity modes, maturity curve
3 key points
A managed platform runs everything from the GPUs up to the model, so the customer builds only the app. Its shared (serverless) models charge per token, not per hour of GPU time.
The 5 layers, bottom to top: GPUs and data centres; the serving engine (the program that runs the model); scaling and routing (adding copies as traffic grows, and sending each request to one); the model; then the app. A raw GPU cloud rents only the bottom layer and leaves the rest to you.
Four ways to buy computing: shared models billed per token, your own rented GPUs, batch for work that can wait, and GPUs reserved on a long contract.
Shared per-token models (serverless) suit traffic that comes in sudden peaks. Your own GPUs (dedicated) suit steady volume or strict speed. Batch, this step, costs about half of serverless because answers come later. Reserved GPUs cost less per hour but bill for the whole contract.
Do the sums before renting a GPU: most products start far below the point where a dedicated one pays.
The video’s app uses 20 million tokens a day: about $146 a month on serverless. One dedicated H100 (a top data-centre GPU) running all month costs $5,840, 40 times more.
Used in the course
- Day 6 · Step 2 · Batch API
- Day 10 · Step 1 · Sizing memo and 10-minute talk (rewatch, 1:18 to 1:59)
The title card in the video says “Video 8 of 18”: that is the file order. The course plays the videos in the order of its days.