Skip to content

The Spark track

course 16 of 22

lesson 14 · lessons/14-spark-track

2 h 30 Free (your hardware) DGX Spark Optional: needs a DGX Spark

What you will do and why

Optional (needs a DGX Spark): you run the server software real services use, vLLM and SGLang, in containers on NVIDIA hardware. Your own test scripts from earlier days then measure it, so its numbers compare directly with your Mac’s.

Why it matters: If a team asks whether a small NVIDIA desktop can serve its internal users, test it with the same questions and timing script you used on your Mac. This step is optional because it needs a DGX Spark. Its published memory speed resembles an M4 Pro, but the measured reply speed and number of users it supports must be checked on the actual machine.

You are done when: vLLM and SGLang both answered from the Spark, 2_bench.sh printed a sweep with a capacity line, SGLang’s usage line shows reused tokens, and 9_teardown.sh freed the GPU. No Spark: read the page, answer the checks, then tap Skip.

Rewatch from Day 2: Serving With vLLM, 0:48 to 1:48ShowHide

The Mac taught the ideas. The Spark runs the production server software (the programs live services use): vLLM and SGLang on an NVIDIA GPU. You start each one in a container on the Spark and drive it from your Mac with the same scripts as before, so the numbers line up with your Mac’s.

Picture it

Same recipe, same measuring cups, but a professional kitchen. Because you measure the same way, any difference in the result comes from the kitchen, not from how you measured.

With real numberslesson 14 and the lab book’s DGX Spark page

  • DGX Spark: 128 GB of memory shared by its CPU and GPU, about 115 GB usable for the model and its notes.
  • Its memory moves 273 GB a second, like an M4 Pro Mac. Llama 3.1 8B at 8 bits is 8 GB, so one person’s limit is 273 ÷ 8 = 34 tokens a second.
  • The lab book’s measurement: 20.5 tokens a second for 1 person, 60% of that limit (20.5 ÷ 34).
  • With 32 people the same Spark wrote 368 tokens a second in total: 368 ÷ 20.5 = 18 times one person’s output.
  • An H100 moves 3,350 GB a second, 12 times the Spark: the Spark is for learning the server stack, not for data-center speed.

Words to know

DGX Spark
NVIDIA’s desktop AI computer, with 128 GB of memory shared by its CPU and GPU.
Container
A sealed package holding a program and the exact software versions it needs. Example: vLLM’s image on the Spark.
vLLM
An open-source server for NVIDIA GPUs built for many users at once.
SGLang
Another open-source server for NVIDIA GPUs, strong at reusing shared prompt openings.

From the lesson

lessons/14-spark-track/README.md

The Mac covers the concepts. The Spark gives you the production server stack the job description names: vLLM and SGLang on an NVIDIA GPU, in containers, benchmarked with the same harness you built in lessons 01, 03 and 06 (Days 1, 3 and 5). The DGX Spark (GB10, 128 GB unified memory, ~273 GB/s) has about the same bandwidth as an M4 Pro, so your ceiling maths carries straight over. The difference is the software: CUDA, FP8/NVFP4 and the engines customers run.

Mac ──(same python scripts, --target spark)──► Spark :8000 vLLM / :30000 SGLang
script where it runs does
remote.sh X.sh Mac runs any script below on the Spark over SSH
0_check.sh Spark GPU, driver, CUDA, memory, and whether Docker can see the GPU
1_vllm.sh [model] Spark vLLM container with prefix caching; prints the KV pool size
2_bench.sh Mac lessons 01, 03 and 06 (Days 1, 3 and 5) against the Spark, with Mac and Spark curves on one chart
2b_vllm_bench.sh Spark vllm bench serve as a cross-check on your own numbers
3_sglang.sh [model] Spark SGLang with --enable-cache-report; RadixAttention hits
9_teardown.sh Spark frees the GPU
TWO_BOX.md – fabric first (iperf, NCCL busbw), then TP=2 or PP=2
Terminal window
# .env: SPARK_IP=…, HF_TOKEN=… (for gated models)
cd lessons/14-spark-track
bash remote.sh 0_check.sh
bash remote.sh 1_vllm.sh nvidia/Llama-3.1-8B-Instruct-FP8
bash 2_bench.sh
bash remote.sh 2b_vllm_bench.sh
bash remote.sh 3_sglang.sh Qwen/Qwen3-8B
SPARK_URL=http://$SPARK_IP:30000/v1 SPARK_MODEL=Qwen/Qwen3-8B \
python ../06-prefix-caching/prefix_bench.py --target spark --usage
bash remote.sh 9_teardown.sh

What each command does

  1. bash remote.sh 0_check.sh

    Runs the check script on the Spark over SSH: your Mac logs in and runs it there. It needs SPARK_IP in 03-labs/.env and SSH keys set up. Look for the GPU name, driver and CUDA version, the memory, the chip type (arm64, printed as aarch64), and a GPU listed from inside a test container: proof that Docker can see the GPU.

  2. bash remote.sh 1_vllm.sh nvidia/Llama-3.1-8B-Instruct-FP8

    Starts vLLM in a container on the Spark with Llama 3.1 8B stored at 8 bits: conversations capped at 8,192 tokens, up to 80% of memory, prefix caching on (Day 5). The first run downloads the model. Look for ready at http://…:8000/v1, then the log lines about the KV cache and Maximum concurrency for 8192 tokens per request: Day 4’s notes calculation, done by the server.

  3. bash 2_bench.sh

    Runs on your Mac, not the Spark. It points three of your own scripts at the Spark: Day 1’s timing test (5 runs); Day 3’s sweep from 1 to 64 people with a 1,000 ms promise, which it then draws with your earlier sweeps on one chart; and Day 5’s shared-opening test. Look for the sweep’s Capacity line: and the chart in results/03-sweep.png.

  4. bash remote.sh 2b_vllm_bench.sh

    Runs vLLM’s own load tester inside the container: 64 made-up requests of 512 tokens in and 128 out, sent at 1, 4 and 16 requests a second, then all at once. Look for Output token throughput (total tokens written a second) and P99 TTFT (99 of 100 requests got their first token within this time): they should tell the same story as your own sweep.

  5. bash remote.sh 3_sglang.sh Qwen/Qwen3-8B

    Stops vLLM (one server at a time on one GPU) and starts SGLang with the Qwen3 8B model, set to report reused tokens (--enable-cache-report). Look for ready. From the Mac: followed by the exact command to run next, with the Spark’s address filled in.

  6. SPARK_URL=http://$SPARK_IP:30000/v1 SPARK_MODEL=Qwen/Qwen3-8B

    Day 5’s shared-opening test against SGLang on port 30000, plus the extra call that prints token counts. Your terminal fills in $SPARK_IP itself and does not read .env, so run export SPARK_IP=<the Spark's address> first, or paste the command 3_sglang.sh printed. Look for the usage: line, with cached_tokens close to prompt_tokens on the same line: nearly the whole opening (the script estimates about 3,028 tokens) was reused.

  7. bash remote.sh 9_teardown.sh

    Stops both servers and frees the GPU; downloaded models stay on the Spark for next time. Look for no vllm or sglang in the container list. The GPU memory line may read N/A: the Spark’s memory is shared, so the lab book says to use free -h there instead.

Container tags move quickly. If an image tag isn’t found, look up the current GB10/arm64 tag (links are in the script comments) and replace the tag after :- in the IMAGE= line near the top of 1_vllm.sh or 3_sglang.sh, in your copy on the Mac (remote.sh sends that file to the Spark). (remote.sh passes only HF_TOKEN to the Spark, so VLLM_IMAGE or SGLANG_IMAGE set on your Mac never arrives.)

How to read it

First, vLLM’s Maximum concurrency for 8192 tokens per request: …x line is Day 4’s notes calculation, done by the server: how many full 8,192-token conversations fit in its memory for notes. Second, the Spark’s sweep curve keeps rising well past your Mac’s: vLLM lets requests join at every step and hands out note memory in small pages. Third, SGLang’s usage line shows cached_tokens close to prompt_tokens: proof it reused the shared opening.

  • The vLLM log line Maximum concurrency for 8192 tokens per request: …x is lesson 05’s KV maths (Day 4) done by the engine.
  • The Spark’s sweep curve keeps rising well past the Mac’s, because vLLM’s continuous batching and paged KV cache handle many users far better than a laptop server.
  • The SGLang usage block shows cached_tokens close to the preamble size on the extra call the script makes after its 12 timed calls.

3 questions. Say your answer out loud, then tap to check it.

vLLM’s startup log says its memory for notes holds 18 full conversations of 8K tokens (8,192, about 6,000 words) at once: maximum concurrency 18x. How do you double that?Show answerHide

In plain words

Make each conversation’s notes smaller, or give the notes more memory. Store them in 8 bits instead of 16, lower the cap on conversation length, or use a smaller or more compressed model.

Picture it

A wardrobe holds 18 coats. Vacuum-pack each coat to half its size, or allow jackets instead of long coats, and 36 fit. Or take out the big suitcase that was using some of the space.

With real numbersDay 4’s calculator for Llama 3.1 8B, and the vLLM settings from the Serving With vLLM video

  • Notes for one 8,192-token conversation at 16 bits: 1.07 GB (Day 4).
  • Say the server has 20 GB for notes, as in Day 4’s example: 20 ÷ 1.07 = 18.6, so vLLM reports about 18.6x (18 whole conversations): the check’s 18x.
  • 8-bit notes (--kv-cache-dtype fp8): 1.07 ÷ 2 = 0.54 GB each, so 20 ÷ 0.54 = 37.
  • Or cap conversations at 4,096 tokens (--max-model-len 4096): also 0.54 GB each, also 37, if real requests fit in 4,096.
  • Or a smaller or more compressed model: its weights shrink, and vLLM gives the freed memory to notes.

Words to know

Maximum concurrency
vLLM’s count of how many full-length conversations fit in its memory for notes.
KV cache dtype
The number format the notes are stored in; 8-bit instead of 16-bit halves their memory.
Max model length (--max-model-len)
vLLM’s cap on tokens per conversation, which caps that conversation’s notes.
FP8
An 8-bit number format, 1 byte per number.
Go deeper: the engineer version

The kit's question

The vLLM log says max concurrency is 18× at 8k context. How do you double it?

The kit's answer

Use an FP8 KV cache, a lower --max-model-len or a smaller or quantized model.

More detail: vLLM’s figure is KV-cache tokens ÷ max_model_len. The Serving With vLLM video’s sample log: 8,234 blocks × 16 tokens ≈ 131,000 tokens, about 16 requests at 8K. FP8 KV halves bytes per token, so it doubles the figure exactly (the video: “roughly doubles concurrency”). Halving --max-model-len also doubles it, but is only safe if the real p95 fits (Day 4). A smaller or quantized model frees memory inside --gpu-memory-utilization for the KV pool, which raises the figure by however much it frees.

When would you pick SGLang over vLLM, the two main open-source servers for NVIDIA GPUs?Show answerHide

In plain words

When many requests share long openings or branch from one conversation, as agents do, or when most output must follow a strict format such as JSON (a common format for structured data). Either way, test both on the customer’s real traffic before choosing.

Picture it

Two delivery firms: one is best at many different parcels, the other at many parcels to the same street. Which is cheaper depends on your parcels, so you give each a trial week.

With real numbersDay 5’s prefix test and this step’s SGLang run

  • Day 5’s shared opening: about 3,028 tokens by the script’s estimate, sent with each of 12 questions.
  • Reusing it made the first word 20 times sooner in the lesson’s sample: 740 ms down to 36 ms (thousandths of a second).
  • SGLang keeps saved notes like a family tree (RadixAttention): one shared opening at the root, a branch for each follow-up, so every branch reuses the common start.
  • This step’s proof: SGLang’s usage: line shows cached_tokens close to prompt_tokens on the same line: nearly the whole opening was reused.

Words to know

SGLang
An open-source server for NVIDIA GPUs, strong at reusing shared prompt openings.
RadixAttention
SGLang’s store of saved notes, kept as a tree of shared openings.
Agent
A program that calls a model many times, using tools, to finish a task.
Structured generation
Making every reply follow a fixed format, such as JSON.
Go deeper: the engineer version

The kit's question

When would you pick SGLang over vLLM?

The kit's answer

For heavy shared-prefix or branching agent workloads and structured-generation-heavy pipelines. Benchmark both on the real traffic.

More detail: SGLang’s RadixAttention keeps the KV cache in a radix tree keyed by token prefixes and schedules requests to maximise prefix hits, which suits multi-turn and branching agent traffic; its constrained decoding is strong for JSON-heavy pipelines. vLLM’s automatic prefix caching also reuses identical block prefixes, so the gap depends on the traffic: benchmark both on the customer’s prompts, with their cache hit rates.

Why run vLLM and SGLang on the Spark only inside containers, never installed straight onto the machine with pip (Python’s package installer)?Show answerHide

In plain words

The GPU’s driver, NVIDIA’s CUDA software and the server must be exactly matching versions. A container ships a set known to work together; installing with pip on the machine mixes versions and can break them.

Picture it

A meal kit comes with exactly the ingredients the recipe was tested with. Shop for each item yourself and you end up with a different flour, and the cake fails.

With real numbersthe Serving With vLLM video (from Day 2), the lab book’s DGX Spark page and lesson 14’s scripts

  • The Serving With vLLM video: when versions clash, the usual workaround switches off a GPU speed-up (CUDA graphs) and loses 20 to 30% of total output.
  • The lab book: the pip route for vLLM on the Spark’s arm64 chip has broken for people (mismatched versions of PyTorch, the Python library vLLM is built on, and CUDA).
  • NVIDIA ships about 65 DGX Spark setup guides (playbooks) that fix container versions known to work.
  • 1_vllm.sh pins one image tag, nvcr.io/nvidia/vllm:25.09-py3; swap it only for another tag built for the Spark (GB10/arm64).

Words to know

Container
A sealed package holding a program and the exact software versions it needs, so it runs the same anywhere.
CUDA
NVIDIA’s software layer that lets programs run on its GPUs.
Driver
The software that lets the operating system use a piece of hardware such as a GPU.
pip
Python’s tool for installing packages. Example: pip install vllm, which the lesson forbids on the Spark.
Go deeper: the engineer version

The kit's question

Why containers only?

The kit's answer

CUDA, driver and engine versions have to match exactly. Containers pin them, and host pip installs break them.

More detail: On GB10 (arm64 with a Blackwell GPU) the host OS, driver and CUDA move together on NVIDIA’s schedule. Prebuilt wheels can expect a different CUDA than the host has, and the usual workaround, eager mode without CUDA graphs (--enforce-eager), costs 20 to 30%. NGC and vendor containers pin CUDA, PyTorch and the engine together, so a customer can reproduce your numbers with the same tag.

Question What do you test models on, and how do you keep the numbers honest?

One clear answer

Mac for iteration, Spark for the server stack. It’s the same harness pointed at both, so the numbers are comparable, and I validate the fabric before I scale across boxes.

What this means

  • “Mac for iteration”: I try ideas on my Mac first: free, quick and private.
  • “Spark for the server stack”: I run the production server software, vLLM and SGLang in containers, on an NVIDIA DGX Spark, which a Mac cannot run.
  • “It’s the same harness pointed at both”: The same test scripts from Days 1, 3 and 5 run against both machines; only the target changes (--target spark).
  • “so the numbers are comparable”: Any difference comes from the machine and the server software, not from how I measured. The Spark and an M4 Pro Mac both move 273 GB a second, so one person’s writing speed should be close: a gap there is mostly the software. Reading long prompts and serving many people also need math power, where the Spark is far stronger (the lab book: compute-rich, bandwidth-poor).
  • “and I validate the fabric before I scale across boxes”: Before I split one model over two linked Sparks, I test the link (the fabric) on its own, with a network speed test and NVIDIA’s GPU-to-GPU test (NCCL). A slow link wipes out the gain.

Your numbersSaved on this device and collected in the Day 9 wrap-up.

Hint: From the Maximum concurrency for 8192 tokens per request: log line that 1_vllm.sh prints.

Hint: The Capacity line: from 2_bench.sh gives the users; read agg_tok_s on that row of the sweep table.

Hint: The usage: line of the prefix test against port 30000.

vLLM and SGLang both answered from the Spark, 2_bench.sh printed a sweep with a capacity line, SGLang’s usage line shows reused tokens, and 9_teardown.sh freed the GPU. No Spark: read the page, answer the checks, then tap Skip.

set SPARK_IP in .env.
Add SPARK_IP=<the Spark's address> to the .env file in 03-labs, plus SPARK_USER=<your login> if it is not nvidia.
SSH asks for a password, or says Permission denied (publickey).
remote.sh needs SSH keys. Set them up once with ssh-copy-id nvidia@<the Spark's address> (with your own login if it differs), then try again.
An image tag is not found.
Container tags change often. Look up the current GB10/arm64 tag (links are in the script comments). Then replace the tag after :- in the IMAGE= line near the top of 1_vllm.sh or 3_sglang.sh, in your copy on the Mac (remote.sh sends that file to the Spark). remote.sh passes only HF_TOKEN, so VLLM_IMAGE or SGLANG_IMAGE set on your Mac never arrives.
A model download fails with an access error.
It may be a gated model, one you must request on Hugging Face first. Request access, then put your token in .env as HF_TOKEN=…; remote.sh passes it to the Spark.
2_bench.sh says python: command not found or No module named 'felab'.
It runs on your Mac. From the 03-labs folder, run source .venv/bin/activate first.
The prefix test cannot connect to http://:30000/v1.
$SPARK_IP was empty in your terminal. Run export SPARK_IP=<the Spark's address>, or paste the command 3_sglang.sh printed.
nvidia-smi shows memory as N/A.
Normal on a Spark: its CPU and GPU share one memory, so the usual fields are blank. Use free -h on the Spark instead.
TWO_BOX.md Notes

Two Sparks: validate the fabric before you trust tensor parallelism

Section titled “Two Sparks: validate the fabric before you trust tensor parallelism”

Two DGX Sparks can be linked through their ConnectX-7 ports to serve a model neither can hold alone (for example a 200B+ model at 4-bit). Scaling out adds a new failure mode: the network. Tensor parallelism does an all-reduce inside every layer for every token, so a slow or misconfigured link wipes out the gain.

The field-engineer habit is to measure the pipe first, then the model.

  • Cable the ConnectX-7 ports directly (QSFP). Give each side an IP on a private subnet.
  • ip addr, ethtool <iface> | grep Speed should report the expected link speed.
  • Set up passwordless SSH both ways, and use the same container image and model path on both boxes.
Terminal window
# box A # box B
iperf3 -s iperf3 -c <A_IP> -P 8 -t 20 # TCP sanity check
# NCCL all-reduce across both boxes (from the NGC PyTorch container, nccl-tests built in or compiled):
mpirun -np 2 -H <A_IP>,<B_IP> ... all_reduce_perf -b 8M -e 1G -f 2 -g 1

Read busbw at large message sizes. If it is far below link speed, fix the network (interface pinning with NCCL_SOCKET_IFNAME, RDMA/RoCE settings, MTU) before touching vLLM.

  • Use vLLM with a Ray cluster spanning both boxes (ray start --head on A, ray start --address=<A_IP>:6379 on B), then vllm serve <model> --tensor-parallel-size 2 or --pipeline-parallel-size 2.
  • Try PP=2 as well as TP=2. Over a network link, pipeline parallelism (one hand-off per layer block) often beats tensor parallelism (an all-reduce every layer).
  • Re-run 2_bench.sh. Compare single-box performance on a model that fits with two-box performance on the same model. That overhead is the cost of scale-out.

“Before tensor parallel across nodes I validate the fabric with NCCL all-reduce busbw. Across a network I’d try pipeline parallel first; TP wants NVLink-class bandwidth.”

0_check.sh Bash · 10 lines
#!/usr/bin/env bash
# Lesson 14 · step 0 — what is this box? (runs on the Spark)
# DGX Spark = GB10 Grace-Blackwell, 128 GB unified memory at ~273 GB/s, arm64, CUDA.
# Rule: everything runs in CONTAINERS. Never `pip install vllm` on the host.
set -uo pipefail
echo "▸ GPU / driver / CUDA"; nvidia-smi --query-gpu=name,driver_version --format=csv,noheader; nvidia-smi | grep -i "cuda version"
echo "▸ memory"; free -g | head -2
echo "▸ arch"; uname -m
echo "▸ docker + GPU access"; docker run --rm --gpus all ubuntu:24.04 nvidia-smi -L 2>&1 | tail -1
echo "▸ running containers"; docker ps --format '{{.Names}} {{.Image}} {{.Status}}'
1_vllm.sh Bash · 27 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 14 · step 1 — vLLM in a container on the Spark (the production-grade server).
#
# Image: NVIDIA's NGC vLLM build supports the GB10 (arm64 + Blackwell). Pick the newest
# tag at https://catalog.ngc.nvidia.com (search "vllm"); override with VLLM_IMAGE.
#
# --gpus all --ipc host GPU + shared memory for the engine's workers
# -v ~/.cache/huggingface keep downloads between runs
# --max-model-len 8192 caps KV per sequence (lesson 05)
# --gpu-memory-utilization 0.8 share of the 128 GB vLLM may claim (weights + KV pool)
# --enable-prefix-caching automatic prefix caching (lesson 06)
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
MODEL=${1:-nvidia/Llama-3.1-8B-Instruct-FP8}
IMAGE=${VLLM_IMAGE:-nvcr.io/nvidia/vllm:25.09-py3}
docker rm -f vllm 2>/dev/null || true
docker run -d --name vllm --gpus all --ipc host -p 8000:8000 \
-e HF_TOKEN="${HF_TOKEN:-}" -v ~/.cache/huggingface:/root/.cache/huggingface \
"$IMAGE" vllm serve "$MODEL" --host 0.0.0.0 --port 8000 \
--max-model-len 8192 --gpu-memory-utilization 0.8 --enable-prefix-caching
echo "▸ loading $MODEL (first run downloads it) — waiting for /health"
until curl -sf localhost:8000/health >/dev/null; do sleep 5; echo -n "."; done
echo " ready at http://$(hostname -I | awk '{print $1}'):8000/v1"
docker logs vllm 2>&1 | grep -iE "KV cache|maximum concurrency" | tail -3
2_bench.sh Bash · 13 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 14 · step 2 — the SAME harnesses from lessons 01/03/06, pointed at the Spark.
# Run this ON YOUR MAC (it talks to the Spark over the network).
# Then, for comparison, vLLM's own benchmark from inside the container (remote.sh 2b_vllm_bench.sh).
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
cd "$(dirname "$0")"
L=..
python $L/01-latency-harness/bench_ttft.py --target spark --runs 5
python $L/03-concurrency-sweep/sweep.py --target spark --levels 1,2,4,8,16,32,64 --slo-ms 1000
python $L/03-concurrency-sweep/plot_sweep.py # Mac and Spark curves on one chart
python $L/06-prefix-caching/prefix_bench.py --target spark
2b_vllm_bench.sh Bash · 12 lines
#!/usr/bin/env bash
# Lesson 14 · step 2b — vLLM's built-in load generator, run inside the container (on the Spark).
# Random 512-in / 128-out prompts at rising request rates; read "Output token throughput"
# and "P99 TTFT" and compare with your own sweep.py numbers — they should tell the same story.
set -euo pipefail
MODEL=$(curl -s localhost:8000/v1/models | python3 -c "import sys,json;print(json.load(sys.stdin)['data'][0]['id'])")
for rate in 1 4 16 inf; do
echo "▸ request rate $rate"
docker exec vllm vllm bench serve --model "$MODEL" --dataset-name random \
--random-input-len 512 --random-output-len 128 --num-prompts 64 --request-rate $rate 2>&1 \
| grep -E "Output token throughput|Mean TTFT|P99 TTFT|Mean ITL"
done
3_sglang.sh Bash · 24 lines
#!/usr/bin/env bash
# ─────────────────────────────────────────────────────────────────────────────
# Lesson 14 · step 3 — SGLang: same model, different engine, RadixAttention prefix cache.
# Stops vLLM first (one engine at a time on one GPU). Serves on :30000.
#
# Image: SGLang publishes Spark/arm64 builds — check https://hub.docker.com/r/lmsysorg/sglang/tags
# for the current Spark tag and override with SGLANG_IMAGE.
# --mem-fraction-static 0.75 share of memory for weights + KV pool
# --enable-cache-report adds cached_tokens to usage → proves prefix hits
# ─────────────────────────────────────────────────────────────────────────────
set -euo pipefail
MODEL=${1:-Qwen/Qwen3-8B}
IMAGE=${SGLANG_IMAGE:-lmsysorg/sglang:spark}
docker rm -f vllm sglang 2>/dev/null || true
docker run -d --name sglang --gpus all --ipc host -p 30000:30000 \
-e HF_TOKEN="${HF_TOKEN:-}" -v ~/.cache/huggingface:/root/.cache/huggingface \
"$IMAGE" python3 -m sglang.launch_server --model-path "$MODEL" --host 0.0.0.0 --port 30000 \
--mem-fraction-static 0.75 --enable-cache-report
until curl -sf localhost:30000/health >/dev/null; do sleep 5; echo -n "."; done
echo " ready. From the Mac:"
echo " SPARK_URL=http://$(hostname -I | awk '{print $1}'):30000/v1 SPARK_MODEL=$MODEL \\"
echo " python ../06-prefix-caching/prefix_bench.py --target spark --usage"
9_teardown.sh Bash · 3 lines
#!/usr/bin/env bash
# Lesson 14 — free the GPU (runs on the Spark). Downloads stay in ~/.cache/huggingface.
docker rm -f vllm sglang 2>/dev/null; docker ps; nvidia-smi --query-gpu=memory.used --format=csv
remote.sh Bash · 11 lines
#!/usr/bin/env bash
# Run any script in this folder ON the Spark, from your Mac:
# bash remote.sh 0_check.sh
# bash remote.sh 1_vllm.sh nvidia/Llama-3.1-8B-Instruct-FP8
# Needs SPARK_IP (and optionally SPARK_USER) in .env or the environment, and SSH keys set up.
set -euo pipefail
cd "$(dirname "$0")"
[[ -f ../../.env ]] && set -a && source ../../.env && set +a
: "${SPARK_IP:?set SPARK_IP in .env}"
script=$1; shift
ssh "${SPARK_USER:-nvidia}@$SPARK_IP" "HF_TOKEN=${HF_TOKEN:-} bash -s -- $*" < "$script"