The Spark track
course 16 of 22
lesson 14 · lessons/14-spark-track
2 h 30 Free (your hardware) DGX Spark Optional: needs a DGX Spark
What you will do and why
Optional (needs a DGX Spark): you run the server software real services use, vLLM and SGLang, in containers on NVIDIA hardware. Your own test scripts from earlier days then measure it, so its numbers compare directly with your Mac’s.
Why it matters: If a team asks whether a small NVIDIA desktop can serve its internal users, test it with the same questions and timing script you used on your Mac. This step is optional because it needs a DGX Spark. Its published memory speed resembles an M4 Pro, but the measured reply speed and number of users it supports must be checked on the actual machine.
You are done when: vLLM and SGLang both answered from the Spark, 2_bench.sh printed a sweep with a capacity line, SGLang’s usage line shows reused tokens, and 9_teardown.sh freed the GPU. No Spark: read the page, answer the checks, then tap Skip.
Rewatch from Day 2: Serving With vLLM, 0:48 to 1:48ShowHide
In plain words
Section titled “In plain words”The Mac taught the ideas. The Spark runs the production server software (the programs live services use): vLLM and SGLang on an NVIDIA GPU. You start each one in a container on the Spark and drive it from your Mac with the same scripts as before, so the numbers line up with your Mac’s.
Picture it
Same recipe, same measuring cups, but a professional kitchen. Because you measure the same way, any difference in the result comes from the kitchen, not from how you measured.
With real numberslesson 14 and the lab book’s DGX Spark page
- DGX Spark: 128 GB of memory shared by its CPU and GPU, about 115 GB usable for the model and its notes.
- Its memory moves 273 GB a second, like an M4 Pro Mac. Llama 3.1 8B at 8 bits is 8 GB, so one person’s limit is 273 ÷ 8 = 34 tokens a second.
- The lab book’s measurement: 20.5 tokens a second for 1 person, 60% of that limit (20.5 ÷ 34).
- With 32 people the same Spark wrote 368 tokens a second in total: 368 ÷ 20.5 = 18 times one person’s output.
- An H100 moves 3,350 GB a second, 12 times the Spark: the Spark is for learning the server stack, not for data-center speed.
Words to know
- DGX Spark
- NVIDIA’s desktop AI computer, with 128 GB of memory shared by its CPU and GPU.
- Container
- A sealed package holding a program and the exact software versions it needs. Example: vLLM’s image on the Spark.
- vLLM
- An open-source server for NVIDIA GPUs built for many users at once.
- SGLang
- Another open-source server for NVIDIA GPUs, strong at reusing shared prompt openings.
From the lesson
lessons/14-spark-track/README.md
What and why
Section titled “What and why”The Mac covers the concepts. The Spark gives you the production server stack the job description names: vLLM and SGLang on an NVIDIA GPU, in containers, benchmarked with the same harness you built in lessons 01, 03 and 06 (Days 1, 3 and 5). The DGX Spark (GB10, 128 GB unified memory, ~273 GB/s) has about the same bandwidth as an M4 Pro, so your ceiling maths carries straight over. The difference is the software: CUDA, FP8/NVFP4 and the engines customers run.
Mac ──(same python scripts, --target spark)──► Spark :8000 vLLM / :30000 SGLangThe scripts
Section titled “The scripts”| script | where it runs | does |
|---|---|---|
remote.sh X.sh |
Mac | runs any script below on the Spark over SSH |
0_check.sh |
Spark | GPU, driver, CUDA, memory, and whether Docker can see the GPU |
1_vllm.sh [model] |
Spark | vLLM container with prefix caching; prints the KV pool size |
2_bench.sh |
Mac | lessons 01, 03 and 06 (Days 1, 3 and 5) against the Spark, with Mac and Spark curves on one chart |
2b_vllm_bench.sh |
Spark | vllm bench serve as a cross-check on your own numbers |
3_sglang.sh [model] |
Spark | SGLang with --enable-cache-report; RadixAttention hits |
9_teardown.sh |
Spark | frees the GPU |
TWO_BOX.md |
– | fabric first (iperf, NCCL busbw), then TP=2 or PP=2 |
# .env: SPARK_IP=…, HF_TOKEN=… (for gated models)cd lessons/14-spark-trackbash remote.sh 0_check.shbash remote.sh 1_vllm.sh nvidia/Llama-3.1-8B-Instruct-FP8bash 2_bench.shbash remote.sh 2b_vllm_bench.shbash remote.sh 3_sglang.sh Qwen/Qwen3-8BSPARK_URL=http://$SPARK_IP:30000/v1 SPARK_MODEL=Qwen/Qwen3-8B \ python ../06-prefix-caching/prefix_bench.py --target spark --usagebash remote.sh 9_teardown.shWhat each command does
bash remote.sh 0_check.shRuns the check script on the Spark over SSH: your Mac logs in and runs it there. It needs
SPARK_IPin03-labs/.envand SSH keys set up. Look for the GPU name, driver and CUDA version, the memory, the chip type (arm64, printed asaarch64), and a GPU listed from inside a test container: proof that Docker can see the GPU.bash remote.sh 1_vllm.sh nvidia/Llama-3.1-8B-Instruct-FP8Starts vLLM in a container on the Spark with Llama 3.1 8B stored at 8 bits: conversations capped at 8,192 tokens, up to 80% of memory, prefix caching on (Day 5). The first run downloads the model. Look for
ready at http://…:8000/v1, then the log lines about the KV cache andMaximum concurrency for 8192 tokens per request: Day 4’s notes calculation, done by the server.bash 2_bench.shRuns on your Mac, not the Spark. It points three of your own scripts at the Spark: Day 1’s timing test (5 runs); Day 3’s sweep from 1 to 64 people with a 1,000 ms promise, which it then draws with your earlier sweeps on one chart; and Day 5’s shared-opening test. Look for the sweep’s
Capacity line:and the chart inresults/03-sweep.png.bash remote.sh 2b_vllm_bench.shRuns vLLM’s own load tester inside the container: 64 made-up requests of 512 tokens in and 128 out, sent at 1, 4 and 16 requests a second, then all at once. Look for
Output token throughput(total tokens written a second) andP99 TTFT(99 of 100 requests got their first token within this time): they should tell the same story as your own sweep.bash remote.sh 3_sglang.sh Qwen/Qwen3-8BStops vLLM (one server at a time on one GPU) and starts SGLang with the Qwen3 8B model, set to report reused tokens (
--enable-cache-report). Look forready. From the Mac:followed by the exact command to run next, with the Spark’s address filled in.SPARK_URL=http://$SPARK_IP:30000/v1 SPARK_MODEL=Qwen/Qwen3-8BDay 5’s shared-opening test against SGLang on port 30000, plus the extra call that prints token counts. Your terminal fills in
$SPARK_IPitself and does not read.env, so runexport SPARK_IP=<the Spark's address>first, or paste the command3_sglang.shprinted. Look for theusage:line, with cached_tokens close to prompt_tokens on the same line: nearly the whole opening (the script estimates about 3,028 tokens) was reused.bash remote.sh 9_teardown.shStops both servers and frees the GPU; downloaded models stay on the Spark for next time. Look for no vllm or sglang in the container list. The GPU memory line may read N/A: the Spark’s memory is shared, so the lab book says to use
free -hthere instead.
Container tags move quickly. If an image tag isn’t found, look up the current GB10/arm64
tag (links are in the script comments) and replace the tag after :- in the IMAGE= line near the top of 1_vllm.sh or 3_sglang.sh, in your copy on the Mac (remote.sh sends that file to the Spark). (remote.sh passes only HF_TOKEN to the Spark, so VLLM_IMAGE or SGLANG_IMAGE set on your Mac never arrives.)
What you should see
Section titled “What you should see”How to read it
First, vLLM’s Maximum concurrency for 8192 tokens per request: …x line is Day 4’s notes calculation, done by the server: how many full 8,192-token conversations fit in its memory for notes. Second, the Spark’s sweep curve keeps rising well past your Mac’s: vLLM lets requests join at every step and hands out note memory in small pages. Third, SGLang’s usage line shows cached_tokens close to prompt_tokens: proof it reused the shared opening.
- The vLLM log line
Maximum concurrency for 8192 tokens per request: …xis lesson 05’s KV maths (Day 4) done by the engine. - The Spark’s sweep curve keeps rising well past the Mac’s, because vLLM’s continuous batching and paged KV cache handle many users far better than a laptop server.
- The SGLang usage block shows
cached_tokensclose to the preamble size on the extra call the script makes after its 12 timed calls.
Check yourself
Section titled “Check yourself”3 questions. Say your answer out loud, then tap to check it.
vLLM’s startup log says its memory for notes holds 18 full conversations of 8K tokens (8,192, about 6,000 words) at once: maximum concurrency 18x. How do you double that?Show answerHide
In plain words
Make each conversation’s notes smaller, or give the notes more memory. Store them in 8 bits instead of 16, lower the cap on conversation length, or use a smaller or more compressed model.
Picture it
A wardrobe holds 18 coats. Vacuum-pack each coat to half its size, or allow jackets instead of long coats, and 36 fit. Or take out the big suitcase that was using some of the space.
With real numbersDay 4’s calculator for Llama 3.1 8B, and the vLLM settings from the Serving With vLLM video
- Notes for one 8,192-token conversation at 16 bits: 1.07 GB (Day 4).
- Say the server has 20 GB for notes, as in Day 4’s example: 20 ÷ 1.07 = 18.6, so vLLM reports about 18.6x (18 whole conversations): the check’s 18x.
- 8-bit notes (
--kv-cache-dtype fp8): 1.07 ÷ 2 = 0.54 GB each, so 20 ÷ 0.54 = 37. - Or cap conversations at 4,096 tokens (
--max-model-len 4096): also 0.54 GB each, also 37, if real requests fit in 4,096. - Or a smaller or more compressed model: its weights shrink, and vLLM gives the freed memory to notes.
Words to know
- Maximum concurrency
- vLLM’s count of how many full-length conversations fit in its memory for notes.
- KV cache dtype
- The number format the notes are stored in; 8-bit instead of 16-bit halves their memory.
- Max model length (--max-model-len)
- vLLM’s cap on tokens per conversation, which caps that conversation’s notes.
- FP8
- An 8-bit number format, 1 byte per number.
Go deeper: the engineer version
The kit's question
The vLLM log says max concurrency is 18× at 8k context. How do you double it?
The kit's answer
Use an FP8 KV cache, a lower
--max-model-lenor a smaller or quantized model.
More detail: vLLM’s figure is KV-cache tokens ÷ max_model_len. The Serving With vLLM video’s sample log: 8,234 blocks × 16 tokens ≈ 131,000 tokens, about 16 requests at 8K. FP8 KV halves bytes per token, so it doubles the figure exactly (the video: “roughly doubles concurrency”). Halving --max-model-len also doubles it, but is only safe if the real p95 fits (Day 4). A smaller or quantized model frees memory inside --gpu-memory-utilization for the KV pool, which raises the figure by however much it frees.
When would you pick SGLang over vLLM, the two main open-source servers for NVIDIA GPUs?Show answerHide
In plain words
When many requests share long openings or branch from one conversation, as agents do, or when most output must follow a strict format such as JSON (a common format for structured data). Either way, test both on the customer’s real traffic before choosing.
Picture it
Two delivery firms: one is best at many different parcels, the other at many parcels to the same street. Which is cheaper depends on your parcels, so you give each a trial week.
With real numbersDay 5’s prefix test and this step’s SGLang run
- Day 5’s shared opening: about 3,028 tokens by the script’s estimate, sent with each of 12 questions.
- Reusing it made the first word 20 times sooner in the lesson’s sample: 740 ms down to 36 ms (thousandths of a second).
- SGLang keeps saved notes like a family tree (RadixAttention): one shared opening at the root, a branch for each follow-up, so every branch reuses the common start.
- This step’s proof: SGLang’s
usage:line shows cached_tokens close to prompt_tokens on the same line: nearly the whole opening was reused.
Words to know
- SGLang
- An open-source server for NVIDIA GPUs, strong at reusing shared prompt openings.
- RadixAttention
- SGLang’s store of saved notes, kept as a tree of shared openings.
- Agent
- A program that calls a model many times, using tools, to finish a task.
- Structured generation
- Making every reply follow a fixed format, such as JSON.
Go deeper: the engineer version
The kit's question
When would you pick SGLang over vLLM?
The kit's answer
For heavy shared-prefix or branching agent workloads and structured-generation-heavy pipelines. Benchmark both on the real traffic.
More detail: SGLang’s RadixAttention keeps the KV cache in a radix tree keyed by token prefixes and schedules requests to maximise prefix hits, which suits multi-turn and branching agent traffic; its constrained decoding is strong for JSON-heavy pipelines. vLLM’s automatic prefix caching also reuses identical block prefixes, so the gap depends on the traffic: benchmark both on the customer’s prompts, with their cache hit rates.
Why run vLLM and SGLang on the Spark only inside containers, never installed straight onto the machine with pip (Python’s package installer)?Show answerHide
In plain words
The GPU’s driver, NVIDIA’s CUDA software and the server must be exactly matching versions. A container ships a set known to work together; installing with pip on the machine mixes versions and can break them.
Picture it
A meal kit comes with exactly the ingredients the recipe was tested with. Shop for each item yourself and you end up with a different flour, and the cake fails.
With real numbersthe Serving With vLLM video (from Day 2), the lab book’s DGX Spark page and lesson 14’s scripts
- The Serving With vLLM video: when versions clash, the usual workaround switches off a GPU speed-up (CUDA graphs) and loses 20 to 30% of total output.
- The lab book: the pip route for vLLM on the Spark’s arm64 chip has broken for people (mismatched versions of PyTorch, the Python library vLLM is built on, and CUDA).
- NVIDIA ships about 65 DGX Spark setup guides (playbooks) that fix container versions known to work.
1_vllm.shpins one image tag,nvcr.io/nvidia/vllm:25.09-py3; swap it only for another tag built for the Spark (GB10/arm64).
Words to know
- Container
- A sealed package holding a program and the exact software versions it needs, so it runs the same anywhere.
- CUDA
- NVIDIA’s software layer that lets programs run on its GPUs.
- Driver
- The software that lets the operating system use a piece of hardware such as a GPU.
- pip
- Python’s tool for installing packages. Example:
pip install vllm, which the lesson forbids on the Spark.
Go deeper: the engineer version
The kit's question
Why containers only?
The kit's answer
CUDA, driver and engine versions have to match exactly. Containers pin them, and host pip installs break them.
More detail: On GB10 (arm64 with a Blackwell GPU) the host OS, driver and CUDA move together on NVIDIA’s schedule. Prebuilt wheels can expect a different CUDA than the host has, and the usual workaround, eager mode without CUDA graphs (--enforce-eager), costs 20 to 30%. NGC and vendor containers pin CUDA, PyTorch and the engine together, so a customer can reproduce your numbers with the same tag.
Explain what you learned
Section titled “Explain what you learned”Question What do you test models on, and how do you keep the numbers honest?
One clear answer
Mac for iteration, Spark for the server stack. It’s the same harness pointed at both, so the numbers are comparable, and I validate the fabric before I scale across boxes.
What this means
- “Mac for iteration”: I try ideas on my Mac first: free, quick and private.
- “Spark for the server stack”: I run the production server software, vLLM and SGLang in containers, on an NVIDIA DGX Spark, which a Mac cannot run.
- “It’s the same harness pointed at both”: The same test scripts from Days 1, 3 and 5 run against both machines; only the target changes (
--target spark). - “so the numbers are comparable”: Any difference comes from the machine and the server software, not from how I measured. The Spark and an M4 Pro Mac both move 273 GB a second, so one person’s writing speed should be close: a gap there is mostly the software. Reading long prompts and serving many people also need math power, where the Spark is far stronger (the lab book: compute-rich, bandwidth-poor).
- “and I validate the fabric before I scale across boxes”: Before I split one model over two linked Sparks, I test the link (the fabric) on its own, with a network speed test and NVIDIA’s GPU-to-GPU test (NCCL). A slow link wipes out the gain.
Your numbersSaved on this device and collected in the Day 9 wrap-up.
Hint: From the Maximum concurrency for 8192 tokens per request: log line that 1_vllm.sh prints.
Hint: The Capacity line: from 2_bench.sh gives the users; read agg_tok_s on that row of the sweep table.
Hint: The usage: line of the prefix test against port 30000.
Done when
Section titled “Done when”vLLM and SGLang both answered from the Spark, 2_bench.sh printed a sweep with a capacity line, SGLang’s usage line shows reused tokens, and 9_teardown.sh freed the GPU. No Spark: read the page, answer the checks, then tap Skip.
Stuck?
Section titled “Stuck?”set SPARK_IP in .env.- Add
SPARK_IP=<the Spark's address>to the.envfile in03-labs, plusSPARK_USER=<your login>if it is notnvidia. - SSH asks for a password, or says
Permission denied (publickey). remote.shneeds SSH keys. Set them up once withssh-copy-id nvidia@<the Spark's address>(with your own login if it differs), then try again.- An image tag is not found.
- Container tags change often. Look up the current GB10/arm64 tag (links are in the script comments). Then replace the tag after
:-in theIMAGE=line near the top of1_vllm.shor3_sglang.sh, in your copy on the Mac (remote.shsends that file to the Spark).remote.shpasses onlyHF_TOKEN, soVLLM_IMAGEorSGLANG_IMAGEset on your Mac never arrives. - A model download fails with an access error.
- It may be a gated model, one you must request on Hugging Face first. Request access, then put your token in
.envasHF_TOKEN=…;remote.shpasses it to the Spark. 2_bench.shsayspython: command not foundorNo module named 'felab'.- It runs on your Mac. From the
03-labsfolder, runsource .venv/bin/activatefirst. - The prefix test cannot connect to
http://:30000/v1. $SPARK_IPwas empty in your terminal. Runexport SPARK_IP=<the Spark's address>, or paste the command3_sglang.shprinted.nvidia-smishows memory as N/A.- Normal on a Spark: its CPU and GPU share one memory, so the usual fields are blank. Use
free -hon the Spark instead.
Code in this step
Section titled “Code in this step”TWO_BOX.md Notes
Two Sparks: validate the fabric before you trust tensor parallelism
Section titled “Two Sparks: validate the fabric before you trust tensor parallelism”Two DGX Sparks can be linked through their ConnectX-7 ports to serve a model neither can hold alone (for example a 200B+ model at 4-bit). Scaling out adds a new failure mode: the network. Tensor parallelism does an all-reduce inside every layer for every token, so a slow or misconfigured link wipes out the gain.
The field-engineer habit is to measure the pipe first, then the model.
1. Link up
Section titled “1. Link up”- Cable the ConnectX-7 ports directly (QSFP). Give each side an IP on a private subnet.
ip addr,ethtool <iface> | grep Speedshould report the expected link speed.- Set up passwordless SSH both ways, and use the same container image and model path on both boxes.
2. Raw bandwidth, then NCCL
Section titled “2. Raw bandwidth, then NCCL”# box A # box Biperf3 -s iperf3 -c <A_IP> -P 8 -t 20 # TCP sanity check# NCCL all-reduce across both boxes (from the NGC PyTorch container, nccl-tests built in or compiled):mpirun -np 2 -H <A_IP>,<B_IP> ... all_reduce_perf -b 8M -e 1G -f 2 -g 1Read busbw at large message sizes. If it is far below link speed, fix the network
(interface pinning with NCCL_SOCKET_IFNAME, RDMA/RoCE settings, MTU) before touching vLLM.
3. Only then: the model across two boxes
Section titled “3. Only then: the model across two boxes”- Use vLLM with a Ray cluster spanning both boxes (
ray start --headon A,ray start --address=<A_IP>:6379on B), thenvllm serve <model> --tensor-parallel-size 2or--pipeline-parallel-size 2. - Try PP=2 as well as TP=2. Over a network link, pipeline parallelism (one hand-off per layer block) often beats tensor parallelism (an all-reduce every layer).
- Re-run
2_bench.sh. Compare single-box performance on a model that fits with two-box performance on the same model. That overhead is the cost of scale-out.
What to say
Section titled “What to say”“Before tensor parallel across nodes I validate the fabric with NCCL all-reduce busbw. Across a network I’d try pipeline parallel first; TP wants NVLink-class bandwidth.”
0_check.sh Bash · 10 lines
#!/usr/bin/env bash# Lesson 14 · step 0 — what is this box? (runs on the Spark)# DGX Spark = GB10 Grace-Blackwell, 128 GB unified memory at ~273 GB/s, arm64, CUDA.# Rule: everything runs in CONTAINERS. Never `pip install vllm` on the host.set -uo pipefailecho "▸ GPU / driver / CUDA"; nvidia-smi --query-gpu=name,driver_version --format=csv,noheader; nvidia-smi | grep -i "cuda version"echo "▸ memory"; free -g | head -2echo "▸ arch"; uname -mecho "▸ docker + GPU access"; docker run --rm --gpus all ubuntu:24.04 nvidia-smi -L 2>&1 | tail -1echo "▸ running containers"; docker ps --format '{{.Names}} {{.Image}} {{.Status}}'1_vllm.sh Bash · 27 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 14 · step 1 — vLLM in a container on the Spark (the production-grade server).## Image: NVIDIA's NGC vLLM build supports the GB10 (arm64 + Blackwell). Pick the newest# tag at https://catalog.ngc.nvidia.com (search "vllm"); override with VLLM_IMAGE.## --gpus all --ipc host GPU + shared memory for the engine's workers# -v ~/.cache/huggingface keep downloads between runs# --max-model-len 8192 caps KV per sequence (lesson 05)# --gpu-memory-utilization 0.8 share of the 128 GB vLLM may claim (weights + KV pool)# --enable-prefix-caching automatic prefix caching (lesson 06)# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailMODEL=${1:-nvidia/Llama-3.1-8B-Instruct-FP8}IMAGE=${VLLM_IMAGE:-nvcr.io/nvidia/vllm:25.09-py3}
docker rm -f vllm 2>/dev/null || truedocker run -d --name vllm --gpus all --ipc host -p 8000:8000 \ -e HF_TOKEN="${HF_TOKEN:-}" -v ~/.cache/huggingface:/root/.cache/huggingface \ "$IMAGE" vllm serve "$MODEL" --host 0.0.0.0 --port 8000 \ --max-model-len 8192 --gpu-memory-utilization 0.8 --enable-prefix-caching
echo "▸ loading $MODEL (first run downloads it) — waiting for /health"until curl -sf localhost:8000/health >/dev/null; do sleep 5; echo -n "."; doneecho " ready at http://$(hostname -I | awk '{print $1}'):8000/v1"docker logs vllm 2>&1 | grep -iE "KV cache|maximum concurrency" | tail -32_bench.sh Bash · 13 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 14 · step 2 — the SAME harnesses from lessons 01/03/06, pointed at the Spark.# Run this ON YOUR MAC (it talks to the Spark over the network).# Then, for comparison, vLLM's own benchmark from inside the container (remote.sh 2b_vllm_bench.sh).# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailcd "$(dirname "$0")"L=..python $L/01-latency-harness/bench_ttft.py --target spark --runs 5python $L/03-concurrency-sweep/sweep.py --target spark --levels 1,2,4,8,16,32,64 --slo-ms 1000python $L/03-concurrency-sweep/plot_sweep.py # Mac and Spark curves on one chartpython $L/06-prefix-caching/prefix_bench.py --target spark2b_vllm_bench.sh Bash · 12 lines
#!/usr/bin/env bash# Lesson 14 · step 2b — vLLM's built-in load generator, run inside the container (on the Spark).# Random 512-in / 128-out prompts at rising request rates; read "Output token throughput"# and "P99 TTFT" and compare with your own sweep.py numbers — they should tell the same story.set -euo pipefailMODEL=$(curl -s localhost:8000/v1/models | python3 -c "import sys,json;print(json.load(sys.stdin)['data'][0]['id'])")for rate in 1 4 16 inf; do echo "▸ request rate $rate" docker exec vllm vllm bench serve --model "$MODEL" --dataset-name random \ --random-input-len 512 --random-output-len 128 --num-prompts 64 --request-rate $rate 2>&1 \ | grep -E "Output token throughput|Mean TTFT|P99 TTFT|Mean ITL"done3_sglang.sh Bash · 24 lines
#!/usr/bin/env bash# ─────────────────────────────────────────────────────────────────────────────# Lesson 14 · step 3 — SGLang: same model, different engine, RadixAttention prefix cache.# Stops vLLM first (one engine at a time on one GPU). Serves on :30000.## Image: SGLang publishes Spark/arm64 builds — check https://hub.docker.com/r/lmsysorg/sglang/tags# for the current Spark tag and override with SGLANG_IMAGE.# --mem-fraction-static 0.75 share of memory for weights + KV pool# --enable-cache-report adds cached_tokens to usage → proves prefix hits# ─────────────────────────────────────────────────────────────────────────────set -euo pipefailMODEL=${1:-Qwen/Qwen3-8B}IMAGE=${SGLANG_IMAGE:-lmsysorg/sglang:spark}
docker rm -f vllm sglang 2>/dev/null || truedocker run -d --name sglang --gpus all --ipc host -p 30000:30000 \ -e HF_TOKEN="${HF_TOKEN:-}" -v ~/.cache/huggingface:/root/.cache/huggingface \ "$IMAGE" python3 -m sglang.launch_server --model-path "$MODEL" --host 0.0.0.0 --port 30000 \ --mem-fraction-static 0.75 --enable-cache-report
until curl -sf localhost:30000/health >/dev/null; do sleep 5; echo -n "."; doneecho " ready. From the Mac:"echo " SPARK_URL=http://$(hostname -I | awk '{print $1}'):30000/v1 SPARK_MODEL=$MODEL \\"echo " python ../06-prefix-caching/prefix_bench.py --target spark --usage"9_teardown.sh Bash · 3 lines
#!/usr/bin/env bash# Lesson 14 — free the GPU (runs on the Spark). Downloads stay in ~/.cache/huggingface.docker rm -f vllm sglang 2>/dev/null; docker ps; nvidia-smi --query-gpu=memory.used --format=csvremote.sh Bash · 11 lines
#!/usr/bin/env bash# Run any script in this folder ON the Spark, from your Mac:# bash remote.sh 0_check.sh# bash remote.sh 1_vllm.sh nvidia/Llama-3.1-8B-Instruct-FP8# Needs SPARK_IP (and optionally SPARK_USER) in .env or the environment, and SSH keys set up.set -euo pipefailcd "$(dirname "$0")"[[ -f ../../.env ]] && set -a && source ../../.env && set +a: "${SPARK_IP:?set SPARK_IP in .env}"script=$1; shiftssh "${SPARK_USER:-nvidia}@$SPARK_IP" "HF_TOKEN=${HF_TOKEN:-} bash -s -- $*" < "$script"