vLLM
vLLM is the default runtime in sparkrun. It provides first-class support for solo and multi-node inference with two clustering variants:
vllm-distributed— uses vLLM’s built-in distributed backend (the default)vllm-ray— uses Ray head/worker orchestration
You can write runtime: vllm in a recipe and sparkrun resolves it automatically — see Runtime alias resolution below.
Features
Section titled “Features”- Solo and multi-node tensor parallelism across DGX Spark nodes
- Broad model support (most HuggingFace models)
- PagedAttention for efficient KV cache management
- Tool calling support
- Two multi-node strategies: vLLM native distributed or Ray clustering
Container images
Section titled “Container images”Both vllm-distributed and vllm-ray use the same container images. sparkrun works with ready-built images including:
- Spark Arena Official Images (built from @eugr’s repo):
ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latestandghcr.io/spark-arena/dgx-vllm-eugr-nightly-tf5:latest - Self-built community images from eugr/spark-vllm-docker (often the first to have the latest fixes)
- Official NVIDIA Images
We recommend that you use the Spark Arena images for sparkrun recipes; when you use the official Spark Arena images, sparkrun will automatically manage using the latest eugr/spark-vllm-docker.
Sentinel images
Section titled “Sentinel images”The Spark Arena nightly image names — ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest, its -tf5 variant, and the Docker Hub eugr/spark-vllm image — are sentinel images. When you reference one at :latest, sparkrun’s eugr builder resolves it to the authoritative published nightly and pulls it.
This is the default for all official v2 recipes: it gives you a stable, single name to reference while sparkrun handles versioning. The -tf5 and non-tf5 nightlies are identical now, so both resolve the same way.
To pin a specific image rather than tracking the nightly, use any non-:latest tag (e.g. ghcr.io/spark-arena/dgx-vllm-eugr-nightly:2026-03-01) — sparkrun pulls that immutable tag as-is.
See Phase 2: Building image for the full builder lifecycle.
Speculative decoding
Section titled “Speculative decoding”When a recipe sets speculative_config in defaults and the JSON blob has a top-level model field, sparkrun automatically pre-syncs that draft model to all target hosts alongside the primary model. This happens during the distribution phase so the container never has to fetch the draft from HuggingFace at launch time (which would fail under sparkrun’s offline-hub default).
defaults: speculative_config: '{"method": "dflash", "model": "z-lab/Qwen3.6-35B-A3B-DFlash", "num_speculative_tokens": 5}'Both vllm-distributed and vllm-ray honor this — they share the same VllmMixin.detect_spec_config_draft_model resolver. The full speculative_config JSON is passed through to vllm serve via the --speculative-config flag; sparkrun only inspects the model field to wire up pre-sync.
Runtime alias resolution
Section titled “Runtime alias resolution”When a recipe specifies runtime: vllm (or omits runtime entirely), sparkrun resolves it to a concrete runtime in this order:
- If the recipe hints at Ray usage (
distributed_executor_backend: rayin defaults, or--distributed-executor-backend rayin the command template) →vllm-ray - Otherwise →
vllm-distributed(the default)
You can also set runtime: vllm-distributed or runtime: vllm-ray explicitly to bypass this resolution.
Example recipe
Section titled “Example recipe”This recipe uses runtime: vllm, which resolves to vllm-distributed by default:
recipe_version: "2"model: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4runtime: vllmmax_nodes: 1container: ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest
metadata: description: Nemotron-3-Nano NVFP4 — single node only
mods: # this is pulled via eugr/spark-vllm-docker - mods/nemotron-nano
defaults: port: 8000 host: 0.0.0.0 tensor_parallel: 1 gpu_memory_utilization: 0.7 max_model_len: 262144
command: | vllm serve {model} \ --moe-backend cutlass \ --max-model-len {max_model_len} \ --host {host} \ --port {port} \ --trust-remote-code \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --reasoning-parser-plugin nano_v3_reasoning_parser.py \ --reasoning-parser nano_v3 \ --kv-cache-dtype fp8 \ --enable-prefix-caching \ --load-format fastsafetensors \ --gpu-memory-utilization {gpu_memory_utilization}Multi-node with vllm-distributed (default)
Section titled “Multi-node with vllm-distributed (default)”vllm-distributed uses vLLM’s built-in multi-node support. sparkrun handles container lifecycle, InfiniBand detection, and node coordination automatically:
sparkrun run qwen3-1.7b-vllm --tp 2In this case, the recipe default is for single node with tensor_parallel of 1, but we override tensor parallel to leverage 2 DGX Sparks.
No Ray cluster is involved — vLLM coordinates the nodes directly via its native distributed backend. See Execution Flow for the full multi-node orchestration details.
Tensor, pipeline, and data parallelism
Section titled “Tensor, pipeline, and data parallelism”vllm-distributed supports three parallelism regimes on DGX Spark (one GPU per node, so each parallel dimension consumes physical hosts):
| Mode | Defaults | Behavior |
|---|---|---|
| Cross-node TP/PP, single replica | tensor_parallel: N or pipeline_parallel: N (DP=1) | sparkrun appends --nnodes, --node-rank, --master-addr, --master-port; worker ranks get --headless. |
| Pure data parallelism | data_parallel: N (TP=PP=1) | Each node is its own replica; sparkrun appends --data-parallel-size, --data-parallel-rank, --data-parallel-address, --data-parallel-rpc-port. |
| Hybrid (TP/PP × DP) | tensor_parallel: T, data_parallel: D | sparkrun maps replicas to host groups of size T·PP; --master-addr is the first host of each replica; both flag sets are emitted. |
Required node count = tensor_parallel × pipeline_parallel × data_parallel. The data_parallel_rpc_port defaults to 13345 and can be overridden in recipe defaults.
Multi-node with vllm-ray
Section titled “Multi-node with vllm-ray”vllm-ray forms a Ray cluster first, then runs the serve command on the head node. sparkrun handles Ray cluster setup, InfiniBand detection, and container lifecycle automatically.
To use vllm-ray, either set runtime: vllm-ray explicitly or add the Ray hint to defaults:
runtime: vllmdefaults: distributed_executor_backend: rayRay-specific options
Section titled “Ray-specific options”These options apply only to vllm-ray:
| Option | Default | Description |
|---|---|---|
--ray-port | 46379 | Ray GCS port |
--dashboard | off | Enable Ray dashboard |
--dashboard-port | 8265 | Ray dashboard port |
Choosing between variants
Section titled “Choosing between variants”Both variants support the same models and vLLM flags. The difference is in multi-node orchestration:
vllm-distributedis simpler — no Ray dependency, each node runs independently with vLLM’s built-in coordination. This is the default.vllm-rayprovides Ray’s cluster management, dashboard, and monitoring. Use it when you need Ray-specific features.
For single-node (solo) workloads, both variants behave identically.