Skip to content

vLLM

vLLM is the default runtime in sparkrun. It provides first-class support for solo and multi-node inference with two clustering variants:

  • vllm-distributed — uses vLLM’s built-in distributed backend (the default)
  • vllm-ray — uses Ray head/worker orchestration

You can write runtime: vllm in a recipe and sparkrun resolves it automatically — see Runtime alias resolution below.

  • Solo and multi-node tensor parallelism across DGX Spark nodes
  • Broad model support (most HuggingFace models)
  • PagedAttention for efficient KV cache management
  • Tool calling support
  • Two multi-node strategies: vLLM native distributed or Ray clustering

Both vllm-distributed and vllm-ray use the same container images. sparkrun works with ready-built images including:

We recommend that you use the Spark Arena images for sparkrun recipes; when you use the official Spark Arena images, sparkrun will automatically manage using the latest eugr/spark-vllm-docker.

The Spark Arena nightly image names — ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest, its -tf5 variant, and the Docker Hub eugr/spark-vllm image — are sentinel images. When you reference one at :latest, sparkrun’s eugr builder resolves it to the authoritative published nightly and pulls it.

This is the default for all official v2 recipes: it gives you a stable, single name to reference while sparkrun handles versioning. The -tf5 and non-tf5 nightlies are identical now, so both resolve the same way.

To pin a specific image rather than tracking the nightly, use any non-:latest tag (e.g. ghcr.io/spark-arena/dgx-vllm-eugr-nightly:2026-03-01) — sparkrun pulls that immutable tag as-is.

See Phase 2: Building image for the full builder lifecycle.

When a recipe sets speculative_config in defaults and the JSON blob has a top-level model field, sparkrun automatically pre-syncs that draft model to all target hosts alongside the primary model. This happens during the distribution phase so the container never has to fetch the draft from HuggingFace at launch time (which would fail under sparkrun’s offline-hub default).

defaults:
speculative_config: '{"method": "dflash", "model": "z-lab/Qwen3.6-35B-A3B-DFlash", "num_speculative_tokens": 5}'

Both vllm-distributed and vllm-ray honor this — they share the same VllmMixin.detect_spec_config_draft_model resolver. The full speculative_config JSON is passed through to vllm serve via the --speculative-config flag; sparkrun only inspects the model field to wire up pre-sync.

When a recipe specifies runtime: vllm (or omits runtime entirely), sparkrun resolves it to a concrete runtime in this order:

  1. If the recipe hints at Ray usage (distributed_executor_backend: ray in defaults, or --distributed-executor-backend ray in the command template) → vllm-ray
  2. Otherwise → vllm-distributed (the default)

You can also set runtime: vllm-distributed or runtime: vllm-ray explicitly to bypass this resolution.

This recipe uses runtime: vllm, which resolves to vllm-distributed by default:

recipe_version: "2"
model: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4
runtime: vllm
max_nodes: 1
container: ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest
metadata:
description: Nemotron-3-Nano NVFP4 — single node only
mods: # this is pulled via eugr/spark-vllm-docker
- mods/nemotron-nano
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.7
max_model_len: 262144
command: |
vllm serve {model} \
--moe-backend cutlass \
--max-model-len {max_model_len} \
--host {host} \
--port {port} \
--trust-remote-code \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser-plugin nano_v3_reasoning_parser.py \
--reasoning-parser nano_v3 \
--kv-cache-dtype fp8 \
--enable-prefix-caching \
--load-format fastsafetensors \
--gpu-memory-utilization {gpu_memory_utilization}

Multi-node with vllm-distributed (default)

Section titled “Multi-node with vllm-distributed (default)”

vllm-distributed uses vLLM’s built-in multi-node support. sparkrun handles container lifecycle, InfiniBand detection, and node coordination automatically:

Terminal window
sparkrun run qwen3-1.7b-vllm --tp 2

In this case, the recipe default is for single node with tensor_parallel of 1, but we override tensor parallel to leverage 2 DGX Sparks.

No Ray cluster is involved — vLLM coordinates the nodes directly via its native distributed backend. See Execution Flow for the full multi-node orchestration details.

vllm-distributed supports three parallelism regimes on DGX Spark (one GPU per node, so each parallel dimension consumes physical hosts):

ModeDefaultsBehavior
Cross-node TP/PP, single replicatensor_parallel: N or pipeline_parallel: N (DP=1)sparkrun appends --nnodes, --node-rank, --master-addr, --master-port; worker ranks get --headless.
Pure data parallelismdata_parallel: N (TP=PP=1)Each node is its own replica; sparkrun appends --data-parallel-size, --data-parallel-rank, --data-parallel-address, --data-parallel-rpc-port.
Hybrid (TP/PP × DP)tensor_parallel: T, data_parallel: Dsparkrun maps replicas to host groups of size T·PP; --master-addr is the first host of each replica; both flag sets are emitted.

Required node count = tensor_parallel × pipeline_parallel × data_parallel. The data_parallel_rpc_port defaults to 13345 and can be overridden in recipe defaults.

vllm-ray forms a Ray cluster first, then runs the serve command on the head node. sparkrun handles Ray cluster setup, InfiniBand detection, and container lifecycle automatically.

To use vllm-ray, either set runtime: vllm-ray explicitly or add the Ray hint to defaults:

runtime: vllm
defaults:
distributed_executor_backend: ray

These options apply only to vllm-ray:

OptionDefaultDescription
--ray-port46379Ray GCS port
--dashboardoffEnable Ray dashboard
--dashboard-port8265Ray dashboard port

Both variants support the same models and vLLM flags. The difference is in multi-node orchestration:

  • vllm-distributed is simpler — no Ray dependency, each node runs independently with vLLM’s built-in coordination. This is the default.
  • vllm-ray provides Ray’s cluster management, dashboard, and monitoring. Use it when you need Ray-specific features.

For single-node (solo) workloads, both variants behave identically.