Runtimes Overview
sparkrun supports multiple inference runtimes through a plugin system. Each runtime knows how to launch a specific inference engine, configure networking, and manage containers.
Runtime comparison
Section titled “Runtime comparison”| Runtime | Engine | Multi-Node | Use Case |
|---|---|---|---|
vllm-distributed | vLLM | Yes (native) | Default vLLM variant. Uses vLLM’s built-in distributed backend. |
vllm-ray | vLLM | Yes (Ray) | vLLM with Ray head/worker orchestration. |
sglang | SGLang | Yes (native) | High-throughput inference, structured generation, experimental GGUF |
llama-cpp | llama.cpp | Experimental | GGUF quantized models, lightweight deployment |
trtllm | TensorRT-LLM | Yes (MPI) | NVIDIA TensorRT-LLM with MPI-based multi-node orchestration |
atlas | Atlas | Yes (native) | Pure-Rust inference server with native NCCL distribution. Composes tensor and expert parallelism. |
modular-max | Modular MAX | No | Single-node only. tensor_parallel maps to local --devices. |
Automatic runtime detection
Section titled “Automatic runtime detection”Explicitly setting runtime in recipes is recommended, but when it is omitted sparkrun can infer the runtime from the command field prefix:
| Command prefix | Detected runtime |
|---|---|
vllm serve ... | vllm |
sglang serve ... | sglang |
python[3] -m sglang.launch_server ... | sglang |
llama-server ... | llama-cpp |
trtllm-serve ... | trtllm |
mpirun...trtllm | trtllm |
If the command doesn’t match any known pattern (or there is no command), the recipe falls through to vllm behavior. An explicit runtime field always takes precedence over command-based detection.
See the recipe format docs for the full detection rules.
The vllm alias
Section titled “The vllm alias”You can write runtime: vllm in a recipe (or let it be auto-detected from a vllm serve command) — sparkrun treats it as a virtual alias and resolves it to a concrete runtime:
- If the recipe hints at Ray usage (
distributed_executor_backend: rayin defaults, or--distributed-executor-backend rayin the command template) →vllm-ray - Otherwise →
vllm-distributed(the default)
Both vllm-distributed and vllm-ray use the same container images and support the same vLLM flags — they differ only in how multi-node clustering is orchestrated.
How runtimes work
Section titled “How runtimes work”Recipes declare which runtime to use via the runtime field. The runtime handles:
- Container launch configuration (volumes, networking, GPU access)
- InfiniBand/RDMA detection and NCCL environment variables
- Multi-node clustering (native distributed for
vllm-distributedandsglang, Ray forvllm-ray, MPI fortrtllm) - Model pre-sync and cache path resolution
- Command generation when no explicit
commandtemplate is provided
Choosing a runtime
Section titled “Choosing a runtime”- vLLM is the default and most widely supported. Write
runtime: vllmand sparkrun picks the right variant automatically (see thevllmalias above).- vllm-distributed (default) uses vLLM’s built-in distributed backend — each node runs the full serve command with node-rank arguments (
--nnodes,--node-rank,--master-addr). - vllm-ray uses Ray head/worker orchestration — a Ray cluster is formed first, then the serve command runs on the head node.
- vllm-distributed (default) uses vLLM’s built-in distributed backend — each node runs the full serve command with node-rank arguments (
- SGLang offers competitive performance, good structured generation support, and experimental GGUF quantized model serving.
- llama.cpp is ideal for GGUF quantized models or when you want a lightweight alternative.
- TensorRT-LLM provides NVIDIA’s optimized inference engine with MPI-based multi-node orchestration.
- Atlas is a pure-Rust inference server (see Atlas on GitHub) with native NCCL distribution. It composes tensor parallelism (
tensor_parallel) and expert parallelism (ep_size) on either an overlapping mesh (tp == ep) or an orthogonal one (tp * ep), and derives the world size from the two. - Modular MAX is single-node only:
tensor_parallelmaps to local--devicesrather than to host count, and the runtime refuses a multi-host placement rather than silently truncating it.
Maintained container images
Section titled “Maintained container images”The Spark Arena team maintains several specialized container images for DGX Spark.
| Image | Description | Link |
|---|---|---|
ghcr.io/spark-arena/dgx-vllm-eugr-nightly | Optimized eugr vLLM for DGX Spark | GitHub |
ghcr.io/spark-arena/dgx-vllm-eugr-nightly-tf5 | Optimized eugr vLLM for DGX Spark with Transformers 5.x | GitHub |
dgx-spark-sglang | Optimized SGLang for DGX Spark | Link |
ghcr.io/spark-arena/dgx-llama-cpp | Optimized llama.cpp for DGX Spark Mirror of scitrera/dgx-spark-llama-cpp | Docker GitHub |