Skip to content

Runtimes Overview

sparkrun supports multiple inference runtimes through a plugin system. Each runtime knows how to launch a specific inference engine, configure networking, and manage containers.

RuntimeEngineMulti-NodeUse Case
vllm-distributedvLLMYes (native)Default vLLM variant. Uses vLLM’s built-in distributed backend.
vllm-rayvLLMYes (Ray)vLLM with Ray head/worker orchestration.
sglangSGLangYes (native)High-throughput inference, structured generation, experimental GGUF
llama-cppllama.cppExperimentalGGUF quantized models, lightweight deployment
trtllmTensorRT-LLMYes (MPI)NVIDIA TensorRT-LLM with MPI-based multi-node orchestration
atlasAtlasYes (native)Pure-Rust inference server with native NCCL distribution. Composes tensor and expert parallelism.
modular-maxModular MAXNoSingle-node only. tensor_parallel maps to local --devices.

Explicitly setting runtime in recipes is recommended, but when it is omitted sparkrun can infer the runtime from the command field prefix:

Command prefixDetected runtime
vllm serve ...vllm
sglang serve ...sglang
python[3] -m sglang.launch_server ...sglang
llama-server ...llama-cpp
trtllm-serve ...trtllm
mpirun...trtllmtrtllm

If the command doesn’t match any known pattern (or there is no command), the recipe falls through to vllm behavior. An explicit runtime field always takes precedence over command-based detection.

See the recipe format docs for the full detection rules.

You can write runtime: vllm in a recipe (or let it be auto-detected from a vllm serve command) — sparkrun treats it as a virtual alias and resolves it to a concrete runtime:

  1. If the recipe hints at Ray usage (distributed_executor_backend: ray in defaults, or --distributed-executor-backend ray in the command template) → vllm-ray
  2. Otherwise → vllm-distributed (the default)

Both vllm-distributed and vllm-ray use the same container images and support the same vLLM flags — they differ only in how multi-node clustering is orchestrated.

Recipes declare which runtime to use via the runtime field. The runtime handles:

  • Container launch configuration (volumes, networking, GPU access)
  • InfiniBand/RDMA detection and NCCL environment variables
  • Multi-node clustering (native distributed for vllm-distributed and sglang, Ray for vllm-ray, MPI for trtllm)
  • Model pre-sync and cache path resolution
  • Command generation when no explicit command template is provided
  • vLLM is the default and most widely supported. Write runtime: vllm and sparkrun picks the right variant automatically (see the vllm alias above).
    • vllm-distributed (default) uses vLLM’s built-in distributed backend — each node runs the full serve command with node-rank arguments (--nnodes, --node-rank, --master-addr).
    • vllm-ray uses Ray head/worker orchestration — a Ray cluster is formed first, then the serve command runs on the head node.
  • SGLang offers competitive performance, good structured generation support, and experimental GGUF quantized model serving.
  • llama.cpp is ideal for GGUF quantized models or when you want a lightweight alternative.
  • TensorRT-LLM provides NVIDIA’s optimized inference engine with MPI-based multi-node orchestration.
  • Atlas is a pure-Rust inference server (see Atlas on GitHub) with native NCCL distribution. It composes tensor parallelism (tensor_parallel) and expert parallelism (ep_size) on either an overlapping mesh (tp == ep) or an orthogonal one (tp * ep), and derives the world size from the two.
  • Modular MAX is single-node only: tensor_parallel maps to local --devices rather than to host count, and the runtime refuses a multi-host placement rather than silently truncating it.

The Spark Arena team maintains several specialized container images for DGX Spark.

ImageDescriptionLink
ghcr.io/spark-arena/dgx-vllm-eugr-nightlyOptimized eugr vLLM for DGX SparkGitHub
ghcr.io/spark-arena/dgx-vllm-eugr-nightly-tf5Optimized eugr vLLM for DGX Spark with Transformers 5.xGitHub
dgx-spark-sglangOptimized SGLang for DGX SparkLink
ghcr.io/spark-arena/dgx-llama-cppOptimized llama.cpp for DGX Spark
Mirror of scitrera/dgx-spark-llama-cpp
Docker
GitHub