sparkrun run
sparkrun run <recipe> [options]Launch an inference workload on one or more DGX Spark systems. sparkrun resolves the recipe, distributes container images and models, configures networking, and starts containers on all target hosts.
Jobs launch in the background (detached containers) and sparkrun follows logs automatically. Ctrl+C detaches from logs — it never kills the inference job.
Options
Section titled “Options”| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list (first = head) |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster by name |
--port | Override serve port |
--tp / --tensor-parallel | Override tensor parallelism |
--pp / --pipeline-parallel | Override pipeline parallelism |
--gpu-mem / --gpu-memory-utilization / --mem-fraction-static | Override GPU memory utilization (0.0–1.0) |
--max-model-len | Override maximum model context length |
--served-model-name | Override served model name |
--image | Override container image (not recommended) |
-o / --option | Override any recipe default: -o key=value (repeatable) |
--dry-run / -n | Show what would be done without executing |
--foreground | Run in foreground (don’t detach) |
--ensure | Only launch if not already running; exit 0 if already up |
--no-follow | Don’t follow container logs after launch |
--no-sync-tuning | Skip syncing tuning configs from registries |
--no-rm | Don’t auto-remove containers on exit (keeps containers after stop) |
--memory-limit | Container memory limit (e.g. 32G) |
--label | Set metadata on the container, e.g. --label com.example.key=value (repeatable) |
--rootful | Run with --privileged as root inside container (legacy behavior) |
Use -o key=value to override recipe defaults — that is the supported way to
pass engine-level settings through to the serve command.
Recipe sources
Section titled “Recipe sources”The <recipe> argument can be:
- A local recipe name — matched against configured registries (e.g.
qwen3-1.7b-vllm) - A file path — a local
.yamlfile (e.g../my-recipe.yaml) - A Spark Arena shortcut —
@spark-arena/<recipe-id>to run directly from Spark Arena - A URL — any
https://URL pointing to a recipe YAML
Examples
Section titled “Examples”# Single nodesparkrun run qwen3-1.7b-vllm
# Two-node tensor parallelsparkrun run qwen3-1.7b-vllm --tp 2
# Specific hosts with overridessparkrun run qwen3-1.7b-vllm -H 192.168.11.13,192.168.11.14 -o max_model_len=8192
# Run a recipe from Spark Arenasparkrun run @spark-arena/<recipe-id>
# Dry run to preview commandssparkrun run nemotron3-nano-30b-nvfp4-vllm --tp 2 --dry-runAdvanced options
Section titled “Advanced options”These options are hidden from --help but available for advanced use cases:
| Option | Description |
|---|---|
--solo | Force single-host mode regardless of how many hosts are available |
--dp / --data-parallel | Override data parallelism |
--scheduler | Placement scheduler: occupancy-sparse, occupancy-dense, or greedy. See Schedulers & Placement. |
--transfer-mode | Resource transfer mode: auto, local, push, or delegated (overrides the cluster setting) |
--collect-diagnostics <path> | Capture the full run lifecycle into an NDJSON file for post-mortem analysis. See Run Diagnostics. |
--restart <policy> | Docker restart policy (no, always, unless-stopped, on-failure[:N]) |
--trust | Trust lifecycle hooks from third-party registries without confirmation. See Security & Trust. |
--rebuild / --no-rebuild | Force the builder to produce a fresh image. No-op for docker-pull; forces a rebuild or fresh pull for eugr. Overrides the recipe’s builder_config.rebuild. |
--container-name | Override the deterministic cluster ID (i.e. use a static container name) |
--executor-args | Arguments passed directly to the container executor, e.g. docker run (repeatable) |
--ray-port | Ray GCS port — default 46379 (Ray runtimes only) |
--init-port | vLLM/SGLang distributed init port — default 25000 |
--dashboard / --no-dashboard | Enable or disable the Ray dashboard on the head node (Ray runtimes only; binds 0.0.0.0 when on) |
--dashboard-port | Ray dashboard port — default 8265 |
# Capture full run diagnostics for debuggingsparkrun run my-recipe --collect-diagnostics run_diag.ndjson
# Set container restart policysparkrun run my-recipe --restart unless-stopped
# Pin placement behavior for this launchsparkrun run my-recipe --scheduler occupancy-dense