Skip to content

sparkrun run

Terminal window
sparkrun run <recipe> [options]

Launch an inference workload on one or more DGX Spark systems. sparkrun resolves the recipe, distributes container images and models, configures networking, and starts containers on all target hosts.

Jobs launch in the background (detached containers) and sparkrun follows logs automatically. Ctrl+C detaches from logs — it never kills the inference job.

OptionDescription
--hosts / -HComma-separated host list (first = head)
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster by name
--portOverride serve port
--tp / --tensor-parallelOverride tensor parallelism
--pp / --pipeline-parallelOverride pipeline parallelism
--gpu-mem / --gpu-memory-utilization / --mem-fraction-staticOverride GPU memory utilization (0.0–1.0)
--max-model-lenOverride maximum model context length
--served-model-nameOverride served model name
--imageOverride container image (not recommended)
-o / --optionOverride any recipe default: -o key=value (repeatable)
--dry-run / -nShow what would be done without executing
--foregroundRun in foreground (don’t detach)
--ensureOnly launch if not already running; exit 0 if already up
--no-followDon’t follow container logs after launch
--no-sync-tuningSkip syncing tuning configs from registries
--no-rmDon’t auto-remove containers on exit (keeps containers after stop)
--memory-limitContainer memory limit (e.g. 32G)
--labelSet metadata on the container, e.g. --label com.example.key=value (repeatable)
--rootfulRun with --privileged as root inside container (legacy behavior)

Use -o key=value to override recipe defaults — that is the supported way to pass engine-level settings through to the serve command.

The <recipe> argument can be:

  • A local recipe name — matched against configured registries (e.g. qwen3-1.7b-vllm)
  • A file path — a local .yaml file (e.g. ./my-recipe.yaml)
  • A Spark Arena shortcut — @spark-arena/<recipe-id> to run directly from Spark Arena
  • A URL — any https:// URL pointing to a recipe YAML
Terminal window
# Single node
sparkrun run qwen3-1.7b-vllm
# Two-node tensor parallel
sparkrun run qwen3-1.7b-vllm --tp 2
# Specific hosts with overrides
sparkrun run qwen3-1.7b-vllm -H 192.168.11.13,192.168.11.14 -o max_model_len=8192
# Run a recipe from Spark Arena
sparkrun run @spark-arena/<recipe-id>
# Dry run to preview commands
sparkrun run nemotron3-nano-30b-nvfp4-vllm --tp 2 --dry-run

These options are hidden from --help but available for advanced use cases:

OptionDescription
--soloForce single-host mode regardless of how many hosts are available
--dp / --data-parallelOverride data parallelism
--schedulerPlacement scheduler: occupancy-sparse, occupancy-dense, or greedy. See Schedulers & Placement.
--transfer-modeResource transfer mode: auto, local, push, or delegated (overrides the cluster setting)
--collect-diagnostics <path>Capture the full run lifecycle into an NDJSON file for post-mortem analysis. See Run Diagnostics.
--restart <policy>Docker restart policy (no, always, unless-stopped, on-failure[:N])
--trustTrust lifecycle hooks from third-party registries without confirmation. See Security & Trust.
--rebuild / --no-rebuildForce the builder to produce a fresh image. No-op for docker-pull; forces a rebuild or fresh pull for eugr. Overrides the recipe’s builder_config.rebuild.
--container-nameOverride the deterministic cluster ID (i.e. use a static container name)
--executor-argsArguments passed directly to the container executor, e.g. docker run (repeatable)
--ray-portRay GCS port — default 46379 (Ray runtimes only)
--init-portvLLM/SGLang distributed init port — default 25000
--dashboard / --no-dashboardEnable or disable the Ray dashboard on the head node (Ray runtimes only; binds 0.0.0.0 when on)
--dashboard-portRay dashboard port — default 8265
Terminal window
# Capture full run diagnostics for debugging
sparkrun run my-recipe --collect-diagnostics run_diag.ndjson
# Set container restart policy
sparkrun run my-recipe --restart unless-stopped
# Pin placement behavior for this launch
sparkrun run my-recipe --scheduler occupancy-dense