Skip to content

Cluster Commands

Clusters are named groups of DGX Spark hosts. Save your hosts once and reference them by name in every command.

The first host in a cluster definition is used as the head node for multi-node jobs.

Terminal window
sparkrun cluster create <name> --hosts <ip1>,<ip2>,... [options]
OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line)
-d / --descriptionCluster description
--user / -uSSH username for this cluster
--cache-dirHuggingFace cache directory on cluster hosts (useful when control machine and cluster have different paths, e.g. macOS → Linux)
--transfer-modeResource transfer mode: auto, local, push, or delegated
--transfer-interfaceNetwork interface for transfers: auto, cx7, or mgmt
--schedulerPlacement scheduler (default occupancy-sparse). See Schedulers & Placement.
--max-gpu-mem-utilFraction of each accelerator’s memory treated as usable for placement
--executorPin the launch executor for this cluster
--executor-opt / -oPer-executor config, -o key=value (repeatable)
--defaultSet this cluster as the default after creating it
Terminal window
sparkrun cluster create mylab \
--hosts 10.24.11.13,10.24.11.14 \
--user dgxuser \
-d "2-node DGX Spark cluster"
# With transfer mode for external control machines
sparkrun cluster create mylab \
--hosts 10.24.11.13,10.24.11.14 \
--transfer-mode push

See Resource Transfer Modes for details on transfer mode options.

Terminal window
sparkrun cluster update <name> [options]

All options are the same as create. Only provided options are updated — omitted fields are left unchanged. Pass --user "" or --cache-dir "" to explicitly clear a value.

update additionally accepts:

OptionDescription
--add-hostAdd a host without restating the whole list (repeatable)
--remove-hostRemove a host (repeatable)
--topologySet the cluster’s network topology
--infer-hardwareProbe hosts and record detected accelerator metadata
--clear-executor-configDrop any stored per-executor config
Terminal window
sparkrun cluster update mylab --hosts 10.24.11.13,10.24.11.14,10.24.11.15
sparkrun cluster update mylab --add-host 10.24.11.15
sparkrun cluster update mylab --transfer-mode delegated
sparkrun cluster update mylab --user newuser
# Adopt the 0.3.x default placement on a cluster created by sparkrun 0.2.x
sparkrun cluster update mylab --scheduler occupancy-sparse

See Resource Transfer Modes for details on transfer mode options.

Terminal window
sparkrun cluster list

Shows all saved clusters with their hosts and descriptions. The default cluster is marked with *.

Terminal window
sparkrun cluster show <name>

Displays name, description, user, cache directory, transfer mode, default status, and all hosts.

Terminal window
sparkrun cluster delete <name>
sparkrun cluster delete <name> --force # skip confirmation prompt
Terminal window
sparkrun cluster inspect [name]

Shows the resolved cluster settings — transfer mode, transfer interface, topology, SSH user, and cache directories — and checks whether those cache directories actually exist on each remote host. Useful for diagnosing configuration, transfer, or permission problems without running a job.

Terminal window
sparkrun cluster import svd <path-to-.env>

Imports a spark-vllm-docker .env file into a sparkrun cluster, mapping CLUSTER_NODES to hosts, ETH_IF to fabric interfaces, and CONTAINER_* to env references. The env references are resolved from the file at launch time, so secrets stay out of the cluster YAML. Re-running syncs in place. Aliased as sparkrun cluster import eugr.

Terminal window
sparkrun cluster set-default <name>
sparkrun cluster unset-default
sparkrun cluster default # show current default

When a default cluster is set, commands like sparkrun run and sparkrun stop use it automatically without needing --cluster or --hosts.

Terminal window
sparkrun cluster status
sparkrun status # alias

Shows running sparkrun containers, pending operations (downloads/distributions in progress), and idle hosts across the cluster. See Cluster Status for full details.

$ sparkrun cluster status
Job: qwen3.5-35b-a3b-fp8-sglang (tp=2) (2 container(s))
node_0 10.24.11.13 (ib: 192.168.11.13) Up 35 seconds scitrera/dgx-spark-sglang:0.5.9-dev1-329817e2-t5
node_1 10.24.11.14 (ib: 192.168.11.14) Up 9 seconds scitrera/dgx-spark-sglang:0.5.9-dev1-329817e2-t5
logs: sparkrun logs qwen3.5-35b-a3b-fp8-sglang --hosts 10.24.11.13,10.24.11.14 --tp 2
stop: sparkrun stop qwen3.5-35b-a3b-fp8-sglang --hosts 10.24.11.13,10.24.11.14 --tp 2
Idle hosts (no sparkrun containers):
10.24.11.16
10.24.11.17
Total: 2 container(s) across 4 host(s)
Terminal window
sparkrun cluster monitor [options]

Live-monitor CPU, RAM, and GPU metrics across cluster hosts. By default launches an interactive TUI; pass --simple for plain-text output. Press q (TUI) or Ctrl-C to stop.

Cluster monitor TUI

OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line)
--clusterUse a saved cluster by name
--interval / -iSampling interval in seconds (default: 2)
--simpleUse plain-text output instead of TUI
--jsonEmit telemetry samples as JSON
--dry-run / -nShow what would be done
Terminal window
sparkrun cluster monitor --cluster mylab
sparkrun cluster monitor --cluster mylab --interval 5
sparkrun cluster monitor --cluster mylab --simple