Skip to content

Setup Commands

Terminal window
sparkrun setup wizard [options]

Guided interactive wizard that walks through the full cluster setup process in one session. When you run bare sparkrun setup and no default cluster exists, the wizard launches automatically.

Phases:

  1. Install check — detects whether sparkrun is installed as a uv tool; installs it if not
  2. Cluster creation — discovers CX7 peers (when running on a DGX Spark), prompts for hosts, creates/updates a named cluster, and sets it as default
  3. SSH access — every later phase assumes passwordless control→host SSH, so the wizard establishes it first: it probes each host, and for hosts that answer but reject the key it offers to install one (generating ~/.ssh/sparkrun_ed25519 if you have no identity). Runs right after the cluster-name and SSH-username prompts, before any other probe.
  4. SSH mesh — establishes passwordless SSH between all cluster hosts as well as the control machine
  5. CX7 configuration — detects ConnectX-7 interfaces, selects conflict-free subnets, assigns static IPs with jumbo frames (MTU 9000), and distributes host keys. If CX7 IPs are added or changed, the SSH mesh is automatically re-run to ensure full connectivity across all interfaces.
  6. Docker group — ensures the SSH user is in the docker group on all hosts so containers can run without sudo
  7. NVIDIA CDI — generates /etc/cdi/nvidia.yaml so Docker can resolve the GPU. sparkrun requests GPUs via CDI rather than --gpus.
  8. Sudoers entries — installs scoped sudoers rules for fix-permissions and clear-cache (no broad sudo)
  9. earlyoom — installs and configures OOM protection with inference-aware defaults

Each phase is optional — you can skip any step when prompted, and configure it later with the individual subcommand.

Terminal window
sparkrun setup wizard
sparkrun setup wizard --hosts 10.24.11.13,10.24.11.14 --cluster mylab
sparkrun setup wizard --dry-run
sparkrun setup wizard --yes --hosts 10.24.11.13,10.24.11.14
OptionDescription
--hosts / -HPre-populate host list (comma-separated)
--clusterCluster name (default: prompted, or default with --yes)
--user / -uSSH username (default: current OS user)
--dry-run / -nPreview all phases without executing
--yes / -yAccept all defaults for non-interactive mode
Terminal window
sparkrun setup check [--cluster mylab] [--json]

Probes each host for the things the wizard configures and reports gaps — without changing anything. Use it to verify a cluster before a first run, or to find what drifted after one.

Checks, in order: docker_installed, docker_group, docker_usable, nvidia_ctk, cdi_spec, host_ipc, earlyoom, sudoers, ssh_mesh, cx7, rdma.

Each result is OK, WARN, FAIL, or SKIP, and a non-OK result names the command that would fix it. --json emits the same results machine-readably.

The rdma check reports whether the RDMA devices are present and ACTIVE. It never sends traffic over the fabric — “the link is configured” and “the link performs” are different questions, and the second one is sparkrun setup rdma-test.

Terminal window
sparkrun setup features list
sparkrun setup features enable executor.local
sparkrun setup features disable executor.local
sparkrun setup features reset executor.local

Manage the gates on experimental capabilities. See Feature Flags for the full list and resolution order.

Terminal window
sparkrun setup telemetry # show the current setting
sparkrun setup telemetry --disable # opt out
sparkrun setup telemetry --enable # opt back in

See Updates & Telemetry for what is collected.

Terminal window
sparkrun setup install

Installs sparkrun on PATH as a managed uv tool with tab completion. This is the recommended installation method.

Terminal window
sparkrun setup completion # auto-detects your shell
sparkrun setup completion --shell zsh # specify shell explicitly

After restarting your shell, recipe names, cluster names, and subcommands all tab-complete.

Terminal window
sparkrun update

Updates sparkrun to the latest version in your managed environment and refreshes all recipe registries. Use --no-update-registries to skip the registry update.

sparkrun update is a top-level alias for sparkrun setup update.

Terminal window
sparkrun setup ssh [options]

Set up passwordless SSH across your cluster hosts. See the SSH Setup guide for full details.

OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster
--user / -uSSH user (default: current OS user)
--extra-hostsAdditional hosts beyond the cluster
--no-include-selfExclude the control machine from the mesh
--discover-ips / --no-discover-ipsAfter meshing, discover IB/CX7 IPs and distribute host keys (default: on)
--diagnoseRun SSH diagnostics instead of mesh setup (see SSH Setup — Cross-user SSH)
--dry-run / -nShow what would be done without executing
Terminal window
sparkrun setup cx7 [options]

Configure ConnectX-7 network interfaces on cluster hosts for high-speed NCCL/RDMA communication. Detects CX-7 interfaces via SSH, selects two conflict-free /24 subnets, assigns static IPs with jumbo frames (MTU 9000), and applies netplan configuration.

Existing valid configurations are preserved unless --force is used. IP addresses are derived from each host’s management IP last octet (e.g., management 10.24.11.13 → CX-7 192.168.11.13 and 192.168.12.13).

Terminal window
sparkrun setup cx7 --hosts 10.24.11.13,10.24.11.14
sparkrun setup cx7 --cluster mylab --dry-run
sparkrun setup cx7 --cluster mylab --subnet1 192.168.11.0/24 --subnet2 192.168.12.0/24
sparkrun setup cx7 --cluster mylab --force
OptionDescription
--hosts / -HComma-separated host list (management IPs or hostnames)
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster definition
--user / -uSSH username (default: from cluster config or current user)
--subnet1Override subnet for CX-7 partition 1 (e.g., 192.168.11.0/24)
--subnet2Override subnet for CX-7 partition 2 (e.g., 192.168.12.0/24)
--mtuMTU for CX-7 interfaces (default: 9000)
--forceReconfigure even if existing config is valid
--dry-run / -nShow the plan without making changes
Terminal window
sparkrun setup rdma-test [options]

Measures what setup cx7 configured: RDMA latency and bandwidth per link, and optionally an NCCL collective across the cluster. Host pairs are derived from the configured CX-7 subnets, so only links that physically exist are tested.

By default it runs the perftest suite only — host-native on DGX OS, no image pull, done in seconds. Add --suite all for the NCCL collective, which fetches the test image onto every host and takes several minutes.

Terminal window
sparkrun setup rdma-test --cluster mylab
sparkrun setup rdma-test --cluster mylab --suite all
sparkrun setup rdma-test --cluster mylab --dry-run
sparkrun setup rdma-test --hosts 10.24.11.13,10.24.11.14 --json
OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster definition
--user / -uSSH username
--suiteperftest, nccl, or all (default: perftest)
--imageTest container image (default: the pinned sparkrun image)
--duration / -DSeconds per bandwidth test (default: 10)
--sizeNCCL message size (default: 16G)
--queue-pairs / -qperftest queue pairs (default: 4)
--gid-indexForce a RoCE GID index (normally automatic)
--link-typeperftest --force-link: IB, Ethernet, or auto
--force-containerUse the image even when perftest is on the host
--keep-containersLeave test containers running for debugging
--jsonEmit the full results as JSON
--dry-run / -nShow which links would be tested, without sending traffic

Each host pair is measured link by link and then with all its links driven concurrently — on DGX Spark a single QSFP112 cable presents as two RDMA devices, so a per-device figure is about half the cable’s real throughput.

Tooling. The perftest suite runs ib_write_bw / ib_write_lat from the host, which DGX OS ships preinstalled, so it pulls no image. The nccl suite needs ghcr.io/spark-arena/sparkrun-rdma-test, which bundles NCCL and nccl-tests; its tag tracks the NCCL release it was built from. Use --image to point at a mirror.

Exit status. Underperformance is a warning and exits 0; only a test that could not run at all is a failure and exits 1. See Networking Best Practices for the numbers to expect.

Terminal window
sparkrun setup fix-permissions [options]

Fix file ownership in the HuggingFace cache on cluster hosts. Docker containers create cache files as root, leaving the normal user unable to manage them. This command runs chown across all target hosts to restore ownership.

Tries non-interactive sudo first on all hosts in parallel, then falls back to password-based sudo for any that fail. Use --save-sudo to avoid password prompts on subsequent runs.

Terminal window
sparkrun setup fix-permissions --cluster mylab
sparkrun setup fix-permissions --cluster mylab --cache-dir /data/hf-cache
sparkrun setup fix-permissions --cluster mylab --save-sudo
sparkrun setup fix-permissions --cluster mylab --dry-run
OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster definition
--user / -uTarget owner for chown (default: SSH user)
--cache-dirCache directory path (default: HF_HOME, typically ~/.cache/huggingface)
--save-sudoInstall a scoped sudoers entry for passwordless chown (requires sudo once)
--dry-run / -nShow what would be done without executing

Use --save-sudo to install a minimal sudoers rule that permits only chown on the cache directory — no broader privileges are granted. After this, sparkrun automatically fixes cache ownership before every model distribution without prompting for a password.

See HuggingFace cache owned by root for background and prevention tips.

Terminal window
sparkrun setup clear-cache [options]

Drop the Linux page cache on cluster hosts. Runs sync followed by writing 3 to /proc/sys/vm/drop_caches on each host, freeing cached file data so inference containers have maximum available memory on DGX Spark’s unified CPU/GPU memory.

Tries non-interactive sudo first on all hosts in parallel, then falls back to password-based sudo for any that fail. Use --save-sudo to avoid password prompts on subsequent runs.

Terminal window
sparkrun setup clear-cache --cluster mylab
sparkrun setup clear-cache --cluster mylab --save-sudo
sparkrun setup clear-cache --cluster mylab --dry-run
OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster definition
--user / -uTarget user for sudoers entry (default: SSH user)
--save-sudoInstall a scoped sudoers entry for passwordless cache clearing (requires sudo once)
--dry-run / -nShow what would be done without executing

Use --save-sudo to install a minimal sudoers rule that permits only writing to /proc/sys/vm/drop_caches — no broader privileges are granted. After this, sparkrun automatically clears the page cache before every container launch without prompting for a password.

See Linux page cache consuming memory for background on why this matters.

Terminal window
sparkrun setup earlyoom [options]

Install and configure earlyoom on cluster hosts. earlyoom monitors available memory and proactively kills processes before the kernel OOM killer triggers, preventing system hangs when large inference models exhaust DGX Spark’s unified CPU/GPU memory.

sparkrun configures earlyoom with inference-aware defaults:

  • Prefer killing (on OOM): vllm, sglang, llama-server, llama-cli, trtllm, tritonserver, python
  • Avoid killing (protect): systemd, sshd, dockerd, containerd, dbus-daemon, NetworkManager
  • Trigger: 2% available memory (~2.5 GB on 128 GB DGX Spark), swap below 80% (emphasis is on unified memory)
Terminal window
sparkrun setup earlyoom --cluster mylab
sparkrun setup earlyoom --hosts 10.24.11.13,10.24.11.14
sparkrun setup earlyoom --cluster mylab --prefer "my-worker,my-app"
sparkrun setup earlyoom --cluster mylab --dry-run
OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster definition
--user / -uSSH username (default: from cluster config or current user)
--preferAdditional comma-separated process patterns to prefer killing on OOM
--avoidAdditional comma-separated process patterns to avoid killing on OOM
--dry-run / -nShow what would be done without executing

The command installs earlyoom via apt-get, writes /etc/default/earlyoom with the configured --prefer and --avoid patterns, sets up a systemd override for CAP_KILL capabilities, and enables the service. Requires sudo on target hosts — you may be prompted for a sudo password if passwordless sudo is not configured.

Use --prefer and --avoid to add custom process patterns to the defaults. For example, --prefer "my-worker" adds my-worker to the list of processes earlyoom will kill first on OOM.

Terminal window
sparkrun setup diagnose [options]

Collect hardware, firmware, network, and Docker diagnostics from cluster hosts. Useful for debugging issues, filing support tickets, or verifying hardware configuration across your cluster.

Collected information:

  • OS, kernel, CPU, memory, and disk details
  • GPU information (nvidia-smi)
  • Network interfaces and configuration
  • Docker status and version
  • Firmware/BIOS data (with --sudo — uses dmidecode)

Output is written as NDJSON (one JSON record per line) for easy parsing and processing.

Terminal window
sparkrun setup diagnose --cluster mylab
sparkrun setup diagnose --hosts 10.24.11.13,10.24.11.14
sparkrun setup diagnose --cluster mylab --output diag.ndjson
sparkrun setup diagnose --cluster mylab --sudo
sparkrun setup diagnose --cluster mylab --json
OptionDescription
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster definition
--output / -oOutput NDJSON file path (default: spark_diag_<timestamp>.ndjson)
--jsonAlso print a summary to stdout as JSON
--sudoCollect additional sudo-only diagnostics (dmidecode for BIOS, system, baseboard, memory details)
--dry-run / -nShow what would be done without executing