Setup Commands
Setup wizard
Section titled “Setup wizard”sparkrun setup wizard [options]Guided interactive wizard that walks through the full cluster setup process in one session. When you run bare sparkrun setup and no default cluster exists, the wizard launches automatically.
Phases:
- Install check — detects whether sparkrun is installed as a uv tool; installs it if not
- Cluster creation — discovers CX7 peers (when running on a DGX Spark), prompts for hosts, creates/updates a named cluster, and sets it as default
- SSH access — every later phase assumes passwordless control→host SSH, so the wizard establishes it first: it probes each host, and for hosts that answer but reject the key it offers to install one (generating
~/.ssh/sparkrun_ed25519if you have no identity). Runs right after the cluster-name and SSH-username prompts, before any other probe. - SSH mesh — establishes passwordless SSH between all cluster hosts as well as the control machine
- CX7 configuration — detects ConnectX-7 interfaces, selects conflict-free subnets, assigns static IPs with jumbo frames (MTU 9000), and distributes host keys. If CX7 IPs are added or changed, the SSH mesh is automatically re-run to ensure full connectivity across all interfaces.
- Docker group — ensures the SSH user is in the
dockergroup on all hosts so containers can run without sudo - NVIDIA CDI — generates
/etc/cdi/nvidia.yamlso Docker can resolve the GPU. sparkrun requests GPUs via CDI rather than--gpus. - Sudoers entries — installs scoped sudoers rules for
fix-permissionsandclear-cache(no broad sudo) - earlyoom — installs and configures OOM protection with inference-aware defaults
Each phase is optional — you can skip any step when prompted, and configure it later with the individual subcommand.
sparkrun setup wizardsparkrun setup wizard --hosts 10.24.11.13,10.24.11.14 --cluster mylabsparkrun setup wizard --dry-runsparkrun setup wizard --yes --hosts 10.24.11.13,10.24.11.14| Option | Description |
|---|---|
--hosts / -H | Pre-populate host list (comma-separated) |
--cluster | Cluster name (default: prompted, or default with --yes) |
--user / -u | SSH username (default: current OS user) |
--dry-run / -n | Preview all phases without executing |
--yes / -y | Accept all defaults for non-interactive mode |
Check readiness
Section titled “Check readiness”sparkrun setup check [--cluster mylab] [--json]Probes each host for the things the wizard configures and reports gaps — without changing anything. Use it to verify a cluster before a first run, or to find what drifted after one.
Checks, in order: docker_installed, docker_group, docker_usable,
nvidia_ctk, cdi_spec, host_ipc, earlyoom, sudoers, ssh_mesh, cx7,
rdma.
Each result is OK, WARN, FAIL, or SKIP, and a non-OK result names the
command that would fix it. --json emits the same results machine-readably.
The rdma check reports whether the RDMA devices are present and ACTIVE. It
never sends traffic over the fabric — “the link is configured” and “the link
performs” are different questions, and the second one is
sparkrun setup rdma-test.
Feature flags
Section titled “Feature flags”sparkrun setup features listsparkrun setup features enable executor.localsparkrun setup features disable executor.localsparkrun setup features reset executor.localManage the gates on experimental capabilities. See Feature Flags for the full list and resolution order.
Telemetry
Section titled “Telemetry”sparkrun setup telemetry # show the current settingsparkrun setup telemetry --disable # opt outsparkrun setup telemetry --enable # opt back inSee Updates & Telemetry for what is collected.
Install sparkrun
Section titled “Install sparkrun”sparkrun setup installInstalls sparkrun on PATH as a managed uv tool with tab completion. This is the recommended installation method.
Shell tab-completion
Section titled “Shell tab-completion”sparkrun setup completion # auto-detects your shellsparkrun setup completion --shell zsh # specify shell explicitlyAfter restarting your shell, recipe names, cluster names, and subcommands all tab-complete.
Update sparkrun
Section titled “Update sparkrun”sparkrun updateUpdates sparkrun to the latest version in your managed environment and refreshes all recipe registries. Use --no-update-registries to skip the registry update.
sparkrun update is a top-level alias for sparkrun setup update.
SSH mesh setup
Section titled “SSH mesh setup”sparkrun setup ssh [options]Set up passwordless SSH across your cluster hosts. See the SSH Setup guide for full details.
| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster |
--user / -u | SSH user (default: current OS user) |
--extra-hosts | Additional hosts beyond the cluster |
--no-include-self | Exclude the control machine from the mesh |
--discover-ips / --no-discover-ips | After meshing, discover IB/CX7 IPs and distribute host keys (default: on) |
--diagnose | Run SSH diagnostics instead of mesh setup (see SSH Setup — Cross-user SSH) |
--dry-run / -n | Show what would be done without executing |
CX-7 network configuration
Section titled “CX-7 network configuration”sparkrun setup cx7 [options]Configure ConnectX-7 network interfaces on cluster hosts for high-speed NCCL/RDMA communication. Detects CX-7 interfaces via SSH, selects two conflict-free /24 subnets, assigns static IPs with jumbo frames (MTU 9000), and applies netplan configuration.
Existing valid configurations are preserved unless --force is used. IP addresses are derived from each host’s management IP last octet (e.g., management 10.24.11.13 → CX-7 192.168.11.13 and 192.168.12.13).
sparkrun setup cx7 --hosts 10.24.11.13,10.24.11.14sparkrun setup cx7 --cluster mylab --dry-runsparkrun setup cx7 --cluster mylab --subnet1 192.168.11.0/24 --subnet2 192.168.12.0/24sparkrun setup cx7 --cluster mylab --force| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list (management IPs or hostnames) |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster definition |
--user / -u | SSH username (default: from cluster config or current user) |
--subnet1 | Override subnet for CX-7 partition 1 (e.g., 192.168.11.0/24) |
--subnet2 | Override subnet for CX-7 partition 2 (e.g., 192.168.12.0/24) |
--mtu | MTU for CX-7 interfaces (default: 9000) |
--force | Reconfigure even if existing config is valid |
--dry-run / -n | Show the plan without making changes |
RDMA performance test
Section titled “RDMA performance test”sparkrun setup rdma-test [options]Measures what setup cx7 configured: RDMA latency and bandwidth per link, and optionally an NCCL
collective across the cluster. Host pairs are derived from the configured CX-7 subnets, so only
links that physically exist are tested.
By default it runs the perftest suite only — host-native on DGX OS, no image pull, done in
seconds. Add --suite all for the NCCL collective, which fetches the test image onto every host and
takes several minutes.
sparkrun setup rdma-test --cluster mylabsparkrun setup rdma-test --cluster mylab --suite allsparkrun setup rdma-test --cluster mylab --dry-runsparkrun setup rdma-test --hosts 10.24.11.13,10.24.11.14 --json| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster definition |
--user / -u | SSH username |
--suite | perftest, nccl, or all (default: perftest) |
--image | Test container image (default: the pinned sparkrun image) |
--duration / -D | Seconds per bandwidth test (default: 10) |
--size | NCCL message size (default: 16G) |
--queue-pairs / -q | perftest queue pairs (default: 4) |
--gid-index | Force a RoCE GID index (normally automatic) |
--link-type | perftest --force-link: IB, Ethernet, or auto |
--force-container | Use the image even when perftest is on the host |
--keep-containers | Leave test containers running for debugging |
--json | Emit the full results as JSON |
--dry-run / -n | Show which links would be tested, without sending traffic |
Each host pair is measured link by link and then with all its links driven concurrently — on DGX Spark a single QSFP112 cable presents as two RDMA devices, so a per-device figure is about half the cable’s real throughput.
Tooling. The perftest suite runs ib_write_bw / ib_write_lat from the host, which DGX OS
ships preinstalled, so it pulls no image. The nccl suite needs
ghcr.io/spark-arena/sparkrun-rdma-test, which bundles NCCL and nccl-tests; its tag tracks the NCCL
release it was built from. Use --image to point at a mirror.
Exit status. Underperformance is a warning and exits 0; only a test that could not run at all
is a failure and exits 1. See Networking Best Practices for the
numbers to expect.
Fix cache permissions
Section titled “Fix cache permissions”sparkrun setup fix-permissions [options]Fix file ownership in the HuggingFace cache on cluster hosts. Docker containers create cache files as root, leaving the normal user unable to manage them. This command runs chown across all target hosts to restore ownership.
Tries non-interactive sudo first on all hosts in parallel, then falls back to password-based sudo for any that fail. Use --save-sudo to avoid password prompts on subsequent runs.
sparkrun setup fix-permissions --cluster mylabsparkrun setup fix-permissions --cluster mylab --cache-dir /data/hf-cachesparkrun setup fix-permissions --cluster mylab --save-sudosparkrun setup fix-permissions --cluster mylab --dry-run| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster definition |
--user / -u | Target owner for chown (default: SSH user) |
--cache-dir | Cache directory path (default: HF_HOME, typically ~/.cache/huggingface) |
--save-sudo | Install a scoped sudoers entry for passwordless chown (requires sudo once) |
--dry-run / -n | Show what would be done without executing |
Use --save-sudo to install a minimal sudoers rule that permits only chown on the cache directory — no broader privileges are granted. After this, sparkrun automatically fixes cache ownership before every model distribution without prompting for a password.
See HuggingFace cache owned by root for background and prevention tips.
Clear page cache
Section titled “Clear page cache”sparkrun setup clear-cache [options]Drop the Linux page cache on cluster hosts. Runs sync followed by writing 3 to /proc/sys/vm/drop_caches on each host, freeing cached file data so inference containers have maximum available memory on DGX Spark’s unified CPU/GPU memory.
Tries non-interactive sudo first on all hosts in parallel, then falls back to password-based sudo for any that fail. Use --save-sudo to avoid password prompts on subsequent runs.
sparkrun setup clear-cache --cluster mylabsparkrun setup clear-cache --cluster mylab --save-sudosparkrun setup clear-cache --cluster mylab --dry-run| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster definition |
--user / -u | Target user for sudoers entry (default: SSH user) |
--save-sudo | Install a scoped sudoers entry for passwordless cache clearing (requires sudo once) |
--dry-run / -n | Show what would be done without executing |
Use --save-sudo to install a minimal sudoers rule that permits only writing to /proc/sys/vm/drop_caches — no broader privileges are granted. After this, sparkrun automatically clears the page cache before every container launch without prompting for a password.
See Linux page cache consuming memory for background on why this matters.
Install earlyoom OOM killer
Section titled “Install earlyoom OOM killer”sparkrun setup earlyoom [options]Install and configure earlyoom on cluster hosts. earlyoom monitors available memory and proactively kills processes before the kernel OOM killer triggers, preventing system hangs when large inference models exhaust DGX Spark’s unified CPU/GPU memory.
sparkrun configures earlyoom with inference-aware defaults:
- Prefer killing (on OOM): vllm, sglang, llama-server, llama-cli, trtllm, tritonserver, python
- Avoid killing (protect): systemd, sshd, dockerd, containerd, dbus-daemon, NetworkManager
- Trigger: 2% available memory (~2.5 GB on 128 GB DGX Spark), swap below 80% (emphasis is on unified memory)
sparkrun setup earlyoom --cluster mylabsparkrun setup earlyoom --hosts 10.24.11.13,10.24.11.14sparkrun setup earlyoom --cluster mylab --prefer "my-worker,my-app"sparkrun setup earlyoom --cluster mylab --dry-run| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster definition |
--user / -u | SSH username (default: from cluster config or current user) |
--prefer | Additional comma-separated process patterns to prefer killing on OOM |
--avoid | Additional comma-separated process patterns to avoid killing on OOM |
--dry-run / -n | Show what would be done without executing |
The command installs earlyoom via apt-get, writes /etc/default/earlyoom with the configured --prefer and --avoid patterns, sets up a systemd override for CAP_KILL capabilities, and enables the service. Requires sudo on target hosts — you may be prompted for a sudo password if passwordless sudo is not configured.
Use --prefer and --avoid to add custom process patterns to the defaults. For example, --prefer "my-worker" adds my-worker to the list of processes earlyoom will kill first on OOM.
Collect diagnostics
Section titled “Collect diagnostics”sparkrun setup diagnose [options]Collect hardware, firmware, network, and Docker diagnostics from cluster hosts. Useful for debugging issues, filing support tickets, or verifying hardware configuration across your cluster.
Collected information:
- OS, kernel, CPU, memory, and disk details
- GPU information (nvidia-smi)
- Network interfaces and configuration
- Docker status and version
- Firmware/BIOS data (with
--sudo— uses dmidecode)
Output is written as NDJSON (one JSON record per line) for easy parsing and processing.
sparkrun setup diagnose --cluster mylabsparkrun setup diagnose --hosts 10.24.11.13,10.24.11.14sparkrun setup diagnose --cluster mylab --output diag.ndjsonsparkrun setup diagnose --cluster mylab --sudosparkrun setup diagnose --cluster mylab --json| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster definition |
--output / -o | Output NDJSON file path (default: spark_diag_<timestamp>.ndjson) |
--json | Also print a summary to stdout as JSON |
--sudo | Collect additional sudo-only diagnostics (dmidecode for BIOS, system, baseboard, memory details) |
--dry-run / -n | Show what would be done without executing |