Setup Wizard Walkthrough
The setup wizard is the easiest way to get your DGX Spark cluster ready for multi-node inference. It walks you through every step interactively — just answer the prompts and sparkrun handles the rest.
sparkrun setupWhat the wizard does
Section titled “What the wizard does”Each phase is optional — you can skip any step and come back to it later with the individual sparkrun setup <command> subcommands.
| Phase | What it does | Why it matters |
|---|---|---|
| 0. Install check | Detects whether sparkrun is installed as a uv tool | Installs it if not, so sparkrun is on PATH afterwards |
| 1. Cluster setup | Saves your DGX Spark hosts by name | So you don’t repeat host lists on every command |
| — SSH access | Establishes passwordless control→host SSH | Every later phase assumes it, so it is settled first |
| 2. SSH mesh | Sets up passwordless SSH between all hosts | Required for sparkrun to orchestrate containers remotely |
| 3. CX7 networking | Configures high-speed ConnectX-7 interfaces | Enables fast NCCL communication and data transfers between nodes |
| 3b. SSH re-mesh | Re-runs the mesh when CX7 changed any IPs | Keeps every interface reachable, automatically |
| 4. Docker group | Adds your user to the docker group | So you can run containers without sudo |
| 4b. NVIDIA CDI | Generates /etc/cdi/nvidia.yaml | sparkrun requests GPUs via CDI, so containers get no GPU without it |
| 5. Sudoers entries | Installs scoped sudo rules | Allows automatic cache permission fixes and page cache clearing |
| 6. earlyoom | Installs OOM protection | Prevents system hangs when models approach the 128 GB memory limit |
Before you begin
Section titled “Before you begin”You’ll need:
- sparkrun installed (Installation)
- The IP address or hostname of your DGX Spark system(s)
- SSH access to each Spark (you’ll be prompted for passwords during setup)
Phase 1: Cluster setup
Section titled “Phase 1: Cluster setup”The wizard starts by discovering your DGX Sparks and creating a named cluster.
If you’re running on a DGX Spark
Section titled “If you’re running on a DGX Spark”sparkrun detects the local ConnectX-7 interfaces and scans for peer Sparks on the same CX7 subnets:
Detecting CX7 interfaces on this machine... CX7 detected! This machine is a DGX Spark. Scanning CX7 subnets for peer Sparks... Found 1 peer(s) on CX7: 192.168.11.14
Note: These are CX7 addresses. Enter management IPs below if your hosts use separate management networking.
Enter host IPs/hostnames (comma-separated) [10.24.11.13]: 10.24.11.13,10.24.11.14The wizard suggests hosts it found, but you should enter management network IPs (the 10 Gbps built-in Ethernet), not CX7 IPs. sparkrun uses management IPs for SSH orchestration and discovers CX7 interfaces automatically.
If you’re running from an external machine
Section titled “If you’re running from an external machine”Detecting CX7 interfaces on this machine... No CX7 interfaces detected on this machine.
Enter DGX Spark host IPs/hostnames (comma-separated): 10.24.11.13,10.24.11.14Enter the management IPs of your DGX Sparks, separated by commas.
Naming your cluster
Section titled “Naming your cluster”Cluster name [default]: mylabSSH username [drew]: dgxuserCreated cluster 'mylab' with 2 host(s), set as default.Pick any name you like. The first host in the list becomes the head node for multi-node jobs. The SSH username should be the same user account on all your Sparks.
SSH access
Section titled “SSH access”Immediately after the cluster name and SSH username prompts — and before any other phase touches a host — the wizard makes sure it can actually reach them:
Checking SSH access to 2 host(s)... 10.24.11.13 reachable, key accepted 10.24.11.14 reachable, password auth (no key installed)
Install an SSH key on 10.24.11.14? [Y/n]: yEvery later phase assumes passwordless control→host SSH, so settling it first turns “the wizard failed halfway through” into a single answerable question.
The probe distinguishes three outcomes, and only the first is offered a key:
- Reachable but rejecting the key — a bootstrap candidate; the wizard offers to install one, generating
~/.ssh/sparkrun_ed25519if you have no identity yet. - Host key changed — never auto-fixed. This can indicate a genuinely different machine, so you resolve it yourself.
- Unreachable — a network or
sshdproblem, not something a key installation would solve.
Success is confirmed by re-probing rather than by trusting the install command’s exit code.
Phase 2: SSH mesh
Section titled “Phase 2: SSH mesh”Phase 2: SSH Mesh------------------------------Set up SSH mesh across 2 host(s) + this machine? [Y/n]: yThis establishes passwordless SSH between all hosts and your control machine. You’ll be prompted for passwords on first connection to each host — after that, everything is passwordless.
The SSH mesh is required for sparkrun to:
- Launch and stop containers on remote hosts
- Sync model files and container images between nodes
- Coordinate multi-node inference
If SSH is already configured between your hosts, this phase detects that and skips the already-working connections.
For more details, see SSH Setup.
Phase 3: CX7 network configuration
Section titled “Phase 3: CX7 network configuration”Phase 3: CX7 Network Configuration------------------------------Configures high-speed CX7 networking between hosts.Configure CX7 networking? [Y/n]: y Subnets: 192.168.11.0/24, 192.168.12.0/24This phase configures the ConnectX-7 high-speed network interfaces on each DGX Spark. It:
- Detects CX7 interfaces on each host
- Selects subnets — picks two
/24subnets that don’t conflict with your existing networks - Assigns static IPs — derives CX7 IPs from each host’s management IP (e.g., management
10.24.11.13→ CX7192.168.11.13and192.168.12.13) - Sets MTU 9000 (jumbo frames) for maximum throughput
- Applies netplan configuration and distributes host keys
If your CX7 interfaces are already configured correctly, this phase skips them automatically.
For a deep dive into DGX Spark networking, see Networking.
Automatic SSH re-mesh after CX7 changes
Section titled “Automatic SSH re-mesh after CX7 changes”When CX7 configuration adds or changes network IPs, the wizard automatically re-runs the SSH mesh to ensure full connectivity across all interfaces (management + CX7). This happens transparently — no extra prompts. If CX7 was skipped or already configured, no re-mesh is needed.
Phase 4: Docker group membership
Section titled “Phase 4: Docker group membership”Phase 4: Docker Group Membership------------------------------Ensures user can run Docker commands without sudo.Add 'dgxuser' to the docker group on all hosts? [Y/n]: ysparkrun runs Docker commands over SSH. Your user needs to be in the docker group on each Spark to run containers without sudo. This phase checks and fixes group membership automatically.
Phase 4b: NVIDIA CDI
Section titled “Phase 4b: NVIDIA CDI”Phase 4b: NVIDIA CDI (Container Device Interface)------------------------------Generates /etc/cdi/nvidia.yaml so Docker can access the GPU(s).Generate the NVIDIA CDI spec on all hosts? [Y/n]: ysparkrun requests GPUs as --device nvidia.com/gpu=all, which Docker resolves
through /etc/cdi/nvidia.yaml. Without that spec a container starts but sees
no GPU, so this phase runs nvidia-ctk cdi generate on each host.
The script skips itself where nvidia-ctk is absent, so a host without the
NVIDIA Container Toolkit reports as skipped rather than failing the wizard.
Phase 5: Sudoers entries
Section titled “Phase 5: Sudoers entries”Phase 5: Sudoers Entries------------------------------Scoped sudoers for fix-permissions + clear-cache (no broad sudo).Install sudoers entries? [Y/n]: yThis installs two minimal, scoped sudoers rules on each host:
- fix-permissions — allows sparkrun to fix HuggingFace cache ownership without a password. Docker containers create cache files as root, which can break subsequent runs. See HuggingFace cache owned by root.
- clear-cache — allows sparkrun to drop the Linux page cache before launching containers, freeing memory for inference. See Linux page cache consuming memory.
These rules grant only the specific operations needed — no broad sudo privileges. You may be prompted for your sudo password once.
Phase 6: earlyoom
Section titled “Phase 6: earlyoom”Phase 6: earlyoom OOM Protection------------------------------Prevents system hangs by proactively managing memory pressure.Install earlyoom? [Y/n]: y earlyoom configured on 2/2 host(s).earlyoom monitors available memory and proactively terminates processes before the system becomes unresponsive. This is especially important on DGX Spark because the 128 GB is shared between CPU and GPU — running a large model near the memory limit without earlyoom can require a hard reboot.
sparkrun configures earlyoom with inference-aware defaults:
- Prefer killing: inference processes (vllm, sglang, llama-server, trtllm)
- Avoid killing: system services (sshd, dockerd, systemd)
Summary
Section titled “Summary”After all phases complete, the wizard shows a summary:
Setup Complete!================================================
Cluster: mylab (2 hosts, set as default) SSH mesh: OK CX7: configured (192.168.11.0/24, 192.168.12.0/24) SSH remesh: OK Docker: OK (2/2) CDI: OK (2/2) Sudoers: installed (fix-permissions, clear-cache) earlyoom: installed
Next steps: sparkrun list # Browse available recipes sparkrun show qwen3-1.7b-vllm # Recipe details + VRAM estimate sparkrun run qwen3-1.7b-vllm --dry-run # Preview a launch sparkrun run qwen3-1.7b-vllm # Launch inferenceYour cluster is ready. Head to the Quick Start to launch your first model.
Skipping phases
Section titled “Skipping phases”Every phase asks for confirmation. Press n to skip any step:
Configure CX7 networking? [Y/n]: nYou can configure skipped steps later with the individual commands:
| Phase | Individual command |
|---|---|
| SSH mesh | sparkrun setup ssh |
| CX7 networking | sparkrun setup cx7 |
| Docker group | sparkrun setup docker-group |
| NVIDIA CDI | sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml (on each host) |
| Sudoers | sparkrun setup fix-permissions --save-sudo and sparkrun setup clear-cache --save-sudo |
| earlyoom | sparkrun setup earlyoom |
To see which of these a cluster is still missing, run sparkrun setup check —
it probes every host non-destructively and names the command that fixes each
gap.
Non-interactive mode
Section titled “Non-interactive mode”For scripted or CI/CD setups, use --yes to accept all defaults:
sparkrun setup wizard --yes --hosts 10.24.11.13,10.24.11.14 --cluster mylabUse --dry-run to preview what would happen without making changes:
sparkrun setup wizard --dry-run --hosts 10.24.11.13,10.24.11.14See Setup Commands for the full reference of all wizard options and individual setup subcommands.