Skip to content

Setup Wizard Walkthrough

The setup wizard is the easiest way to get your DGX Spark cluster ready for multi-node inference. It walks you through every step interactively — just answer the prompts and sparkrun handles the rest.

Terminal window
sparkrun setup

Each phase is optional — you can skip any step and come back to it later with the individual sparkrun setup <command> subcommands.

PhaseWhat it doesWhy it matters
0. Install checkDetects whether sparkrun is installed as a uv toolInstalls it if not, so sparkrun is on PATH afterwards
1. Cluster setupSaves your DGX Spark hosts by nameSo you don’t repeat host lists on every command
— SSH accessEstablishes passwordless control→host SSHEvery later phase assumes it, so it is settled first
2. SSH meshSets up passwordless SSH between all hostsRequired for sparkrun to orchestrate containers remotely
3. CX7 networkingConfigures high-speed ConnectX-7 interfacesEnables fast NCCL communication and data transfers between nodes
3b. SSH re-meshRe-runs the mesh when CX7 changed any IPsKeeps every interface reachable, automatically
4. Docker groupAdds your user to the docker groupSo you can run containers without sudo
4b. NVIDIA CDIGenerates /etc/cdi/nvidia.yamlsparkrun requests GPUs via CDI, so containers get no GPU without it
5. Sudoers entriesInstalls scoped sudo rulesAllows automatic cache permission fixes and page cache clearing
6. earlyoomInstalls OOM protectionPrevents system hangs when models approach the 128 GB memory limit

You’ll need:

  • sparkrun installed (Installation)
  • The IP address or hostname of your DGX Spark system(s)
  • SSH access to each Spark (you’ll be prompted for passwords during setup)

The wizard starts by discovering your DGX Sparks and creating a named cluster.

sparkrun detects the local ConnectX-7 interfaces and scans for peer Sparks on the same CX7 subnets:

Detecting CX7 interfaces on this machine...
CX7 detected! This machine is a DGX Spark.
Scanning CX7 subnets for peer Sparks...
Found 1 peer(s) on CX7: 192.168.11.14
Note: These are CX7 addresses. Enter management IPs below
if your hosts use separate management networking.
Enter host IPs/hostnames (comma-separated) [10.24.11.13]: 10.24.11.13,10.24.11.14

The wizard suggests hosts it found, but you should enter management network IPs (the 10 Gbps built-in Ethernet), not CX7 IPs. sparkrun uses management IPs for SSH orchestration and discovers CX7 interfaces automatically.

If you’re running from an external machine

Section titled “If you’re running from an external machine”
Detecting CX7 interfaces on this machine...
No CX7 interfaces detected on this machine.
Enter DGX Spark host IPs/hostnames (comma-separated): 10.24.11.13,10.24.11.14

Enter the management IPs of your DGX Sparks, separated by commas.

Cluster name [default]: mylab
SSH username [drew]: dgxuser
Created cluster 'mylab' with 2 host(s), set as default.

Pick any name you like. The first host in the list becomes the head node for multi-node jobs. The SSH username should be the same user account on all your Sparks.

Immediately after the cluster name and SSH username prompts — and before any other phase touches a host — the wizard makes sure it can actually reach them:

Checking SSH access to 2 host(s)...
10.24.11.13 reachable, key accepted
10.24.11.14 reachable, password auth (no key installed)
Install an SSH key on 10.24.11.14? [Y/n]: y

Every later phase assumes passwordless control→host SSH, so settling it first turns “the wizard failed halfway through” into a single answerable question.

The probe distinguishes three outcomes, and only the first is offered a key:

  • Reachable but rejecting the key — a bootstrap candidate; the wizard offers to install one, generating ~/.ssh/sparkrun_ed25519 if you have no identity yet.
  • Host key changed — never auto-fixed. This can indicate a genuinely different machine, so you resolve it yourself.
  • Unreachable — a network or sshd problem, not something a key installation would solve.

Success is confirmed by re-probing rather than by trusting the install command’s exit code.

Phase 2: SSH Mesh
------------------------------
Set up SSH mesh across 2 host(s) + this machine? [Y/n]: y

This establishes passwordless SSH between all hosts and your control machine. You’ll be prompted for passwords on first connection to each host — after that, everything is passwordless.

The SSH mesh is required for sparkrun to:

  • Launch and stop containers on remote hosts
  • Sync model files and container images between nodes
  • Coordinate multi-node inference

If SSH is already configured between your hosts, this phase detects that and skips the already-working connections.

For more details, see SSH Setup.

Phase 3: CX7 Network Configuration
------------------------------
Configures high-speed CX7 networking between hosts.
Configure CX7 networking? [Y/n]: y
Subnets: 192.168.11.0/24, 192.168.12.0/24

This phase configures the ConnectX-7 high-speed network interfaces on each DGX Spark. It:

  1. Detects CX7 interfaces on each host
  2. Selects subnets — picks two /24 subnets that don’t conflict with your existing networks
  3. Assigns static IPs — derives CX7 IPs from each host’s management IP (e.g., management 10.24.11.13 → CX7 192.168.11.13 and 192.168.12.13)
  4. Sets MTU 9000 (jumbo frames) for maximum throughput
  5. Applies netplan configuration and distributes host keys

If your CX7 interfaces are already configured correctly, this phase skips them automatically.

For a deep dive into DGX Spark networking, see Networking.

When CX7 configuration adds or changes network IPs, the wizard automatically re-runs the SSH mesh to ensure full connectivity across all interfaces (management + CX7). This happens transparently — no extra prompts. If CX7 was skipped or already configured, no re-mesh is needed.

Phase 4: Docker Group Membership
------------------------------
Ensures user can run Docker commands without sudo.
Add 'dgxuser' to the docker group on all hosts? [Y/n]: y

sparkrun runs Docker commands over SSH. Your user needs to be in the docker group on each Spark to run containers without sudo. This phase checks and fixes group membership automatically.

Phase 4b: NVIDIA CDI (Container Device Interface)
------------------------------
Generates /etc/cdi/nvidia.yaml so Docker can access the GPU(s).
Generate the NVIDIA CDI spec on all hosts? [Y/n]: y

sparkrun requests GPUs as --device nvidia.com/gpu=all, which Docker resolves through /etc/cdi/nvidia.yaml. Without that spec a container starts but sees no GPU, so this phase runs nvidia-ctk cdi generate on each host.

The script skips itself where nvidia-ctk is absent, so a host without the NVIDIA Container Toolkit reports as skipped rather than failing the wizard.

Phase 5: Sudoers Entries
------------------------------
Scoped sudoers for fix-permissions + clear-cache (no broad sudo).
Install sudoers entries? [Y/n]: y

This installs two minimal, scoped sudoers rules on each host:

  • fix-permissions — allows sparkrun to fix HuggingFace cache ownership without a password. Docker containers create cache files as root, which can break subsequent runs. See HuggingFace cache owned by root.
  • clear-cache — allows sparkrun to drop the Linux page cache before launching containers, freeing memory for inference. See Linux page cache consuming memory.

These rules grant only the specific operations needed — no broad sudo privileges. You may be prompted for your sudo password once.

Phase 6: earlyoom OOM Protection
------------------------------
Prevents system hangs by proactively managing memory pressure.
Install earlyoom? [Y/n]: y
earlyoom configured on 2/2 host(s).

earlyoom monitors available memory and proactively terminates processes before the system becomes unresponsive. This is especially important on DGX Spark because the 128 GB is shared between CPU and GPU — running a large model near the memory limit without earlyoom can require a hard reboot.

sparkrun configures earlyoom with inference-aware defaults:

  • Prefer killing: inference processes (vllm, sglang, llama-server, trtllm)
  • Avoid killing: system services (sshd, dockerd, systemd)

After all phases complete, the wizard shows a summary:

Setup Complete!
================================================
Cluster: mylab (2 hosts, set as default)
SSH mesh: OK
CX7: configured (192.168.11.0/24, 192.168.12.0/24)
SSH remesh: OK
Docker: OK (2/2)
CDI: OK (2/2)
Sudoers: installed (fix-permissions, clear-cache)
earlyoom: installed
Next steps:
sparkrun list # Browse available recipes
sparkrun show qwen3-1.7b-vllm # Recipe details + VRAM estimate
sparkrun run qwen3-1.7b-vllm --dry-run # Preview a launch
sparkrun run qwen3-1.7b-vllm # Launch inference

Your cluster is ready. Head to the Quick Start to launch your first model.

Every phase asks for confirmation. Press n to skip any step:

Configure CX7 networking? [Y/n]: n

You can configure skipped steps later with the individual commands:

PhaseIndividual command
SSH meshsparkrun setup ssh
CX7 networkingsparkrun setup cx7
Docker groupsparkrun setup docker-group
NVIDIA CDIsudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml (on each host)
Sudoerssparkrun setup fix-permissions --save-sudo and sparkrun setup clear-cache --save-sudo
earlyoomsparkrun setup earlyoom

To see which of these a cluster is still missing, run sparkrun setup check — it probes every host non-destructively and names the command that fixes each gap.

For scripted or CI/CD setups, use --yes to accept all defaults:

Terminal window
sparkrun setup wizard --yes --hosts 10.24.11.13,10.24.11.14 --cluster mylab

Use --dry-run to preview what would happen without making changes:

Terminal window
sparkrun setup wizard --dry-run --hosts 10.24.11.13,10.24.11.14

See Setup Commands for the full reference of all wizard options and individual setup subcommands.