Skip to content

SSH Setup

All multi-node orchestration in sparkrun relies on SSH. At minimum, you need passwordless SSH from your control machine to every cluster node.

sparkrun pulls container images and models locally and pushes them to each node directly, so node-to-node SSH is not strictly required for the default workflow. However, setting up a full SSH mesh (every host can reach every other host) is recommended.

The easiest approach is sparkrun setup ssh, which creates a full mesh across your cluster hosts and the control machine (included automatically):

Terminal window
# Set up SSH mesh across your cluster + this machine
sparkrun setup ssh --hosts 192.168.11.13,192.168.11.14 --user ubuntu
# Or use a saved cluster
sparkrun setup ssh --cluster mylab
# Or use your default cluster
sparkrun setup ssh
# Add extra hosts beyond the cluster
sparkrun setup ssh --cluster mylab --extra-hosts 10.0.0.99
# Exclude the control machine from the mesh
sparkrun setup ssh --cluster mylab --no-include-self
# Skip IP discovery phase (InfiniBand, CX7)
sparkrun setup ssh --cluster mylab --no-discover-ips

You will be prompted for passwords on first connection to each host. After that, every host in the mesh can SSH to every other host without passwords.

After installing keys, sparkrun verifies that pubkey authentication actually works by testing each host independently (without reusing the password-authenticated connection). If verification fails, it prints diagnostics and actionable remediation steps but continues with the inter-node mesh.

If you prefer to set up SSH yourself:

Terminal window
# Generate a key if you don't have one
ssh-keygen -t ed25519
# Copy to each node
ssh-copy-id 192.168.11.13
ssh-copy-id 192.168.11.14

By default, sparkrun uses your current OS user for SSH. You can configure the user per-cluster:

Terminal window
# Set during cluster creation
sparkrun cluster create mylab --hosts 192.168.11.13,192.168.11.14 --user dgxuser
# Update an existing cluster
sparkrun cluster update mylab --user dgxuser

Or override per-command with --user.

Cross-user SSH (control machine ≠ cluster user)

Section titled “Cross-user SSH (control machine ≠ cluster user)”

A common setup is running sparkrun from a control machine where your local user (e.g., me) differs from the cluster SSH user (e.g., dgxuser). sparkrun handles this automatically — the mesh script installs your local user’s public key into the remote user’s authorized_keys.

If cross-user SSH fails after the mesh completes, the most common cause is home directory permissions. OpenSSH’s StrictModes (enabled by default) rejects keys when the remote user’s home directory is group- or world-writable.

sparkrun automatically fixes this during setup (chmod go-w ~), but if you encounter issues, use --diagnose to pinpoint the problem:

Terminal window
sparkrun setup ssh --cluster mylab --diagnose

This runs comprehensive diagnostics on each host:

  • Tests pubkey authentication independently (without password)
  • Checks home directory, .ssh/, and authorized_keys permissions
  • Inspects sshd_config for AuthorizedKeysFile, StrictModes, AllowUsers/AllowGroups
  • Verifies your local public key is present in the remote authorized_keys
  • Prints pass/fail results with specific remediation steps

For non-default ports, identity files, or jump hosts, use ~/.ssh/config:

Host spark1
HostName 192.168.11.13
User dgxuser
Host spark2
HostName 192.168.11.14
User dgxuser

The SSH user must be a member of the docker group on every cluster node:

Terminal window
sudo usermod -aG docker "$USER"