Networking Best Practices
This guide covers how to configure the high-speed ConnectX-7 network on DGX Spark systems for multi-node clustering. Proper network setup is a prerequisite for performant multi-node inference with NCCL.
DGX Spark networking overview
Section titled “DGX Spark networking overview”Each DGX Spark has two wired network connection types:
| Interface | Speed | Purpose |
|---|---|---|
| Built-in Ethernet (management) | 10 Gbps | SSH, general traffic, management |
| ConnectX-7 (QSFP56 ports) | Up to 200 Gbps | NCCL/RDMA, model sync, container distribution |
The ConnectX-7 NIC has two QSFP56 ports and supports up to 200 Gbps aggregate bandwidth. This can be achieved with a single 200 Gbps cable or two 100 Gbps cables.
The PCIe split
Section titled “The PCIe split”DGX Spark has a unique hardware constraint: the SoC provides a maximum of PCIe 5.0 x4 lanes per device. A single x4 link delivers approximately 100 Gbps of bandwidth — half the NIC’s 200 Gbps capacity.
To work around this, each physical QSFP56 port is split into two PCIe x4 partitions. Each partition appears as a separate network device pair in Linux — one Ethernet interface and one RoCEv2 interface:
# Example output of ibdev2netdev with one cable in the "first" port (p0)rocep1s0f0 port 1 ==> enp1s0f0np0 (Up)rocep1s0f1 port 1 ==> enp1s0f1np1 (Down)roceP2p1s0f0 port 1 ==> enP2p1s0f0np0 (Up)roceP2p1s0f1 port 1 ==> enP2p1s0f1np1 (Down)With a single cable plugged into one port, you get two active Ethernet interfaces (e.g., enp1s0f1np0 and enP2p1s0f1np0) and two active RoCEv2 interfaces (e.g., rocep1s0f0 and roceP2p1s0f0). Each pair represents one PCIe x4 partition capable of ~100 Gbps.
Using two cables (one per physical port) also results in two devices per port — four total — but does not meaningfully increase point-to-point bandwidth between two Sparks. A single cable is sufficient for full 200 Gbps between two nodes.
Network design principles
Section titled “Network design principles”Separate management and cluster traffic
Section titled “Separate management and cluster traffic”The single most important networking decision: keep cluster (NCCL/RDMA) traffic off the management network, and keep management traffic off the CX-7 interfaces.
- Management network (10 Gbps built-in Ethernet): SSH, system administration, general internet access, package updates.
- Cluster network (CX-7, up to 200 Gbps): NCCL communication, model file sync (rsync), container image distribution. Nothing else.
Minimize non-cluster activity on the CX-7 subnets. Stray traffic competes with NCCL and degrades inference performance.
sparkrun follows this principle by default — SSH orchestration and control-plane communication go over the management interface, while model sync, container distribution, and NCCL all use the high-speed CX-7 interfaces. In fact, it would be ideal to NOT do container and model synchronization over CX-7 interfaces; however, it’s realistically the most practical choice to do so.
Always use static IPs
Section titled “Always use static IPs”Use static IP addresses on every CX-7 interface. Do not use DHCP or auto-assigned addresses.
Choosing your subnets
Section titled “Choosing your subnets”You need two private subnets for the CX-7 cluster interfaces — one per PCIe partition. These are isolated point-to-point or switched links with no internet routing, so you must use addresses from the RFC 1918 private ranges:
| Range | CIDR | Addresses | Common name |
|---|---|---|---|
10.0.0.0 – 10.255.255.255 | 10.0.0.0/8 | ~16.7 million | Class A private |
172.16.0.0 – 172.31.255.255 | 172.16.0.0/12 | ~1 million | Class B private |
192.168.0.0 – 192.168.255.255 | 192.168.0.0/16 | ~65,000 | Class C private |
Any of these ranges work. Use a /24 subnet mask (254 usable addresses) — it is the simplest to reason about and more than enough for any Spark cluster. Even an 8-node cluster only needs 8 addresses per subnet.
The key rule: your cluster subnets must not overlap with any network your Sparks are already connected to. Check for conflicts with:
- Your management network / LAN (e.g., if your Sparks get management IPs via DHCP on
192.168.1.0/24, don’t use that range) - VPN address ranges
- Docker’s default bridge networks (
172.17.0.0/16and similar) - Any other subnets routed on the machine
A simple approach: pick two adjacent /24 subnets from a range you know is unused. Throughout this guide we use 192.168.11.0/24 and 192.168.12.0/24 as examples — replace these with whatever is free in your environment.
Assign IPs to both partitions — on different subnets
Section titled “Assign IPs to both partitions — on different subnets”Each physical port exposes two Ethernet interfaces (the “twin” pair from the PCIe split). Assign an IP address to both interfaces, but they must be on different subnets:
| Interface | Subnet | Example IP |
|---|---|---|
enp1s0f1np0 (partition 1) | 192.168.11.0/24 | 192.168.11.13 |
enP2p1s0f1np0 (partition 2) | 192.168.12.0/24 | 192.168.12.13 |
Assigning IPs to both partitions allows tools like sparkrun to potentially be able to use both paths for data transfers (model sync, container distribution), maximizing the available 200 Gbps. NCCL uses RoCEv2 directly and will utilize both RoCE interfaces regardless, but having IPs on both enables addressing both interfaces via TCP/IP interfaces.
Jumbo frames (MTU 9000)
Section titled “Jumbo frames (MTU 9000)”MTU is maximum transmission unit. It sets the upper bound on packet size. The default & standard MTU is 1500. Jumbo Frames are an “industry-standard” but technically not part of proper standards (which is why your Internet connection won’t allow jumbo frames).
Set MTU to 9000 on all CX-7 interfaces that are participating in data transfer activity. Jumbo frames reduce per-packet overhead and measurably improve throughput for large data transfers.
# /etc/netplan/40-cx7.yaml — set MTU on both twin interfacesnetwork: version: 2 ethernets: enp1s0f1np0: dhcp4: no mtu: 9000 addresses: [192.168.11.13/24] enP2p1s0f1np0: dhcp4: no mtu: 9000 addresses: [192.168.12.13/24]If your cluster uses a switch, the switch ports should be configured for MTU 9216 to leave room MTU 9000 + overhead.
Setup sequence
Section titled “Setup sequence”The end-to-end setup involves three sparkrun commands. Register the cluster first so subsequent commands can reference it by name:
| Step | Command | Notes |
|---|---|---|
| 1. Register cluster | sparkrun cluster create mylab --hosts <mgmt-IPs> | Save hosts by name using management IPs |
| 2. SSH access | sparkrun setup ssh | Interactive — prompts for passwords; uses default cluster |
| 3. CX-7 networking | sparkrun setup cx7 | Requires passwordless SSH from step 2. Will also prompt for passwords as needed |
sparkrun setup ssh is interactive — it prompts for passwords and doesn’t require pre-existing passwordless SSH. sparkrun setup cx7 runs remote scripts non-interactively, so it does require passwordless SSH to already be in place. That’s why SSH setup comes first.
Your cluster stays registered with the management IPs — there is no need to update it after CX-7 setup. sparkrun automatically discovers CX-7 interfaces on each host and uses them for high-bandwidth operations (model sync, container distribution, NCCL), while SSH orchestration continues over the management network.
# Step 1: Register cluster with management IPssparkrun cluster create mylab \ --hosts 10.24.11.13,10.24.11.14 \ --user dgxuser \ -d "2-node DGX Spark cluster"sparkrun cluster set-default mylab
# Step 2: SSH setup (prompts for passwords, uses default cluster)sparkrun setup ssh
# Step 3: Configure CX-7 interfaces (uses default cluster)sparkrun setup cx7See SSH Setup for more details on SSH configuration options.
Configuring CX-7 interfaces with sparkrun
Section titled “Configuring CX-7 interfaces with sparkrun”sparkrun setup cx7 automates the CX-7 configuration — it SSHs into each host, detects ConnectX-7 interfaces, selects two conflict-free subnets, assigns static IPs with jumbo frames, and applies the netplan configuration. This is the recommended approach.
Basic usage
Section titled “Basic usage”# Specify hosts directly (management IPs)sparkrun setup cx7 --hosts 10.24.11.13,10.24.11.14
# Or use a saved clustersparkrun setup cx7 --cluster mylab
# Preview what would be done without making changessparkrun setup cx7 --cluster mylab --dry-runWhat it does
Section titled “What it does”- Detects CX-7 interfaces on each host via SSH (runs
ibdev2netdev, identifies twin pairs) - Selects subnets — automatically picks two
/24subnets from RFC 1918 space that don’t conflict with any existing networks on the hosts. If hosts already have a valid CX-7 configuration, it preserves those subnets. - Assigns IPs — derives the last octet of each CX-7 IP from the host’s management IP (e.g., management IP
10.24.11.13→ CX-7 IPs192.168.11.13and192.168.12.13), making the addressing scheme easy to remember. - Applies netplan — generates and writes
/etc/netplan/40-cx7.yamlon each host with static IPs, MTU 9000, and runsnetplan apply. Requires sudo on the target hosts.
Existing valid configurations are preserved — if a host already has correct CX-7 netplan config with the right subnets and MTU, it is skipped. Use --force to reconfigure all hosts regardless.
Specifying subnets
Section titled “Specifying subnets”If you prefer specific subnets rather than automatic selection:
sparkrun setup cx7 --cluster mylab \ --subnet1 192.168.11.0/24 \ --subnet2 192.168.12.0/24Both --subnet1 and --subnet2 must be provided together. See Choosing your subnets above for guidance on picking ranges that won’t conflict.
Options reference
Section titled “Options reference”| Option | Description |
|---|---|
--hosts / -H | Comma-separated host list (management IPs or hostnames) |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster definition |
--user / -u | SSH username (default: from cluster config or current user) |
--subnet1 | Override subnet for CX-7 partition 1 (e.g., 192.168.11.0/24) |
--subnet2 | Override subnet for CX-7 partition 2 (e.g., 192.168.12.0/24) |
--mtu | MTU for CX-7 interfaces (default: 9000) |
--dry-run / -n | Show the plan without making changes |
--force | Reconfigure even if existing config is valid |
Verify connectivity
Section titled “Verify connectivity”After sparkrun setup cx7 completes, verify the links are working:
# From spark1 — test both subnetsping -c 3 192.168.11.14ping -c 3 192.168.12.14Verify jumbo frames are effective end-to-end:
ping -M do -s 8972 -c 3 192.168.11.14A successful large-packet ping confirms MTU 9000 is working. If it fails, check MTU on both endpoints and any switch in between.
Manual configuration
Section titled “Manual configuration”If you prefer to configure netplan by hand (or need to troubleshoot), the process is:
- Identify your interfaces with
ibdev2netdev— the “Up” interfaces correspond to your connected port - Create
/etc/netplan/40-cx7.yamlon each node with static IPs on both twin interfaces, different subnets, and MTU 9000:
network: version: 2 ethernets: enp1s0f0np0: dhcp4: no mtu: 9000 addresses: [192.168.11.13/24] enP2p1s0f0np0: dhcp4: no mtu: 9000 addresses: [192.168.12.13/24]- Apply on each node:
sudo chmod 600 /etc/netplan/40-cx7.yamlsudo netplan applyHost preparation checklist
Section titled “Host preparation checklist”Beyond networking, each DGX Spark in the cluster needs some basic setup.
Stay on DGX OS
Section titled “Stay on DGX OS”We recommend sticking with the DGX OS that ships with the Spark rather than switching to another Linux distribution. DGX OS includes NVIDIA drivers, CUDA, container runtime, and firmware that are tested together. Switching distributions risks driver/firmware compatibility issues.
Same username across all nodes
Section titled “Same username across all nodes”Use the same username on every Spark in the cluster. sparkrun connects via SSH and expects a consistent user across hosts.
Ideally, the cluster user would be a dedicated service account rather than your personal login. In practice, most Spark deployments have a single user, and that works fine.
Required packages
Section titled “Required packages”These should already be installed on DGX OS, but verify:
# Confirm these are availablewhich rsync git dockerIf anything is missing:
sudo apt-get update && sudo apt-get install -y rsync gitVerifying RDMA performance
Section titled “Verifying RDMA performance”Once networking is configured, verify that the fabric actually carries traffic:
sparkrun setup features enable cli.setup.rdma_test # experimental; enable oncesparkrun setup rdma-test --cluster myclusterThe command derives host pairs from the CX-7 subnets it configured, so it only tests links that physically exist. For each link it measures latency and bandwidth, then drives every link between a host pair at once — a DGX Spark QSFP112 cable presents as two RDMA devices, and testing them one at a time reports half the cable’s real throughput.
h1 <-> h2 (2 links) [OK] h1:rocep1s0f0 <-> h2:rocep1s0f0 (192.168.11.0/24) — lat 3.02 us, bw 111.7 Gb/s, of 100 Gb/s [OK] h1:roceP2p1s0f0 <-> h2:roceP2p1s0f0 (192.168.12.0/24) — lat 3.32 us, bw 111.7 Gb/s, of 100 Gb/s aggregate (all links concurrently) — 195.7 Gb/s of 200 Gb/s
Results: 2 OK.Expect roughly 100 Gbps per RDMA device and close to 200 Gbps aggregate across the cable.
Note what the per-device expectation is measured against. On DGX Spark the NIC exposes two PCIe
functions on a single physical port, so both devices share one 200 Gbps wire; each is compared
against its fair share of that port, not against the port’s whole rate. sparkrun detects this from
the hardware — devices reporting the same sys_image_guid and phys_port_name share a wire — so
a conventional dual-port card with two cables is still treated as additive.
Underperformance is reported as a warning; only a test that could not run at all fails the command, so it is safe to gate a script on the exit status.
Adding the NCCL collective
Section titled “Adding the NCCL collective”--suite all additionally runs an nccl-tests collective, which is the number that actually predicts
multi-node inference performance. It fetches the test image onto every host and takes several
minutes, which is why it is not the default:
sparkrun setup rdma-test --cluster mycluster --suite all [OK] NCCL all_gather_perf (16G) across 2 host(s) — avg bus bandwidth 23.53 GB/sThis needs passwordless host-to-host SSH across the cluster (sparkrun setup ssh-mesh), because
mpirun reaches each peer through the host’s own sshd before stepping into its container.
Useful options
Section titled “Useful options”| Option | Effect |
|---|---|
--suite all | Add the NCCL collective (pulls the image; takes minutes) |
--suite nccl | The collective only |
-D 30 | Longer bandwidth runs (default 10 seconds) |
--size 4G | Smaller NCCL message size |
--json | Machine-readable results |
--dry-run | Show which links would be tested, without sending traffic |
The perftest suite runs ib_write_bw / ib_write_lat straight from the host, which DGX OS ships
preinstalled. The NCCL suite needs a container image
(ghcr.io/spark-arena/sparkrun-rdma-test) because it bundles NCCL and nccl-tests, which otherwise
take about ten minutes to build on every node. Point --image at a mirror if your nodes cannot
reach GHCR.
Running the tests by hand
Section titled “Running the tests by hand”The same measurements, without sparkrun. On the receiver node:
ib_write_bw -d rocep1s0f1 --report_gbits -q 4 -R --force-link IBOn the sender node:
ib_write_bw 192.168.11.14 -d rocep1s0f1 --report_gbits -q 4 -R --force-link IBAnd for latency, on the receiver then the sender:
ib_write_lat -d rocep1s0f1 --report_gbits -R --force-link IBib_write_lat 192.168.11.14 -d rocep1s0f1 --report_gbits -R --force-link IBYou should see approximately 100 Gbps per RoCE interface and 1.5 microseconds of latency.
Note that a single ib_write_bw run exercises one of the two RDMA devices, so it reports about half
the cable’s capacity; with both twins active via NCCL, aggregate bandwidth approaches 200 Gbps.
Resource transfer modes
Section titled “Resource transfer modes”When launching multi-node inference, sparkrun needs to distribute container images and model files to all hosts. The transfer mode controls how resources flow from the control machine (where you run sparkrun) to the cluster nodes.
The four modes
Section titled “The four modes”auto (default)
Section titled “auto (default)”Automatically selects the best mode based on network topology. If the control machine can reach the cluster’s CX-7 IPs, it uses local. Otherwise, it falls back to push.
The control machine distributes directly to all hosts over the high-speed CX-7 network. This is the fastest option when the control machine is on the same CX-7 subnet as the cluster (e.g., the control machine is itself a DGX Spark in the cluster).
┌─────────────┐│ Control ││ Machine │ ◀───▶ Internet│ (on CX-7) │ (registry/HF)└──────┬──────┘ │ CX-7 (up to 200 Gbps) ├────────────────┐ ▼ ▼┌─────────────┐ ┌─────────────┐│ Head │ │ Worker ││ (Spark 1) │ │ (Spark 2) │└─────────────┘ └─────────────┘The control machine pushes resources to the head node over the management network, then the head node distributes to workers over CX-7. This is the right choice when the control machine is external (e.g., a laptop or desktop not on the CX-7 network).
┌─────────────┐│ Control ││ Machine │ ◀───▶ Internet│ (external) │ (registry/HF)└──────┬──────┘ │ Management (up to 10 Gbps) ▼┌─────────────┐│ Head ││ (Spark 1) │└──────┬──────┘ │ CX-7 (up to 200 Gbps) ▼┌─────────────┐│ Worker ││ (Spark 2) │└─────────────┘delegated
Section titled “delegated”The head node downloads resources directly (e.g., pulls the container image from a registry, downloads the model from HuggingFace) and distributes to workers over CX-7. The control machine transfers nothing — it only orchestrates via SSH.
┌─────────────┐│ Control ││ Machine │─── SSH only (orchestration) ───┐│ (external) │ │└─────────────┘ │ ▼ ┌─────────────┐ ┌─────────────┐ Internet <────> │ Head │ ──────> │ Worker │ (registry/HF) │ (Spark 1) │ CX-7 │ (Spark 2) │ └─────────────┘ └─────────────┘When to use each mode
Section titled “When to use each mode”| Scenario | Recommended mode | Why |
|---|---|---|
| Control machine is a cluster Spark | local or auto | Direct CX-7 transfers to all nodes |
| Control machine is external (laptop, desktop) | push or auto | Avoids slow management-network fan-out |
| CI/CD, External Control Node, etc. | delegated | Head downloads at full speed, skips slow uplink |
| First-time setup, unsure of topology | auto | Auto-detects CX-7 reachability |
Resource handling by mode
Section titled “Resource handling by mode”| Step | local | push | delegated |
|---|---|---|---|
| Image pull/build | Control machine | Control machine | Head node |
| Image → head | Control → head (CX-7) | Control → head (mgmt) | Already on head |
| Image → workers | Control → workers (CX-7) | Head → workers (CX-7) | Head → workers (CX-7) |
| Model download | Control machine | Control machine | Head node |
| Model → head | Control → head (CX-7) | Control → head (mgmt) | Already on head |
| Model → workers | Control → workers (CX-7) | Head → workers (CX-7) | Head → workers (CX-7) |
In all modes, sparkrun checks whether the image and model already exist on each target host before transferring. If a host already has the correct image ID or model files, that host is skipped. This means repeat launches are fast regardless of transfer mode.
Setting the transfer mode
Section titled “Setting the transfer mode”Transfer mode can be set per-cluster (persistent) or per-command (one-off):
# Set on the cluster definition (persists across all runs)sparkrun cluster create mylab \ --hosts 10.24.11.13,10.24.11.14 \ --transfer-mode push
# Or update an existing clustersparkrun cluster update mylab --transfer-mode push
# Override for a single runsparkrun run qwen3-1.7b-vllm --cluster mylab --transfer-mode delegatedCLI --transfer-mode overrides the cluster setting for that invocation.
Quick reference
Section titled “Quick reference”| Item | Recommendation |
|---|---|
| Cable type | QSFP56, 200 Gbps (single cable per node pair is sufficient) |
| IP addressing | Static IPs on both CX-7 partitions, different subnets |
| MTU | 9000 on all CX-7 interfaces (and switch ports) |
| Management traffic | Route over 10 Gbps built-in Ethernet, not CX-7 |
| Inference traffic | Route over 10 Gbps built-in Ethernet, not CX-7 |
| Cluster traffic | NCCL, model sync, container distribution over CX-7 only |
| OS | DGX OS (do not switch distributions) |
| User account | Same username on every node, member of docker group |
| Required tools | rsync, git, docker (pre-installed on DGX OS) |
| Switch (3+ nodes) | Required to maximize bandwidth |