Skip to content

Multi-Node Tensor Parallelism

This tutorial shows how to run models that exceed a single DGX Spark’s 128 GB unified memory by splitting them across multiple nodes with tensor parallelism.

  • 2 or more DGX Sparks with SSH access
  • A cluster created (via sparkrun cluster create or the setup wizard)
  • SSH mesh configured (sparkrun setup ssh)
  • ConnectX-7 networking configured (sparkrun setup cx7) — recommended for fast NCCL communication

If you haven’t done these yet, see Quick Start for the one-time setup steps.

Each DGX Spark has a single GPU with 128 GB of unified memory. Models that exceed this limit — or that benefit from more compute — can be split across multiple Sparks using tensor parallelism (TP).

TP splits model layers across GPUs. With sparkrun, each DGX Spark contributes one GPU, so --tp 2 means 2 hosts, --tp 4 means 4 hosts, and so on.

Use sparkrun show to see whether a model fits on a single node:

Terminal window
sparkrun show qwen3-30b-a3b-vllm

The VRAM estimate tells you the expected memory usage at the default tensor parallelism. To see the estimate at a different TP level, use the -o override:

Terminal window
# Compare VRAM at TP=1 vs TP=2
sparkrun show qwen3-30b-a3b-vllm
sparkrun show qwen3-30b-a3b-vllm -o tensor_parallel=2

If the estimate at TP=1 exceeds 128 GB, you need multi-node.

Run with --tp 2 to spread the model across two Sparks:

Terminal window
sparkrun run qwen3-30b-a3b-vllm --tp 2

This uses your default cluster. You can also specify hosts explicitly:

Terminal window
sparkrun run qwen3-30b-a3b-vllm --tp 2 --hosts 10.24.11.13,10.24.11.14

When you launch a multi-node job, sparkrun:

  1. Detects InfiniBand interfaces on each host for fast inter-node communication
  2. Configures NCCL environment variables to use the high-speed network
  3. Syncs the container image to all hosts (skipped if already present)
  4. Syncs the model to all hosts (uses InfiniBand IPs for fast rsync when available)
  5. Starts containers on all placed hosts — the first host of the placement becomes the head node, the others are workers
  6. Establishes distributed communication between nodes (via Ray, native distributed init, or RPC depending on the runtime)

The head node runs the inference API endpoint. Worker nodes participate in tensor-parallel computation but don’t serve HTTP traffic.

Check that all nodes are running:

Terminal window
sparkrun status

You should see containers on each host — one head and one or more workers — all showing as running.

Watch live resource usage across all nodes:

Terminal window
sparkrun cluster monitor

This opens a TUI (terminal UI) showing CPU, RAM, and GPU utilization on each host. For a simpler non-interactive view:

Terminal window
sparkrun cluster monitor --simple

The API is served on the head node. sparkrun status tells you which host that is:

Terminal window
curl http://<head-node-ip>:8000/v1/models

Usage is identical to single-node — the tensor parallelism is transparent to API clients.

When stopping a multi-node job, match the --tp you used at launch:

Terminal window
sparkrun stop qwen3-30b-a3b-vllm --tp 2
  • Head node selection: The head is the first host of the scheduled placement, which is not necessarily the first host in your cluster definition — see the note below.
  • InfiniBand matters: CX-7 networking provides dramatically faster model syncing and NCCL communication. If multi-node jobs are slow to start or inference latency is high, make sure CX-7 is configured (sparkrun setup cx7).
  • Node selection: If your cluster has more hosts than the requested TP level, sparkrun uses only as many as the job needs. Which ones depends on the scheduler: the default occupancy-sparse picks the least-loaded hosts, while greedy always takes the first N in the definition.
  • Measure performance: Compare single-node vs multi-node throughput with Benchmarking Models
  • Serve multiple models: Use the Proxy Gateway to expose models through a unified API