Multi-Node Tensor Parallelism
This tutorial shows how to run models that exceed a single DGX Spark’s 128 GB unified memory by splitting them across multiple nodes with tensor parallelism.
Prerequisites
Section titled “Prerequisites”- 2 or more DGX Sparks with SSH access
- A cluster created (via
sparkrun cluster createor the setup wizard) - SSH mesh configured (
sparkrun setup ssh) - ConnectX-7 networking configured (
sparkrun setup cx7) — recommended for fast NCCL communication
If you haven’t done these yet, see Quick Start for the one-time setup steps.
Why multi-node?
Section titled “Why multi-node?”Each DGX Spark has a single GPU with 128 GB of unified memory. Models that exceed this limit — or that benefit from more compute — can be split across multiple Sparks using tensor parallelism (TP).
TP splits model layers across GPUs. With sparkrun, each DGX Spark contributes one GPU, so --tp 2 means 2 hosts, --tp 4 means 4 hosts, and so on.
1. Check VRAM requirements
Section titled “1. Check VRAM requirements”Use sparkrun show to see whether a model fits on a single node:
sparkrun show qwen3-30b-a3b-vllmThe VRAM estimate tells you the expected memory usage at the default tensor parallelism. To see the estimate at a different TP level, use the -o override:
# Compare VRAM at TP=1 vs TP=2sparkrun show qwen3-30b-a3b-vllmsparkrun show qwen3-30b-a3b-vllm -o tensor_parallel=2If the estimate at TP=1 exceeds 128 GB, you need multi-node.
2. Launch with tensor parallelism
Section titled “2. Launch with tensor parallelism”Run with --tp 2 to spread the model across two Sparks:
sparkrun run qwen3-30b-a3b-vllm --tp 2This uses your default cluster. You can also specify hosts explicitly:
sparkrun run qwen3-30b-a3b-vllm --tp 2 --hosts 10.24.11.13,10.24.11.143. What happens behind the scenes
Section titled “3. What happens behind the scenes”When you launch a multi-node job, sparkrun:
- Detects InfiniBand interfaces on each host for fast inter-node communication
- Configures NCCL environment variables to use the high-speed network
- Syncs the container image to all hosts (skipped if already present)
- Syncs the model to all hosts (uses InfiniBand IPs for fast rsync when available)
- Starts containers on all placed hosts — the first host of the placement becomes the head node, the others are workers
- Establishes distributed communication between nodes (via Ray, native distributed init, or RPC depending on the runtime)
The head node runs the inference API endpoint. Worker nodes participate in tensor-parallel computation but don’t serve HTTP traffic.
4. Verify the deployment
Section titled “4. Verify the deployment”Check that all nodes are running:
sparkrun statusYou should see containers on each host — one head and one or more workers — all showing as running.
5. Monitor your cluster
Section titled “5. Monitor your cluster”Watch live resource usage across all nodes:
sparkrun cluster monitorThis opens a TUI (terminal UI) showing CPU, RAM, and GPU utilization on each host. For a simpler non-interactive view:
sparkrun cluster monitor --simple6. Test the endpoint
Section titled “6. Test the endpoint”The API is served on the head node. sparkrun status tells you which host that is:
curl http://<head-node-ip>:8000/v1/modelsUsage is identical to single-node — the tensor parallelism is transparent to API clients.
7. Stop the workload
Section titled “7. Stop the workload”When stopping a multi-node job, match the --tp you used at launch:
sparkrun stop qwen3-30b-a3b-vllm --tp 2- Head node selection: The head is the first host of the scheduled placement, which is not necessarily the first host in your cluster definition — see the note below.
- InfiniBand matters: CX-7 networking provides dramatically faster model syncing and NCCL communication. If multi-node jobs are slow to start or inference latency is high, make sure CX-7 is configured (
sparkrun setup cx7). - Node selection: If your cluster has more hosts than the requested TP level, sparkrun uses only as many as the job needs. Which ones depends on the scheduler: the default
occupancy-sparsepicks the least-loaded hosts, whilegreedyalways takes the first N in the definition.
Next steps
Section titled “Next steps”- Measure performance: Compare single-node vs multi-node throughput with Benchmarking Models
- Serve multiple models: Use the Proxy Gateway to expose models through a unified API