Tips & Troubleshooting
HuggingFace cache owned by root
Section titled “HuggingFace cache owned by root”DGX Spark inference containers typically run as root inside rootful Docker. If a container downloads a model from HuggingFace — either intentionally or because a file wasn’t pre-cached (e.g., a tokenizer config) — the downloaded files in the HuggingFace cache end up owned by root. This causes permission errors on subsequent runs when your non-root user can’t read or update the cache.
sparkrun avoids this by downloading models as your user on the host and syncing them to cluster nodes before launching containers. But if models were previously downloaded inside a container (or you ran a container manually that pulled a model), you may already have root-owned files.
Symptoms:
Permission deniedorOSError: [Errno 13]errors referencing paths under your HuggingFace cache (typically~/.cache/huggingface)- A model that worked before suddenly fails on launch
- sparkrun’s model sync step fails with permission errors
Fix:
sparkrun has a built-in command that fixes ownership across all cluster hosts in one shot:
# Fix permissions on your default clustersparkrun setup fix-permissions
# Or target specific hostssparkrun setup fix-permissions --cluster mylab
# Custom cache locationsparkrun setup fix-permissions --cluster mylab --cache-dir /data/hf-cache
# Preview what would be donesparkrun setup fix-permissions --cluster mylab --dry-runThis runs chown on the HuggingFace cache directory across all target hosts. It tries passwordless sudo first; if that fails, it prompts for a password once.
Automatic fix on every run:
sparkrun automatically attempts to fix cache ownership on remote hosts before syncing models. If non-interactive sudo is available (either via general NOPASSWD or a scoped sudoers entry), this happens silently. If sudo isn’t available, sparkrun logs a warning and continues — the rsync may still succeed if permissions aren’t actually broken.
To make this automatic fix work without ever needing a password, run --save-sudo once:
# Install a scoped sudoers entry on all cluster hosts (one-time setup)sparkrun setup fix-permissions --cluster mylab --save-sudoThis installs a minimal sudoers rule on each host that permits only chown -R <user> <cache_dir> without a password — no broader sudo privileges are granted. After this, every future sparkrun run silently fixes cache ownership before model distribution.
# Preview what the sudoers entry would look likesparkrun setup fix-permissions --cluster mylab --save-sudo --dry-runAlternatively, you can fix a single machine manually (adjust the path if you’ve set HF_HOME):
sudo chown -R $USER:$USER ~/.cache/huggingfacePrevention:
- Let sparkrun handle model downloading and distribution — it downloads as your user and rsyncs to target nodes.
- Avoid running
huggingface-cli downloador model download scripts inside inference containers. Download on the host instead. - Run
sparkrun setup fix-permissions --save-sudoonce per cluster so that cache ownership is corrected automatically on every launch.
Linux page cache consuming memory
Section titled “Linux page cache consuming memory”DGX Spark systems have 128 GB of unified CPU/GPU memory. The Linux kernel’s page cache can consume a significant portion of this, reducing the memory available for inference containers. While the kernel will reclaim cached pages under pressure, large cached datasets can cause initial allocation failures or slower model loading.
Symptoms:
- Out-of-memory errors when launching large models that should fit in VRAM
- Model loading is slower than expected on a freshly booted system with prior file I/O
free -hshows large amounts of memory in the “buff/cache” column
Fix:
sparkrun has a built-in command to drop the page cache across all cluster hosts:
# Clear page cache on your default clustersparkrun setup clear-cache
# Or target specific hostssparkrun setup clear-cache --cluster mylab
# Preview what would be donesparkrun setup clear-cache --cluster mylab --dry-runThis runs sync followed by writing 3 to /proc/sys/vm/drop_caches on each host, which safely drops the page cache, dentries, and inodes.
Automatic clearing on every run:
sparkrun automatically attempts to drop the page cache on all target hosts before launching containers. If non-interactive sudo is available (either via general NOPASSWD or a scoped sudoers entry), this happens silently. If sudo isn’t available, sparkrun logs a warning and continues.
To make this automatic clearing work without ever needing a password, run --save-sudo once:
# Install a scoped sudoers entry on all cluster hosts (one-time setup)sparkrun setup clear-cache --cluster mylab --save-sudoThis installs a minimal sudoers rule on each host that permits only writing to /proc/sys/vm/drop_caches without a password — no broader sudo privileges are granted. After this, every future sparkrun run silently clears the page cache before launching containers.
# Preview what the sudoers entry would look likesparkrun setup clear-cache --cluster mylab --save-sudo --dry-runSSH connection failures
Section titled “SSH connection failures”Symptoms:
Permission denied (publickey)when sparkrun tries to connect- SSH mesh completes with “Done” but passwordless SSH still doesn’t work
- SSH connection timeouts
Quick diagnosis:
Run the built-in SSH diagnostics to identify the exact problem:
sparkrun setup ssh --cluster mylab --diagnoseThis checks permissions, sshd_config settings, and key installation on each host, then prints pass/fail results with specific remediation steps.
Common causes (especially for cross-user setups):
- Home directory permissions — the most common cause. OpenSSH’s
StrictModesrejects keys when the remote user’s home directory is group- or world-writable:Terminal window # Fix on the remote hostchmod go-w ~dgxuser AuthorizedKeysFileoverride —sshd_configmay point to a non-default location like/etc/ssh/authorized_keys/%uinstead of~/.ssh/authorized_keysAllowUsers/AllowGroups—sshd_configmay restrict which users can SSH in
Other fixes:
- Check key permissions:
chmod 600 ~/.ssh/id_*andchmod 700 ~/.ssh - Re-establish the SSH mesh across cluster hosts:
Terminal window sparkrun setup ssh --cluster mylab - Verify your SSH agent is running and has the correct key loaded (
ssh-add -l) - Check
~/.ssh/configfor conflictingHostentries that might override settings for your DGX Spark IPs - Debug with verbose SSH output:
ssh -vv <host>to see exactly where authentication fails
Docker permission errors
Section titled “Docker permission errors”Symptoms:
permission deniederrors referencing the Docker socket (/var/run/docker.sock)
Fix:
Ensure the SSH user is in the docker group on each DGX Spark host:
sudo usermod -aG docker $USERYou must reconnect the SSH session after changing groups (or log out and back in on the host). sparkrun’s SSH mesh doesn’t pick up group changes from existing connections.
GPU not visible in containers (missing CDI spec)
Section titled “GPU not visible in containers (missing CDI spec)”Symptoms:
- Containers start but the workload sees no GPU
- Docker errors referring to
nvidia.com/gpuor a CDI device that cannot be resolved sparkrun setup checkreports FAIL onNVIDIA CDI spec (/etc/cdi/nvidia.yaml)with/etc/cdi/nvidia.yaml missing or empty, or WARN withN of M referenced paths missing — spec looks stale- A failed launch prints the fix directly: sparkrun recognizes CDI-shaped Docker errors and appends the command to run
Since 0.3.0 sparkrun requests GPUs through CDI, emitting --device nvidia.com/gpu=all rather than --gpus, so a host without a generated CDI spec cannot hand a GPU to the container. CDI is the portable path: it also works on hosts whose Docker rejects --gpus outright. It requires Docker 25 or newer — on an older daemon the flag is not understood, which produces the same “no GPU” symptom for a different reason.
Check first:
sparkrun setup check --cluster mylabThis probes every host without changing anything and tells you which ones are missing the spec.
Fix — on each affected host:
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yamlVerify the devices are now registered:
nvidia-ctk cdi list # expect nvidia.com/gpu=... entriessparkrun setup wizard performs this as its NVIDIA CDI step, so a host set up through the wizard already has it. Run the command manually when you added a host later, or provisioned one outside the wizard.
Container launch failures
Section titled “Container launch failures”Symptoms:
- Container starts and immediately exits
image not foundormanifest unknownerrors
Fixes:
- Pull the image manually to check availability:
docker pull <image> - Verify the image tag exists — typos in container references are common
- Check Docker disk space on the target host:
Terminal window docker system df - Inspect container logs for the exit reason:
Terminal window docker logs <container-name>
Port conflicts
Section titled “Port conflicts”Symptoms:
address already in useerrors when launching a workload
Fix:
Override the serve port with --port or -o:
sparkrun run my-recipe --port 9001sparkrun run my-recipe -o port=9001To find what’s using the default port:
ss -tlnp | grep 8000Model doesn’t fit in VRAM
Section titled “Model doesn’t fit in VRAM”Symptoms:
- OOM errors during model loading
CUDA out of memoryor similar allocation failures
Fixes:
- Check the VRAM estimate before launching:
Terminal window sparkrun show <recipe> - Lower
gpu_memory_utilizationto leave headroom for the system (e.g. 0.7):Terminal window sparkrun run <recipe> -o gpu_memory_utilization=0.7 - Reduce
max_model_lento shrink KV cache allocation:Terminal window sparkrun run <recipe> -o max_model_len=4096 - Use multi-node tensor parallelism to split the model across hosts:
Terminal window sparkrun run <recipe> --tp 2 - Try a quantized variant of the model (AWQ, GPTQ, FP8, etc.)
InfiniBand not detected
Section titled “InfiniBand not detected”Symptoms:
- NCCL falls back to TCP for multi-node communication
- Slow multi-node inference performance compared to expectations
Fix:
Run the CX7 configuration command to detect and set up ConnectX-7 interfaces:
sparkrun setup cx7 --cluster mylabVerify IB interfaces are up on each host:
ibstatip addr showCheck that CX7 IPs are reachable between all hosts in the cluster.
NCCL errors
Section titled “NCCL errors”Symptoms:
NCCL WARNmessages in container logs- Timeout during distributed initialization
Connection refusederrors between nodes
Fixes:
- Verify InfiniBand subnets match across all hosts — nodes must be on the same subnet to communicate
- Check firewall rules: NCCL and Ray use ports including 25000 and 46379. Ensure these are open between cluster hosts
- If NCCL auto-detection picks the wrong network interface, set it explicitly:
Terminal window sparkrun run <recipe> -o NCCL_SOCKET_IFNAME=enp3s0f0np0 - Enable detailed NCCL logging for debugging:
Terminal window sparkrun run <recipe> -o NCCL_DEBUG=INFO
Multi-node startup issues
Section titled “Multi-node startup issues”Symptoms:
- Containers start on all nodes but the serve command fails
- Workers can’t reach the head node
- Distributed initialization hangs or times out
Fixes:
- Verify the SSH mesh is healthy:
Terminal window sparkrun setup ssh --cluster mylab - Check InfiniBand connectivity between hosts (see InfiniBand not detected above)
- Ensure the same container image is present on all hosts — image ID mismatches can cause subtle failures. sparkrun automatically syncs container images to all target hosts during
sparkrun run. To force a full re-distribution, stop the workload and re-run the recipe. - For very large models, increase the distributed initialization timeout via runtime-specific environment variables
For systematic debugging, use diagnostics to collect detailed host information.
Capturing run diagnostics
Section titled “Capturing run diagnostics”If a launch fails or behaves unexpectedly and you need detailed post-mortem data, capture the full lifecycle:
sparkrun run my-recipe --collect-diagnostics run_diag.ndjsonThis records everything — recipe resolution, SSH commands, phase timing, container logs, health check attempts, and all log output at DEBUG level — into a single NDJSON file, even when console output is at the default verbosity. Attach this file when filing support issues.
For more details, see Run Diagnostics.