Skip to content

Tips & Troubleshooting

DGX Spark inference containers typically run as root inside rootful Docker. If a container downloads a model from HuggingFace — either intentionally or because a file wasn’t pre-cached (e.g., a tokenizer config) — the downloaded files in the HuggingFace cache end up owned by root. This causes permission errors on subsequent runs when your non-root user can’t read or update the cache.

sparkrun avoids this by downloading models as your user on the host and syncing them to cluster nodes before launching containers. But if models were previously downloaded inside a container (or you ran a container manually that pulled a model), you may already have root-owned files.

Symptoms:

  • Permission denied or OSError: [Errno 13] errors referencing paths under your HuggingFace cache (typically ~/.cache/huggingface)
  • A model that worked before suddenly fails on launch
  • sparkrun’s model sync step fails with permission errors

Fix:

sparkrun has a built-in command that fixes ownership across all cluster hosts in one shot:

Terminal window
# Fix permissions on your default cluster
sparkrun setup fix-permissions
# Or target specific hosts
sparkrun setup fix-permissions --cluster mylab
# Custom cache location
sparkrun setup fix-permissions --cluster mylab --cache-dir /data/hf-cache
# Preview what would be done
sparkrun setup fix-permissions --cluster mylab --dry-run

This runs chown on the HuggingFace cache directory across all target hosts. It tries passwordless sudo first; if that fails, it prompts for a password once.

Automatic fix on every run:

sparkrun automatically attempts to fix cache ownership on remote hosts before syncing models. If non-interactive sudo is available (either via general NOPASSWD or a scoped sudoers entry), this happens silently. If sudo isn’t available, sparkrun logs a warning and continues — the rsync may still succeed if permissions aren’t actually broken.

To make this automatic fix work without ever needing a password, run --save-sudo once:

Terminal window
# Install a scoped sudoers entry on all cluster hosts (one-time setup)
sparkrun setup fix-permissions --cluster mylab --save-sudo

This installs a minimal sudoers rule on each host that permits only chown -R <user> <cache_dir> without a password — no broader sudo privileges are granted. After this, every future sparkrun run silently fixes cache ownership before model distribution.

Terminal window
# Preview what the sudoers entry would look like
sparkrun setup fix-permissions --cluster mylab --save-sudo --dry-run

Alternatively, you can fix a single machine manually (adjust the path if you’ve set HF_HOME):

Terminal window
sudo chown -R $USER:$USER ~/.cache/huggingface

Prevention:

  • Let sparkrun handle model downloading and distribution — it downloads as your user and rsyncs to target nodes.
  • Avoid running huggingface-cli download or model download scripts inside inference containers. Download on the host instead.
  • Run sparkrun setup fix-permissions --save-sudo once per cluster so that cache ownership is corrected automatically on every launch.

DGX Spark systems have 128 GB of unified CPU/GPU memory. The Linux kernel’s page cache can consume a significant portion of this, reducing the memory available for inference containers. While the kernel will reclaim cached pages under pressure, large cached datasets can cause initial allocation failures or slower model loading.

Symptoms:

  • Out-of-memory errors when launching large models that should fit in VRAM
  • Model loading is slower than expected on a freshly booted system with prior file I/O
  • free -h shows large amounts of memory in the “buff/cache” column

Fix:

sparkrun has a built-in command to drop the page cache across all cluster hosts:

Terminal window
# Clear page cache on your default cluster
sparkrun setup clear-cache
# Or target specific hosts
sparkrun setup clear-cache --cluster mylab
# Preview what would be done
sparkrun setup clear-cache --cluster mylab --dry-run

This runs sync followed by writing 3 to /proc/sys/vm/drop_caches on each host, which safely drops the page cache, dentries, and inodes.

Automatic clearing on every run:

sparkrun automatically attempts to drop the page cache on all target hosts before launching containers. If non-interactive sudo is available (either via general NOPASSWD or a scoped sudoers entry), this happens silently. If sudo isn’t available, sparkrun logs a warning and continues.

To make this automatic clearing work without ever needing a password, run --save-sudo once:

Terminal window
# Install a scoped sudoers entry on all cluster hosts (one-time setup)
sparkrun setup clear-cache --cluster mylab --save-sudo

This installs a minimal sudoers rule on each host that permits only writing to /proc/sys/vm/drop_caches without a password — no broader sudo privileges are granted. After this, every future sparkrun run silently clears the page cache before launching containers.

Terminal window
# Preview what the sudoers entry would look like
sparkrun setup clear-cache --cluster mylab --save-sudo --dry-run

Symptoms:

  • Permission denied (publickey) when sparkrun tries to connect
  • SSH mesh completes with “Done” but passwordless SSH still doesn’t work
  • SSH connection timeouts

Quick diagnosis:

Run the built-in SSH diagnostics to identify the exact problem:

Terminal window
sparkrun setup ssh --cluster mylab --diagnose

This checks permissions, sshd_config settings, and key installation on each host, then prints pass/fail results with specific remediation steps.

Common causes (especially for cross-user setups):

  1. Home directory permissions — the most common cause. OpenSSH’s StrictModes rejects keys when the remote user’s home directory is group- or world-writable:
    Terminal window
    # Fix on the remote host
    chmod go-w ~dgxuser
  2. AuthorizedKeysFile override — sshd_config may point to a non-default location like /etc/ssh/authorized_keys/%u instead of ~/.ssh/authorized_keys
  3. AllowUsers/AllowGroups — sshd_config may restrict which users can SSH in

Other fixes:

  • Check key permissions: chmod 600 ~/.ssh/id_* and chmod 700 ~/.ssh
  • Re-establish the SSH mesh across cluster hosts:
    Terminal window
    sparkrun setup ssh --cluster mylab
  • Verify your SSH agent is running and has the correct key loaded (ssh-add -l)
  • Check ~/.ssh/config for conflicting Host entries that might override settings for your DGX Spark IPs
  • Debug with verbose SSH output: ssh -vv <host> to see exactly where authentication fails

Symptoms:

  • permission denied errors referencing the Docker socket (/var/run/docker.sock)

Fix:

Ensure the SSH user is in the docker group on each DGX Spark host:

Terminal window
sudo usermod -aG docker $USER

You must reconnect the SSH session after changing groups (or log out and back in on the host). sparkrun’s SSH mesh doesn’t pick up group changes from existing connections.

GPU not visible in containers (missing CDI spec)

Section titled “GPU not visible in containers (missing CDI spec)”

Symptoms:

  • Containers start but the workload sees no GPU
  • Docker errors referring to nvidia.com/gpu or a CDI device that cannot be resolved
  • sparkrun setup check reports FAIL on NVIDIA CDI spec (/etc/cdi/nvidia.yaml) with /etc/cdi/nvidia.yaml missing or empty, or WARN with N of M referenced paths missing — spec looks stale
  • A failed launch prints the fix directly: sparkrun recognizes CDI-shaped Docker errors and appends the command to run

Since 0.3.0 sparkrun requests GPUs through CDI, emitting --device nvidia.com/gpu=all rather than --gpus, so a host without a generated CDI spec cannot hand a GPU to the container. CDI is the portable path: it also works on hosts whose Docker rejects --gpus outright. It requires Docker 25 or newer — on an older daemon the flag is not understood, which produces the same “no GPU” symptom for a different reason.

Check first:

Terminal window
sparkrun setup check --cluster mylab

This probes every host without changing anything and tells you which ones are missing the spec.

Fix — on each affected host:

Terminal window
sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml

Verify the devices are now registered:

Terminal window
nvidia-ctk cdi list # expect nvidia.com/gpu=... entries

sparkrun setup wizard performs this as its NVIDIA CDI step, so a host set up through the wizard already has it. Run the command manually when you added a host later, or provisioned one outside the wizard.

Symptoms:

  • Container starts and immediately exits
  • image not found or manifest unknown errors

Fixes:

  • Pull the image manually to check availability: docker pull <image>
  • Verify the image tag exists — typos in container references are common
  • Check Docker disk space on the target host:
    Terminal window
    docker system df
  • Inspect container logs for the exit reason:
    Terminal window
    docker logs <container-name>

Symptoms:

  • address already in use errors when launching a workload

Fix:

Override the serve port with --port or -o:

Terminal window
sparkrun run my-recipe --port 9001
sparkrun run my-recipe -o port=9001

To find what’s using the default port:

Terminal window
ss -tlnp | grep 8000

Symptoms:

  • OOM errors during model loading
  • CUDA out of memory or similar allocation failures

Fixes:

  • Check the VRAM estimate before launching:
    Terminal window
    sparkrun show <recipe>
  • Lower gpu_memory_utilization to leave headroom for the system (e.g. 0.7):
    Terminal window
    sparkrun run <recipe> -o gpu_memory_utilization=0.7
  • Reduce max_model_len to shrink KV cache allocation:
    Terminal window
    sparkrun run <recipe> -o max_model_len=4096
  • Use multi-node tensor parallelism to split the model across hosts:
    Terminal window
    sparkrun run <recipe> --tp 2
  • Try a quantized variant of the model (AWQ, GPTQ, FP8, etc.)

Symptoms:

  • NCCL falls back to TCP for multi-node communication
  • Slow multi-node inference performance compared to expectations

Fix:

Run the CX7 configuration command to detect and set up ConnectX-7 interfaces:

Terminal window
sparkrun setup cx7 --cluster mylab

Verify IB interfaces are up on each host:

Terminal window
ibstat
ip addr show

Check that CX7 IPs are reachable between all hosts in the cluster.

Symptoms:

  • NCCL WARN messages in container logs
  • Timeout during distributed initialization
  • Connection refused errors between nodes

Fixes:

  • Verify InfiniBand subnets match across all hosts — nodes must be on the same subnet to communicate
  • Check firewall rules: NCCL and Ray use ports including 25000 and 46379. Ensure these are open between cluster hosts
  • If NCCL auto-detection picks the wrong network interface, set it explicitly:
    Terminal window
    sparkrun run <recipe> -o NCCL_SOCKET_IFNAME=enp3s0f0np0
  • Enable detailed NCCL logging for debugging:
    Terminal window
    sparkrun run <recipe> -o NCCL_DEBUG=INFO

Symptoms:

  • Containers start on all nodes but the serve command fails
  • Workers can’t reach the head node
  • Distributed initialization hangs or times out

Fixes:

  • Verify the SSH mesh is healthy:
    Terminal window
    sparkrun setup ssh --cluster mylab
  • Check InfiniBand connectivity between hosts (see InfiniBand not detected above)
  • Ensure the same container image is present on all hosts — image ID mismatches can cause subtle failures. sparkrun automatically syncs container images to all target hosts during sparkrun run. To force a full re-distribution, stop the workload and re-run the recipe.
  • For very large models, increase the distributed initialization timeout via runtime-specific environment variables

For systematic debugging, use diagnostics to collect detailed host information.

If a launch fails or behaves unexpectedly and you need detailed post-mortem data, capture the full lifecycle:

Terminal window
sparkrun run my-recipe --collect-diagnostics run_diag.ndjson

This records everything — recipe resolution, SSH commands, phase timing, container logs, health check attempts, and all log output at DEBUG level — into a single NDJSON file, even when console output is at the default verbosity. Attach this file when filing support issues.

For more details, see Run Diagnostics.