Managing Clusters
List all clusters
Section titled “List all clusters”sparkrun cluster listShows all saved clusters with their hosts and whether they are the default.
Show cluster details
Section titled “Show cluster details”sparkrun cluster show mylabDisplays the cluster’s hosts, description, SSH user, and default status.
Inspect effective configuration
Section titled “Inspect effective configuration”sparkrun cluster inspect mylabShows the resolved settings — transfer mode, transfer interface, topology, SSH user, cache directories — and checks whether those cache directories exist on each host. The fastest way to diagnose a transfer or permission problem without launching a job.
Cluster-level environment
Section titled “Cluster-level environment”A cluster can carry container environment variables that apply to every
workload launched on it. They sit below recipe env and CLI overrides in
precedence: CLI > recipe env > cluster env.
env: HF_TOKEN: "${HF_TOKEN}"env_file: /home/me/.sparkrun.env${VAR} references resolve at launch time from env_file — read locally on
the control machine and emitted per host, so the file never has to exist on the
cluster nodes. Secrets therefore stay in your env file and never land in the
cluster YAML.
This block is typically populated by sparkrun cluster import, which maps a
legacy CONTAINER_* block into it. Import owns the block and rewrites it
wholesale on re-sync.
Delete a cluster
Section titled “Delete a cluster”sparkrun cluster delete mylabRemoves the cluster configuration. Use --force to skip the confirmation prompt. This does not affect running workloads.
Default cluster
Section titled “Default cluster”Set a default
Section titled “Set a default”sparkrun cluster set-default mylabShow the current default
Section titled “Show the current default”sparkrun cluster defaultRemove the default
Section titled “Remove the default”sparkrun cluster unset-defaultWhen a default cluster is set, workload commands (run, stop, logs, status) use it automatically unless --hosts or --cluster is specified.
Cluster status
Section titled “Cluster status”sparkrun cluster statussparkrun cluster status --cluster mylabShows sparkrun containers running on cluster hosts — container names, images, and status.
Cluster monitor
Section titled “Cluster monitor”sparkrun cluster monitor # Interactive TUI (default)sparkrun cluster monitor --simple # Plain text outputsparkrun cluster monitor --json # JSON output for scriptingLive-monitors CPU, RAM, and GPU metrics across all hosts in the cluster. The TUI mode (default) uses Textual for a rich terminal interface with progress bars. Press q to quit.
| Option | Description |
|---|---|
--cluster | Target a specific cluster |
--interval | Refresh interval in seconds (default: 5) |
--simple | Plain text output instead of TUI |
--json | JSON output for scripting |
Check job status
Section titled “Check job status”sparkrun cluster check-job <recipe>Check whether a specific recipe’s containers are still running across cluster hosts.