Skip to content

CLI Automation & Scripting

sparkrun is designed to be scriptable. This page covers flags, output formats, and patterns for using sparkrun in non-interactive contexts.

Several commands support --json for machine-readable output:

Terminal window
# Job health check with JSON output
sparkrun cluster check-job my-recipe --cluster mylab --json
# Cluster monitor (continuous JSON stream)
sparkrun cluster monitor --json

Use sparkrun cluster check-job to verify a running workload is healthy. The target can be a recipe name or a cluster ID (sparkrun_<hex>):

Terminal window
# Basic check — exits 0 if running, non-zero otherwise
sparkrun cluster check-job my-recipe
# JSON output for parsing
sparkrun cluster check-job my-recipe --cluster mylab --json
# Also verify the inference server responds at /v1/models
sparkrun cluster check-job my-recipe --check-http-models

The --ensure flag makes sparkrun run idempotent — if the workload is already running with the same configuration, it’s a no-op:

Terminal window
sparkrun run qwen3-1.7b-vllm --ensure

This is useful in scripts and cron jobs where you want to guarantee a workload is running without duplicating it.

FlagEffect
--no-followLaunch without following logs
--no-rmKeep containers after the workload stops (useful for log inspection)
--ensureIdempotent launch — skip if already running
--dry-run / -nShow what would be done without executing
Terminal window
# Launch in background without following logs
sparkrun run qwen3-1.7b-vllm --no-follow
# Keep containers after stop for debugging
sparkrun run qwen3-1.7b-vllm --no-rm

Generate and deploy systemd unit files for persistent inference services.

The recommended workflow is to first launch the workload manually, verify it works, then create a systemd service from the running job. This ensures the service configuration matches a known-good state:

Terminal window
# 1. Launch and verify the workload works
sparkrun run my-recipe --cluster mylab
# 2. Get the job ID from the running workload
sparkrun status
# 3. Preview the generated service file
sparkrun export systemd <job_id>
# 4. Install the service
sparkrun export systemd <job_id> --install

You can also generate a service directly from a recipe name (with optional overrides like --tp, --port, -o):

Terminal window
# Preview the generated service file
sparkrun export systemd my-recipe --cluster mylab
# Deploy and start immediately
sparkrun export systemd my-recipe --cluster mylab --install --start
# Remove the service
sparkrun export systemd my-recipe --cluster mylab --uninstall

The generated unit file uses sparkrun run --foreground --no-follow with Restart=on-failure, so systemd automatically restarts the workload if the inference process crashes.

CodeMeaning
0Success
1General error
2Invalid arguments or missing recipe

sparkrun supports tiered verbosity via the global -v flag (stackable) and a -q quiet mode for scripting:

FlagLevelOutput
(default)PROGRESSPhase and step output only
-vINFOAdds detail lines (model resolution, host detection, etc.)
-vvVERBOSEAdds timestamps and logger names for each message
-vvvDEBUGFull SSH command output, script content, remote stdout/stderr
-qWARNINGErrors and warnings only — suppresses all progress output
Terminal window
# Scripting: suppress progress output
sparkrun -q run my-recipe --no-follow
# Debugging: full SSH and script output
sparkrun -vvv run my-recipe --dry-run
# Moderate detail: timestamps + logger names
sparkrun -vv setup ssh --cluster mylab

When debugging a failed launch or performance issue, capture the full run lifecycle:

Terminal window
sparkrun run my-recipe --collect-diagnostics run_diag.ndjson

This records recipe resolution, phase timing, SSH commands, container logs, health checks, and errors into a structured NDJSON file. The diagnostics file captures all log output at DEBUG level even when the console shows default (quiet) output.

See Diagnostics for the full record type reference and processing examples.

VariablePurpose
HF_HOMEHuggingFace cache root directory (used by huggingface_hub for model downloads)
HF_HUB_CACHEHuggingFace model cache directory (overrides HF_HOME/hub)
HUGGINGFACE_HUB_CACHELegacy alias for HF_HUB_CACHE

sparkrun reads these via huggingface_hub’s standard resolution. See HuggingFace cache owned by root for the full resolution order.

For SSH user configuration, use sparkrun cluster create --user or --user on individual commands. For verbosity, use the -v / -q flags (see Verbosity levels).

#!/bin/bash
sparkrun run my-recipe --no-follow
sleep 10
sparkrun cluster check-job my-recipe --check-http-models
#!/bin/bash
# Idempotent — safe to run repeatedly
sparkrun run my-recipe --ensure --no-follow