CLI Automation & Scripting
sparkrun is designed to be scriptable. This page covers flags, output formats, and patterns for using sparkrun in non-interactive contexts.
JSON output
Section titled “JSON output”Several commands support --json for machine-readable output:
# Job health check with JSON outputsparkrun cluster check-job my-recipe --cluster mylab --json
# Cluster monitor (continuous JSON stream)sparkrun cluster monitor --jsonHealth checks
Section titled “Health checks”Use sparkrun cluster check-job to verify a running workload is healthy. The target can be a recipe name or a cluster ID (sparkrun_<hex>):
# Basic check — exits 0 if running, non-zero otherwisesparkrun cluster check-job my-recipe
# JSON output for parsingsparkrun cluster check-job my-recipe --cluster mylab --json
# Also verify the inference server responds at /v1/modelssparkrun cluster check-job my-recipe --check-http-modelsIdempotent launches
Section titled “Idempotent launches”The --ensure flag makes sparkrun run idempotent — if the workload is already running with the same configuration, it’s a no-op:
sparkrun run qwen3-1.7b-vllm --ensureThis is useful in scripts and cron jobs where you want to guarantee a workload is running without duplicating it.
Non-interactive flags
Section titled “Non-interactive flags”| Flag | Effect |
|---|---|
--no-follow | Launch without following logs |
--no-rm | Keep containers after the workload stops (useful for log inspection) |
--ensure | Idempotent launch — skip if already running |
--dry-run / -n | Show what would be done without executing |
# Launch in background without following logssparkrun run qwen3-1.7b-vllm --no-follow
# Keep containers after stop for debuggingsparkrun run qwen3-1.7b-vllm --no-rmsystemd services
Section titled “systemd services”Generate and deploy systemd unit files for persistent inference services.
The recommended workflow is to first launch the workload manually, verify it works, then create a systemd service from the running job. This ensures the service configuration matches a known-good state:
# 1. Launch and verify the workload workssparkrun run my-recipe --cluster mylab
# 2. Get the job ID from the running workloadsparkrun status
# 3. Preview the generated service filesparkrun export systemd <job_id>
# 4. Install the servicesparkrun export systemd <job_id> --installYou can also generate a service directly from a recipe name (with optional overrides like --tp, --port, -o):
# Preview the generated service filesparkrun export systemd my-recipe --cluster mylab
# Deploy and start immediatelysparkrun export systemd my-recipe --cluster mylab --install --start
# Remove the servicesparkrun export systemd my-recipe --cluster mylab --uninstallThe generated unit file uses sparkrun run --foreground --no-follow with Restart=on-failure, so systemd automatically restarts the workload if the inference process crashes.
Exit codes
Section titled “Exit codes”| Code | Meaning |
|---|---|
0 | Success |
1 | General error |
2 | Invalid arguments or missing recipe |
Verbosity and quiet mode
Section titled “Verbosity and quiet mode”sparkrun supports tiered verbosity via the global -v flag (stackable) and a -q quiet mode for scripting:
| Flag | Level | Output |
|---|---|---|
| (default) | PROGRESS | Phase and step output only |
-v | INFO | Adds detail lines (model resolution, host detection, etc.) |
-vv | VERBOSE | Adds timestamps and logger names for each message |
-vvv | DEBUG | Full SSH command output, script content, remote stdout/stderr |
-q | WARNING | Errors and warnings only — suppresses all progress output |
# Scripting: suppress progress outputsparkrun -q run my-recipe --no-follow
# Debugging: full SSH and script outputsparkrun -vvv run my-recipe --dry-run
# Moderate detail: timestamps + logger namessparkrun -vv setup ssh --cluster mylabRun diagnostics
Section titled “Run diagnostics”When debugging a failed launch or performance issue, capture the full run lifecycle:
sparkrun run my-recipe --collect-diagnostics run_diag.ndjsonThis records recipe resolution, phase timing, SSH commands, container logs, health checks, and errors into a structured NDJSON file. The diagnostics file captures all log output at DEBUG level even when the console shows default (quiet) output.
See Diagnostics for the full record type reference and processing examples.
Environment variables
Section titled “Environment variables”| Variable | Purpose |
|---|---|
HF_HOME | HuggingFace cache root directory (used by huggingface_hub for model downloads) |
HF_HUB_CACHE | HuggingFace model cache directory (overrides HF_HOME/hub) |
HUGGINGFACE_HUB_CACHE | Legacy alias for HF_HUB_CACHE |
sparkrun reads these via huggingface_hub’s standard resolution. See HuggingFace cache owned by root for the full resolution order.
For SSH user configuration, use sparkrun cluster create --user or --user on individual commands. For verbosity, use the -v / -q flags (see Verbosity levels).
Scripting patterns
Section titled “Scripting patterns”Launch and wait for health
Section titled “Launch and wait for health”#!/bin/bashsparkrun run my-recipe --no-followsleep 10sparkrun cluster check-job my-recipe --check-http-modelsEnsure a workload is running
Section titled “Ensure a workload is running”#!/bin/bash# Idempotent — safe to run repeatedlysparkrun run my-recipe --ensure --no-follow