proxy
The sparkrun proxy commands manage a LiteLLM-powered gateway that discovers running sparkrun inference endpoints and exposes them through a single OpenAI-compatible API.
Overview
Section titled “Overview”The proxy sits in front of one or more inference workloads launched by sparkrun and provides:
- Live endpoint discovery using the same mechanism as
sparkrun cluster status - Auto-discovery background process that periodically re-scans and syncs models
- Health checking via
GET /v1/modelson each discovered endpoint - Deduplication of endpoints reachable on multiple network interfaces (e.g. management IP vs ConnectX-7 IP)
- Model aliases so clients can address a model by a friendly name
- Load/unload models through
sparkrun proxy load/sparkrun proxy unload
The proxy runs LiteLLM via uvx --from 'litellm[proxy]==1.82.6' litellm — no
permanent installation required.
Quick start
Section titled “Quick start”# Launch some inference workloads firstsparkrun run qwen3-1.7b-vllm --cluster mylab
# Start the proxy (discovers endpoints automatically)sparkrun proxy start --cluster mylab
# Query models through the unified APIcurl http://localhost:4000/v1/modelsOr use the proxy to manage the full lifecycle:
# Start the proxysparkrun proxy start
# Load a model through the proxysparkrun proxy load qwen3.5-0.8b-bf16-sglang
# Query modelscurl http://localhost:4000/v1/modelsCommands
Section titled “Commands”sparkrun proxy start
Section titled “sparkrun proxy start”Discovers running endpoints, generates a LiteLLM config, and launches the proxy. A background auto-discover process periodically re-scans and syncs models with the proxy.
sparkrun proxy start # port 4000sparkrun proxy start --host 127.0.0.1 # recommended bind addresssparkrun proxy start --port 8080 # custom portsparkrun proxy start --cluster mylab # discover from cluster hostssparkrun proxy start --hosts 10.0.0.1,10.0.0.2 # explicit host listsparkrun proxy start --foreground # run in foreground (blocking)sparkrun proxy start --master-key sk-mykey # require a bearer tokensparkrun proxy start --no-auto-discover # disable periodic re-scanningsparkrun proxy start --discover-interval 60 # re-scan every 60s (default: 30)sparkrun proxy start --restart # replace an already-running proxysparkrun proxy start --dry-run # show what would be doneBy default the proxy daemonizes in the background. Logs are written to ~/.cache/sparkrun/proxy/litellm.log.
Starting when a proxy is already running is an error; pass --restart to
replace it with one carrying the new settings. Explicitly-supplied settings are
persisted to proxy.yaml either way.
Options
Section titled “Options”| Option | Description |
|---|---|
--port | Proxy listen port (default: 4000) |
--host | Bind address — persisted to proxy.yaml. Recommended: 127.0.0.1. |
--master-key | Bearer token for stateless LiteLLM auth (default: none) |
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line) |
--cluster | Use a saved cluster by name |
--foreground | Run in foreground (blocking) |
--no-auto-discover | Disable periodic endpoint re-scanning |
--discover-interval | Seconds between discovery sweeps (default: 30) |
--restart | If a proxy is already running, stop it and start fresh with the new settings |
--dry-run / -n | Show what would be done without executing |
sparkrun proxy stop
Section titled “sparkrun proxy stop”Sends SIGTERM to the running proxy and its auto-discover process using the stored PIDs.
sparkrun proxy stopsparkrun proxy status
Section titled “sparkrun proxy status”Shows whether the proxy is running, its PID, bind address, auto-discover status, and lists models registered via the LiteLLM management API.
sparkrun proxy statussparkrun proxy models
Section titled “sparkrun proxy models”Lists models currently registered with the running proxy. With --refresh, re-discovers endpoints and syncs the proxy — adding newly available models and removing stale entries whose backends are no longer healthy.
sparkrun proxy modelssparkrun proxy models --refreshsparkrun proxy load <recipe>
Section titled “sparkrun proxy load <recipe>”Launches an inference workload (detached) and registers it with the running proxy.
Unlike plain sparkrun run, proxy load automatically avoids port conflicts. When no --port is specified, it checks the head host over SSH to find the first available port. If the desired port is occupied, it increments until a free port is found:
$ sparkrun proxy load qwen3-1.7b-vllm# Uses port 8000
$ sparkrun proxy load qwen3.5-35b-a3b-fp8-sglang# Note: port 8000 in use on 10.24.11.13, using 8001 insteadThis is intentionally different from sparkrun run, which uses exactly the port specified (or the recipe default) and fails if it’s occupied. The proxy’s load command is designed for managing multiple concurrent models where automatic port assignment is expected.
Options
Section titled “Options”| Option | Description |
|---|---|
RECIPE_NAME | Recipe to load (required) |
--hosts / -H | Comma-separated host list |
--hosts-file | File with hosts (one per line) |
--cluster | Use a saved cluster by name |
--tp / --tensor-parallel | Override tensor parallelism |
--pp / --pipeline-parallel | Override pipeline parallelism |
--gpu-mem | Override GPU memory utilization (0.0–1.0) |
--max-model-len | Override maximum model context length |
-o / --option | Override recipe options: -o key=value (repeatable) |
--image | Override container image |
--solo | Force single-node mode |
--port | Override serve port |
--dry-run / -n | Show what would be done without executing |
sparkrun proxy unload <recipe>
Section titled “sparkrun proxy unload <recipe>”Stops the inference workload containers directly (same logic as sparkrun stop) and syncs the proxy to remove the now-stale model entry.
sparkrun proxy unload qwen3-1.7b-vllm --cluster mylabsparkrun proxy alias
Section titled “sparkrun proxy alias”Manage model aliases so clients can reference models by friendly names.
sparkrun proxy alias add qwen3-small "Qwen/Qwen3-1.7B"sparkrun proxy alias remove qwen3-smallsparkrun proxy alias listAn alias is saved to proxy.yaml immediately. If a proxy is running, the
config is regenerated and the proxy restarted so the alias takes effect. An
alias whose target model has no healthy backend is saved but skipped in the
generated config — it starts working as soon as the target is loaded.
How discovery works
Section titled “How discovery works”- Enumerates persisted job metadata (
~/.cache/sparkrun/jobs/*.yaml) for every workload sparkrun has launched - Takes a live cluster status snapshot to determine which of those are actually running — the same cross-executor source
sparkrun cluster statususes, so native (local) workloads are visible too, not only Docker containers - Normalises host IPs to management IPs (prefers management IPs over InfiniBand IPs)
- Performs parallel health checks via
GET /v1/models(3-second timeout) - Returns only healthy endpoints
When no host context is available, discovery falls back to metadata only and skips the liveness step.
Auto-discovery
Section titled “Auto-discovery”When the proxy starts with auto-discover enabled (the default), a background process runs alongside the proxy:
- Periodically re-discovers endpoints at the configured interval (default: 30 seconds)
- Reconciles models and aliases in one pass — they share a single config file, so applying them separately would restart the proxy twice per sweep
- Rewrites the config and restarts the proxy only when the desired model set actually differs from what is on disk
- Re-reads the proxy PID each sweep, so it follows a restart instead of mistaking it for a shutdown
- Monitors that PID and exits automatically when the proxy dies
Disable with --no-auto-discover or set auto_discover: false in proxy.yaml.
Configuration
Section titled “Configuration”Persistent proxy settings are stored in ~/.config/sparkrun/proxy.yaml:
proxy: port: 4000 host: 127.0.0.1 # recommended; unset keeps the legacy 0.0.0.0 + warning master_key: null # set to require a bearer token (stateless; no DB) auto_discover: true discover_interval: 30 # seconds between re-scans gateway: litellm # optional; pins the gateway implementation
aliases: my-model: "Qwen/Qwen3-1.7B" gpt-4: "Qwen/Qwen3-30B-A3B"CLI flags override config file values for a given invocation, and explicitly supplied values are persisted back.
Gateway selection
Section titled “Gateway selection”LiteLLM is currently the only gateway implementation, and it is enabled on
every channel via the gateway.litellm feature flag.
The proxy.gateway key exists so an alternate implementation can be selected
later; exactly one gateway is used at a time. Pinning a gateway that is
disabled or unknown is an error rather than a silent fallback.
Bringing the proxy up is gated by that flag. Managing an already-running
proxy — stop, status, models, aliases — is not, so a proxy started while
the flag was on stays stoppable if you later turn it off.
State files
Section titled “State files”| Path | Purpose |
|---|---|
~/.config/sparkrun/proxy.yaml | Persistent proxy settings and aliases |
~/.cache/sparkrun/proxy/litellm_config.yaml | Generated LiteLLM config |
~/.cache/sparkrun/proxy/state.yaml | Running proxy state (PID, port, host, gateway, auto-discover PID, start time) |
~/.cache/sparkrun/proxy/litellm.log | Proxy process stdout/stderr |
~/.cache/sparkrun/proxy/autodiscover.yaml | Auto-discover process config |
~/.cache/sparkrun/proxy/autodiscover.log | Auto-discover process stdout/stderr |
~/.cache/sparkrun/jobs/*.yaml | Job metadata used for endpoint discovery |