Skip to content

proxy

The sparkrun proxy commands manage a LiteLLM-powered gateway that discovers running sparkrun inference endpoints and exposes them through a single OpenAI-compatible API.

The proxy sits in front of one or more inference workloads launched by sparkrun and provides:

  • Live endpoint discovery using the same mechanism as sparkrun cluster status
  • Auto-discovery background process that periodically re-scans and syncs models
  • Health checking via GET /v1/models on each discovered endpoint
  • Deduplication of endpoints reachable on multiple network interfaces (e.g. management IP vs ConnectX-7 IP)
  • Model aliases so clients can address a model by a friendly name
  • Load/unload models through sparkrun proxy load / sparkrun proxy unload

The proxy runs LiteLLM via uvx --from 'litellm[proxy]==1.82.6' litellm — no permanent installation required.

Terminal window
# Launch some inference workloads first
sparkrun run qwen3-1.7b-vllm --cluster mylab
# Start the proxy (discovers endpoints automatically)
sparkrun proxy start --cluster mylab
# Query models through the unified API
curl http://localhost:4000/v1/models

Or use the proxy to manage the full lifecycle:

Terminal window
# Start the proxy
sparkrun proxy start
# Load a model through the proxy
sparkrun proxy load qwen3.5-0.8b-bf16-sglang
# Query models
curl http://localhost:4000/v1/models

Discovers running endpoints, generates a LiteLLM config, and launches the proxy. A background auto-discover process periodically re-scans and syncs models with the proxy.

Terminal window
sparkrun proxy start # port 4000
sparkrun proxy start --host 127.0.0.1 # recommended bind address
sparkrun proxy start --port 8080 # custom port
sparkrun proxy start --cluster mylab # discover from cluster hosts
sparkrun proxy start --hosts 10.0.0.1,10.0.0.2 # explicit host list
sparkrun proxy start --foreground # run in foreground (blocking)
sparkrun proxy start --master-key sk-mykey # require a bearer token
sparkrun proxy start --no-auto-discover # disable periodic re-scanning
sparkrun proxy start --discover-interval 60 # re-scan every 60s (default: 30)
sparkrun proxy start --restart # replace an already-running proxy
sparkrun proxy start --dry-run # show what would be done

By default the proxy daemonizes in the background. Logs are written to ~/.cache/sparkrun/proxy/litellm.log.

Starting when a proxy is already running is an error; pass --restart to replace it with one carrying the new settings. Explicitly-supplied settings are persisted to proxy.yaml either way.

OptionDescription
--portProxy listen port (default: 4000)
--hostBind address — persisted to proxy.yaml. Recommended: 127.0.0.1.
--master-keyBearer token for stateless LiteLLM auth (default: none)
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line)
--clusterUse a saved cluster by name
--foregroundRun in foreground (blocking)
--no-auto-discoverDisable periodic endpoint re-scanning
--discover-intervalSeconds between discovery sweeps (default: 30)
--restartIf a proxy is already running, stop it and start fresh with the new settings
--dry-run / -nShow what would be done without executing

Sends SIGTERM to the running proxy and its auto-discover process using the stored PIDs.

Terminal window
sparkrun proxy stop

Shows whether the proxy is running, its PID, bind address, auto-discover status, and lists models registered via the LiteLLM management API.

Terminal window
sparkrun proxy status

Lists models currently registered with the running proxy. With --refresh, re-discovers endpoints and syncs the proxy — adding newly available models and removing stale entries whose backends are no longer healthy.

Terminal window
sparkrun proxy models
sparkrun proxy models --refresh

Launches an inference workload (detached) and registers it with the running proxy.

Unlike plain sparkrun run, proxy load automatically avoids port conflicts. When no --port is specified, it checks the head host over SSH to find the first available port. If the desired port is occupied, it increments until a free port is found:

$ sparkrun proxy load qwen3-1.7b-vllm
# Uses port 8000
$ sparkrun proxy load qwen3.5-35b-a3b-fp8-sglang
# Note: port 8000 in use on 10.24.11.13, using 8001 instead

This is intentionally different from sparkrun run, which uses exactly the port specified (or the recipe default) and fails if it’s occupied. The proxy’s load command is designed for managing multiple concurrent models where automatic port assignment is expected.

OptionDescription
RECIPE_NAMERecipe to load (required)
--hosts / -HComma-separated host list
--hosts-fileFile with hosts (one per line)
--clusterUse a saved cluster by name
--tp / --tensor-parallelOverride tensor parallelism
--pp / --pipeline-parallelOverride pipeline parallelism
--gpu-memOverride GPU memory utilization (0.0–1.0)
--max-model-lenOverride maximum model context length
-o / --optionOverride recipe options: -o key=value (repeatable)
--imageOverride container image
--soloForce single-node mode
--portOverride serve port
--dry-run / -nShow what would be done without executing

Stops the inference workload containers directly (same logic as sparkrun stop) and syncs the proxy to remove the now-stale model entry.

Terminal window
sparkrun proxy unload qwen3-1.7b-vllm --cluster mylab

Manage model aliases so clients can reference models by friendly names.

Terminal window
sparkrun proxy alias add qwen3-small "Qwen/Qwen3-1.7B"
sparkrun proxy alias remove qwen3-small
sparkrun proxy alias list

An alias is saved to proxy.yaml immediately. If a proxy is running, the config is regenerated and the proxy restarted so the alias takes effect. An alias whose target model has no healthy backend is saved but skipped in the generated config — it starts working as soon as the target is loaded.

  1. Enumerates persisted job metadata (~/.cache/sparkrun/jobs/*.yaml) for every workload sparkrun has launched
  2. Takes a live cluster status snapshot to determine which of those are actually running — the same cross-executor source sparkrun cluster status uses, so native (local) workloads are visible too, not only Docker containers
  3. Normalises host IPs to management IPs (prefers management IPs over InfiniBand IPs)
  4. Performs parallel health checks via GET /v1/models (3-second timeout)
  5. Returns only healthy endpoints

When no host context is available, discovery falls back to metadata only and skips the liveness step.

When the proxy starts with auto-discover enabled (the default), a background process runs alongside the proxy:

  • Periodically re-discovers endpoints at the configured interval (default: 30 seconds)
  • Reconciles models and aliases in one pass — they share a single config file, so applying them separately would restart the proxy twice per sweep
  • Rewrites the config and restarts the proxy only when the desired model set actually differs from what is on disk
  • Re-reads the proxy PID each sweep, so it follows a restart instead of mistaking it for a shutdown
  • Monitors that PID and exits automatically when the proxy dies

Disable with --no-auto-discover or set auto_discover: false in proxy.yaml.

Persistent proxy settings are stored in ~/.config/sparkrun/proxy.yaml:

proxy:
port: 4000
host: 127.0.0.1 # recommended; unset keeps the legacy 0.0.0.0 + warning
master_key: null # set to require a bearer token (stateless; no DB)
auto_discover: true
discover_interval: 30 # seconds between re-scans
gateway: litellm # optional; pins the gateway implementation
aliases:
my-model: "Qwen/Qwen3-1.7B"
gpt-4: "Qwen/Qwen3-30B-A3B"

CLI flags override config file values for a given invocation, and explicitly supplied values are persisted back.

LiteLLM is currently the only gateway implementation, and it is enabled on every channel via the gateway.litellm feature flag. The proxy.gateway key exists so an alternate implementation can be selected later; exactly one gateway is used at a time. Pinning a gateway that is disabled or unknown is an error rather than a silent fallback.

Bringing the proxy up is gated by that flag. Managing an already-running proxy — stop, status, models, aliases — is not, so a proxy started while the flag was on stays stoppable if you later turn it off.

PathPurpose
~/.config/sparkrun/proxy.yamlPersistent proxy settings and aliases
~/.cache/sparkrun/proxy/litellm_config.yamlGenerated LiteLLM config
~/.cache/sparkrun/proxy/state.yamlRunning proxy state (PID, port, host, gateway, auto-discover PID, start time)
~/.cache/sparkrun/proxy/litellm.logProxy process stdout/stderr
~/.cache/sparkrun/proxy/autodiscover.yamlAuto-discover process config
~/.cache/sparkrun/proxy/autodiscover.logAuto-discover process stdout/stderr
~/.cache/sparkrun/jobs/*.yamlJob metadata used for endpoint discovery