Proxy Gateway
This tutorial shows how to use sparkrun’s proxy gateway to serve multiple inference models through a single unified API endpoint, powered by LiteLLM.
The sparkrun proxy sits in front of your running inference workloads and provides:
- A single OpenAI-compatible API endpoint (default:
http://localhost:4000/v1) - Automatic discovery of running sparkrun inference endpoints
- Model aliases so clients can use familiar names (e.g.,
gpt-4routed to a local model) - Dynamic loading and unloading of models without restarting the proxy
1. Launch some models
Section titled “1. Launch some models”Start by running a couple of inference workloads:
sparkrun run qwen3-1.7b-vllmsparkrun run qwen3.5-0.8b-bf16-sglang2. Start the proxy
Section titled “2. Start the proxy”Launch the proxy gateway. It automatically discovers your running endpoints:
sparkrun proxy start --host 127.0.0.1You’ll see output like:
Discovering inference endpoints...Discovered 2 healthy endpoint(s): 10.24.11.13:8000 — Qwen/Qwen3-1.7B (vllm) 10.24.11.13:8001 — Qwen/Qwen3.5-0.8B (sglang)
Proxy started on 127.0.0.1:4000. API: http://localhost:4000/v1Auto-discover enabled (every 30s)If you’re using a cluster, pass it to the proxy so it can discover endpoints on those hosts:
sparkrun proxy start --cluster mylab --host 127.0.0.13. Check available models
Section titled “3. Check available models”List the models registered with the proxy:
sparkrun proxy modelsOr query the API directly:
curl http://localhost:4000/v1/models4. Call the API
Section titled “4. Call the API”Send requests to the proxy just like any OpenAI-compatible API. Select which model to use with the model field:
curl http://localhost:4000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen/Qwen3-1.7B", "messages": [{"role": "user", "content": "What is a DGX Spark?"}], "max_tokens": 100 }'To use a different model, change the model field:
curl http://localhost:4000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen/Qwen3.5-0.8B", "messages": [{"role": "user", "content": "Hello!"}], "max_tokens": 50 }'5. Add model aliases
Section titled “5. Add model aliases”Create friendly names so clients can reference models without knowing the exact HuggingFace ID:
sparkrun proxy alias add my-assistant "Qwen/Qwen3-1.7B"Now clients can use my-assistant as the model name:
curl http://localhost:4000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "my-assistant", "messages": [{"role": "user", "content": "Hello!"}], "max_tokens": 50 }'Aliases are applied immediately — no proxy restart needed. List all aliases with:
sparkrun proxy alias listRemove an alias:
sparkrun proxy alias remove my-assistant6. Dynamic loading
Section titled “6. Dynamic loading”Load a new model and register it with the proxy in one command:
sparkrun proxy load qwen3-30b-a3b-vllmThis launches the inference workload (with automatic port conflict avoidance) and registers it with the running proxy. The proxy automatically picks an available port if the default is in use.
7. Unload a model
Section titled “7. Unload a model”Stop a model and remove it from the proxy:
sparkrun proxy unload qwen3-30b-a3b-vllmThis stops the inference containers and syncs the proxy to remove the stale model entry.
8. Check proxy status
Section titled “8. Check proxy status”See the proxy process status and registered models:
sparkrun proxy statusThis shows the proxy PID, bind address, auto-discover status, and lists all registered models with their backend endpoints.
9. Stop the proxy
Section titled “9. Stop the proxy”When you’re done, stop the proxy:
sparkrun proxy stopAuto-discovery
Section titled “Auto-discovery”By default, the proxy runs a background auto-discover process that periodically scans for new or removed inference endpoints (every 30 seconds). When you launch or stop models with sparkrun run or sparkrun stop, the proxy picks up the changes automatically.
To adjust the scan interval:
sparkrun proxy start --discover-interval 60To disable auto-discovery:
sparkrun proxy start --no-auto-discoverWith auto-discovery disabled, use sparkrun proxy models --refresh to manually trigger a re-scan.
Next steps
Section titled “Next steps”- Full reference: See the proxy CLI reference for all options and configuration details
- First model: If you haven’t launched a model yet, start with Your First Model
- Scale up: Run larger models across multiple Sparks with Multi-Node Tensor Parallelism