Skip to content

Proxy Gateway

This tutorial shows how to use sparkrun’s proxy gateway to serve multiple inference models through a single unified API endpoint, powered by LiteLLM.

The sparkrun proxy sits in front of your running inference workloads and provides:

  • A single OpenAI-compatible API endpoint (default: http://localhost:4000/v1)
  • Automatic discovery of running sparkrun inference endpoints
  • Model aliases so clients can use familiar names (e.g., gpt-4 routed to a local model)
  • Dynamic loading and unloading of models without restarting the proxy

Start by running a couple of inference workloads:

Terminal window
sparkrun run qwen3-1.7b-vllm
sparkrun run qwen3.5-0.8b-bf16-sglang

Launch the proxy gateway. It automatically discovers your running endpoints:

Terminal window
sparkrun proxy start --host 127.0.0.1

You’ll see output like:

Discovering inference endpoints...
Discovered 2 healthy endpoint(s):
10.24.11.13:8000 — Qwen/Qwen3-1.7B (vllm)
10.24.11.13:8001 — Qwen/Qwen3.5-0.8B (sglang)
Proxy started on 127.0.0.1:4000. API: http://localhost:4000/v1
Auto-discover enabled (every 30s)

If you’re using a cluster, pass it to the proxy so it can discover endpoints on those hosts:

Terminal window
sparkrun proxy start --cluster mylab --host 127.0.0.1

List the models registered with the proxy:

Terminal window
sparkrun proxy models

Or query the API directly:

Terminal window
curl http://localhost:4000/v1/models

Send requests to the proxy just like any OpenAI-compatible API. Select which model to use with the model field:

Terminal window
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-1.7B",
"messages": [{"role": "user", "content": "What is a DGX Spark?"}],
"max_tokens": 100
}'

To use a different model, change the model field:

Terminal window
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3.5-0.8B",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 50
}'

Create friendly names so clients can reference models without knowing the exact HuggingFace ID:

Terminal window
sparkrun proxy alias add my-assistant "Qwen/Qwen3-1.7B"

Now clients can use my-assistant as the model name:

Terminal window
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "my-assistant",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 50
}'

Aliases are applied immediately — no proxy restart needed. List all aliases with:

Terminal window
sparkrun proxy alias list

Remove an alias:

Terminal window
sparkrun proxy alias remove my-assistant

Load a new model and register it with the proxy in one command:

Terminal window
sparkrun proxy load qwen3-30b-a3b-vllm

This launches the inference workload (with automatic port conflict avoidance) and registers it with the running proxy. The proxy automatically picks an available port if the default is in use.

Stop a model and remove it from the proxy:

Terminal window
sparkrun proxy unload qwen3-30b-a3b-vllm

This stops the inference containers and syncs the proxy to remove the stale model entry.

See the proxy process status and registered models:

Terminal window
sparkrun proxy status

This shows the proxy PID, bind address, auto-discover status, and lists all registered models with their backend endpoints.

When you’re done, stop the proxy:

Terminal window
sparkrun proxy stop

By default, the proxy runs a background auto-discover process that periodically scans for new or removed inference endpoints (every 30 seconds). When you launch or stop models with sparkrun run or sparkrun stop, the proxy picks up the changes automatically.

To adjust the scan interval:

Terminal window
sparkrun proxy start --discover-interval 60

To disable auto-discovery:

Terminal window
sparkrun proxy start --no-auto-discover

With auto-discovery disabled, use sparkrun proxy models --refresh to manually trigger a re-scan.