Skip to content

SGLang

SGLang is a first-class runtime in sparkrun with support for solo and multi-node inference using SGLang’s native distributed backend.

  • Solo and multi-node clustering via SGLang’s built-in distribution
  • Native distributed backend (--dist-init-addr, --nnodes, --node-rank)
  • High-throughput inference
  • Structured generation support
  • Experimental GGUF support (quantized models via SGLang)

sparkrun works with ready-built images including:

model: Qwen/Qwen3-1.7B
runtime: sglang
min_nodes: 1
container: scitrera/dgx-spark-sglang:0.5.8-t5
metadata:
description: Qwen3 1.7B via SGLang
model_params: 1.7B
model_dtype: bf16
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
command: |
python3 -m sglang.launch_server \
--model-path {model} \
--host {host} \
--port {port} \
--tp {tensor_parallel}

For multi-node, sparkrun uses SGLang’s native distribution:

Terminal window
sparkrun run my-sglang-recipe --tp 2
OptionDefaultDescription
--init-port25000SGLang distributed init port

When a recipe sets speculative_draft_model_path (or the alias speculative_draft_model) in defaults, sparkrun automatically pre-syncs that model to all target hosts alongside the primary model. This happens during the distribution phase so the container never has to fetch the draft from HuggingFace at launch time (which would fail under sparkrun’s offline-hub default).

defaults:
speculative_algorithm: EAGLE
speculative_draft_model_path: lmsys/sglang-EAGLE-llama2-chat-7B
speculative_num_steps: 3
speculative_eagle_topk: 4

The value is also rendered into the generated command as --speculative-draft-model-path <model>. The alias key speculative_draft_model is accepted for convenience and normalized to the canonical key at command-render time.

SGLang has experimental support for GGUF quantized models. This lets you run quantized models through SGLang’s high-performance serving stack instead of llama.cpp.

Example: Qwen3.5-397B MoE (Q6_K GGUF, 4-node)

Section titled “Example: Qwen3.5-397B MoE (Q6_K GGUF, 4-node)”
model: unsloth/Qwen3.5-397B-A17B-GGUF:Q6_K
runtime: sglang
min_nodes: 4
container: scitrera/dgx-spark-sglang:0.5.8-t5
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 4
gpu_memory_utilization: 0.8
max_model_len: 200000
served_model_name: qwen3.5-397b
attention_backend: triton
tool_call_parser: qwen3_coder
tokenizer_path: Qwen/Qwen3.5-397B-A17B
command: |
python3 -m sglang.launch_server \
--model-path {model} \
--served-model-name {served_model_name} \
--context-length {max_model_len} \
--mem-fraction-static {gpu_memory_utilization} \
--tp-size {tensor_parallel} \
--host {host} \
--port {port} \
--attention-backend {attention_backend} \
--tokenizer-path {tokenizer_path} \
--tool-call-parser {tool_call_parser}
  • tokenizer_path — GGUF files don’t bundle tokenizer configs, so you must point to the original / non-GGUF HuggingFace repo (e.g. Qwen/Qwen3.5-397B-A17B)
  • Model field uses the same repo:quant syntax as llama-cpp GGUF recipes (e.g. unsloth/Qwen3.5-397B-A17B-GGUF:Q6_K)