SGLang
SGLang is a first-class runtime in sparkrun with support for solo and multi-node inference using SGLang’s native distributed backend.
Features
Section titled “Features”- Solo and multi-node clustering via SGLang’s built-in distribution
- Native distributed backend (
--dist-init-addr,--nnodes,--node-rank) - High-throughput inference
- Structured generation support
- Experimental GGUF support (quantized models via SGLang)
Container images
Section titled “Container images”sparkrun works with ready-built images including:
Example recipe
Section titled “Example recipe”model: Qwen/Qwen3-1.7Bruntime: sglangmin_nodes: 1container: scitrera/dgx-spark-sglang:0.5.8-t5
metadata: description: Qwen3 1.7B via SGLang model_params: 1.7B model_dtype: bf16
defaults: port: 8000 host: 0.0.0.0 tensor_parallel: 1
command: | python3 -m sglang.launch_server \ --model-path {model} \ --host {host} \ --port {port} \ --tp {tensor_parallel}Multi-node inference
Section titled “Multi-node inference”For multi-node, sparkrun uses SGLang’s native distribution:
sparkrun run my-sglang-recipe --tp 2SGLang-specific options
Section titled “SGLang-specific options”| Option | Default | Description |
|---|---|---|
--init-port | 25000 | SGLang distributed init port |
Speculative decoding
Section titled “Speculative decoding”When a recipe sets speculative_draft_model_path (or the alias speculative_draft_model) in defaults, sparkrun automatically pre-syncs that model to all target hosts alongside the primary model. This happens during the distribution phase so the container never has to fetch the draft from HuggingFace at launch time (which would fail under sparkrun’s offline-hub default).
defaults: speculative_algorithm: EAGLE speculative_draft_model_path: lmsys/sglang-EAGLE-llama2-chat-7B speculative_num_steps: 3 speculative_eagle_topk: 4The value is also rendered into the generated command as --speculative-draft-model-path <model>. The alias key speculative_draft_model is accepted for convenience and normalized to the canonical key at command-render time.
Experimental: GGUF support
Section titled “Experimental: GGUF support”SGLang has experimental support for GGUF quantized models. This lets you run quantized models through SGLang’s high-performance serving stack instead of llama.cpp.
Example: Qwen3.5-397B MoE (Q6_K GGUF, 4-node)
Section titled “Example: Qwen3.5-397B MoE (Q6_K GGUF, 4-node)”model: unsloth/Qwen3.5-397B-A17B-GGUF:Q6_Kruntime: sglangmin_nodes: 4container: scitrera/dgx-spark-sglang:0.5.8-t5
defaults: port: 8000 host: 0.0.0.0 tensor_parallel: 4 gpu_memory_utilization: 0.8 max_model_len: 200000 served_model_name: qwen3.5-397b attention_backend: triton tool_call_parser: qwen3_coder tokenizer_path: Qwen/Qwen3.5-397B-A17B
command: | python3 -m sglang.launch_server \ --model-path {model} \ --served-model-name {served_model_name} \ --context-length {max_model_len} \ --mem-fraction-static {gpu_memory_utilization} \ --tp-size {tensor_parallel} \ --host {host} \ --port {port} \ --attention-backend {attention_backend} \ --tokenizer-path {tokenizer_path} \ --tool-call-parser {tool_call_parser}Key notes for SGLang GGUF
Section titled “Key notes for SGLang GGUF”tokenizer_path— GGUF files don’t bundle tokenizer configs, so you must point to the original / non-GGUF HuggingFace repo (e.g.Qwen/Qwen3.5-397B-A17B)- Model field uses the same
repo:quantsyntax as llama-cpp GGUF recipes (e.g.unsloth/Qwen3.5-397B-A17B-GGUF:Q6_K)