Skip to content

sparkrun tune

Terminal window
sparkrun tune sglang <recipe> [options]
sparkrun tune vllm <recipe> [options]

Run kernel autotuning on a single host using a recipe’s container image. This generates optimal tile configurations (BLOCK_M/N/K, warps, stages) for fused MoE kernels — and, for vLLM, also for FP8 dense GEMM kernels — at each tensor parallel size. The resulting configs are saved locally and auto-mounted in future sparkrun run invocations — no manual configuration needed.

Tuning is particularly beneficial for Mixture-of-Experts (MoE) models and FP8 quantized models, where the fused-MoE / FP8 GEMM kernels are performance-critical paths. Default Triton configs are generic; tuning finds the best configuration for your specific hardware and model combination.

sparkrun tune sglang launches a tuning container on the target host, clones SGLang’s benchmark scripts inside it, detects the Triton version, and runs SGLang’s autotuner for each requested TP size before tearing the container down.

sparkrun tune vllm shells out to the vllm-tune backing engine on the target host. Sparkrun ensures the pinned vllm-tune git ref is checked out under ~/.cache/sparkrun/vllm-tune/<ref>/, then invokes vllm-tune.sh --standalone --foreground --mode {moe|fp8|all} once per TP size. vllm-tune manages its own standalone tuning container per TP, covers both fused MoE and FP8 dense GEMM kernels, and writes results into sparkrun’s flat tuning cache via its --export-sparkrun integration point. Sparkrun then rsyncs that cache back to the control machine.

In both cases, on subsequent sparkrun run invocations the configs are auto-detected and mounted with the appropriate environment variable (SGLANG_MOE_CONFIG_DIR for SGLang, VLLM_TUNED_CONFIG_FOLDER for vLLM).

sparkrun tune vllm requires jq, docker, and git on the remote host (the DGX Spark, not the control machine). The tuner runs a preflight check and bails with an actionable error if any are missing.

OptionDescriptionApplies to
--hosts / -HComma-separated host list (only the first host is used)both
--hosts-fileFile with hosts (one per line)both
--clusterUse a saved cluster by nameboth
--tpTP size(s) to tune (repeatable; default: 1,2,4,8)both
--imageOverride container imageboth
--output-dirOverride tuning config output directoryboth
--parallel / -jRun N tuning jobs concurrently (default: 1 = sequential)both
--dry-run / -nShow what would be done without executingboth
--skip-cloneSkip cloning benchmark scripts (if already in image)sglang only
--mode {moe,fp8,all}Kernels to tune (default: all)vllm only
--vllm-tune-refOverride the vllm-tune git ref pinned in configvllm only
Terminal window
# Tune for all default TP sizes (1, 2, 4, 8)
sparkrun tune sglang qwen3.5-35b-bf16-sglang -H 127.0.0.1
# Tune for a specific TP size
sparkrun tune sglang qwen3.5-35b-bf16-sglang -H 127.0.0.1 --tp 2
# Tune for multiple specific TP sizes
sparkrun tune sglang qwen3.5-35b-bf16-sglang -H 127.0.0.1 --tp 1 --tp 2 --tp 4
# Tune with 4 parallel jobs
sparkrun tune sglang qwen3.5-35b-bf16-sglang -H 127.0.0.1 -j4
# Preview without running
sparkrun tune sglang qwen3.5-35b-bf16-sglang -H 127.0.0.1 --dry-run

Requires an SGLang recipe (the recipe’s runtime must be sglang). Configs are saved to ~/.cache/sparkrun/tuning/sglang/ and auto-mounted via SGLANG_MOE_CONFIG_DIR.

Terminal window
# Tune both MoE and FP8 kernels for all default TP sizes (1, 2, 4, 8)
sparkrun tune vllm qwen3-moe-vllm -H 127.0.0.1
# FP8 dense GEMM only — much faster (~25 min) for non-MoE FP8 models
sparkrun tune vllm qwen3-4b-fp8 -H 127.0.0.1 --mode fp8 --tp 1
# MoE only
sparkrun tune vllm qwen3-moe-vllm -H 127.0.0.1 --mode moe --tp 4
# Tune for multiple specific TP sizes
sparkrun tune vllm qwen3-moe-vllm -H 127.0.0.1 --tp 1 --tp 2 --tp 4
# Tune with 4 parallel jobs
sparkrun tune vllm qwen3-moe-vllm -H 127.0.0.1 -j4
# Pin an in-flight vllm-tune branch for testing
sparkrun tune vllm qwen3-moe-vllm -H 127.0.0.1 --vllm-tune-ref my-pr-branch
# Preview without running
sparkrun tune vllm qwen3-moe-vllm -H 127.0.0.1 --dry-run

Accepts any vLLM recipe variant (vllm-ray, vllm-distributed, or eugr-vllm). Configs are saved to ~/.cache/sparkrun/tuning/vllm/ and auto-mounted via VLLM_TUNED_CONFIG_FOLDER.

The vllm-tune git URL and ref can be pinned in ~/.config/sparkrun/config.yaml:

tuning:
vllm_tune_repo: https://github.com/SeraphimSerapis/vllm-tune.git
vllm_tune_ref: main

Configs are stored on the host at:

RuntimeHost pathContainer pathEnv variable
SGLang~/.cache/sparkrun/tuning/sglang//tuning/sglang/configsSGLANG_MOE_CONFIG_DIR
vLLM~/.cache/sparkrun/tuning/vllm//tuning/vllmVLLM_TUNED_CONFIG_FOLDER

vLLM tuning configs are stored flat (one JSON per kernel shape); MoE filenames start with E=..., FP8 dense GEMM filenames with N=....

Once tuning configs exist on a host, every sparkrun run for the matching runtime will mount them automatically. No recipe changes or extra flags are needed.

  • Tuning runs on a single host even if you specify a cluster — only the first host is used.
  • Each TP size produces its own config. Tune for the TP sizes you actually plan to use.
  • Use -j4 to run multiple TP sizes in parallel and reduce total tuning time.
  • Use --dry-run to preview the Docker and tuning commands before committing to a long run.
  • Tuning can take hours per TP size depending on the model.