benchmark
The sparkrun benchmark command runs end-to-end benchmarking: it launches inference, executes a benchmark framework against the running server, collects results, and stops inference — all in one command.
sparkrun benchmark <recipe> [options]Categories
Section titled “Categories”Benchmarks are grouped into categories, each with its own subcommand. Bare
sparkrun benchmark <recipe> falls through to performance:
| Command | Measures |
|---|---|
sparkrun benchmark performance <recipe> (alias perf) | Throughput and latency. The default. |
sparkrun benchmark tools <recipe> | Tool-calling / function-calling behavior. |
sparkrun benchmark run <recipe> | Legacy entry point that imposes no category. |
New categories appear automatically when a benchmarking plugin registers them,
so sparkrun benchmark --help is the authoritative list for your install.
Quick examples
Section titled “Quick examples”# Benchmark on localhost with defaults (uses llama-benchy)sparkrun benchmark qwen3-1.7b-vllm --solo
# Benchmark using a specific profile from a registrysparkrun benchmark qwen3-1.7b-sglang --profile spark-arena-v1
# Benchmark an already-running inference server (skip launch/stop)sparkrun benchmark qwen3-1.7b-vllm --skip-run --hosts 192.168.11.13
# Override benchmark args (use -b for benchmark options)sparkrun benchmark qwen3-1.7b-vllm --solo -b depth=0,4096,16384 -b concurrency=1,2,5
# Override recipe options (use -o for recipe overrides like gpu_memory_utilization)sparkrun benchmark qwen3-1.7b-vllm --solo -o gpu_memory_utilization=0.9
# Dry run — shows what would be donesparkrun benchmark qwen3-1.7b-vllm --solo --dry-runHow it works
Section titled “How it works”The benchmark command follows a three-step flow:
- Launch inference — Starts the inference server using the same logic as
sparkrun run. The server’sserved_model_nameis suppressed so the benchmark framework can address the model by its HuggingFace ID. - Run benchmark — Executes the benchmarking framework (default: llama-benchy) against the running server. Results are streamed live and captured for export.
- Stop inference — Stops the inference server after benchmarking completes.
Each step can be controlled independently:
--skip-runskips step 1 (benchmark an existing server)--no-stopskips step 3 (leave inference running after benchmarking)
Options
Section titled “Options”| Option | Description |
|---|---|
RECIPE_NAME | Recipe to benchmark (required) |
--hosts / -H | Comma-separated host list (first = head) |
--hosts-file | File with hosts (one per line, # comments) |
--cluster | Use a saved cluster by name |
--solo | Force single-node mode |
--tp / --tensor-parallel | Override tensor parallelism |
--pp / --pipeline-parallel | Override pipeline parallelism |
--gpu-mem | Override GPU memory utilization (0.0–1.0) |
--max-model-len | Override maximum model context length |
--port | Override serve port |
--image | Override container image |
-o / --option | Override recipe options: -o key=value (repeatable) |
--profile | Benchmark profile name or file (@registry/name syntax) |
--framework | Override benchmarking framework (default: llama-benchy) |
--output | Output file for results YAML |
-b / --benchmark-option | Override benchmark args: -b key=value (repeatable) |
--api-key-env | Name of the env var to read the API key from (e.g. OPENAI_API_KEY) |
--exit-on-first-fail / --no-exit-on-first-fail | Abort on first benchmark failure (default: enabled) |
--no-stop | Don’t stop inference after benchmarking |
--skip-run | Skip launching inference (benchmark existing instance) |
--sync-tuning | Sync tuning configs from registries before benchmarking |
--rootful | Run with --privileged as root inside container (legacy behavior) |
--timeout | Benchmark timeout in seconds (default: 14400, or from the profile) |
--fresh | Force a fresh start, deleting any prior state |
--resume | Resume prior state without prompting (mutually exclusive with --fresh) |
--arena | Run the opinionated Spark Arena flow and submit results |
--dry-run / -n | Show what would be done without executing |
Benchmark configuration
Section titled “Benchmark configuration”Benchmark arguments are resolved from multiple sources (highest priority first):
- CLI overrides (
-b key=value) - Benchmark profile (
--profile) - Recipe’s
benchmark:block (embedded in recipe YAML) - Framework defaults (e.g. llama-benchy defaults)
Benchmark profiles
Section titled “Benchmark profiles”Profiles are standalone YAML files that define benchmark configurations. They can live in recipe registries alongside recipes.
# List available profilessparkrun registry list-benchmark-profiles
# Show profile detailssparkrun registry show-benchmark-profile spark-arena-v1
# Use a profilesparkrun benchmark my-recipe --profile spark-arena-v1Profiles from registries can use @registry/name syntax:
sparkrun benchmark my-recipe --profile @sparkrun-testing/spark-arena-v1Recipe benchmark blocks
Section titled “Recipe benchmark blocks”Recipes can embed model/recipe-specific default benchmark configuration:
model: Qwen/Qwen3-1.7Bruntime: vllmcontainer: scitrera/dgx-spark-vllm:0.16.0-t5
benchmark: framework: llama-benchy pp: [2048] depth: [0, 4096] prefix_caching: trueResults
Section titled “Results”Results are saved as YAML (and JSON when available) with full metadata:
sparkrun_benchmark: version: "1" timestamp: "2026-02-28T..." recipe: name: qwen3-1.7b-vllm model: Qwen/Qwen3-1.7B runtime: vllm cluster: hosts: ["192.168.11.13"] tp: 1 benchmark: framework: llama-benchy args: pp: [2048] depth: [0] results: rows: - {pp: 2048, tg: 32, depth: 0, ...}The output filename is auto-generated as benchmark_<recipe>_<profile>_tp<N>.yaml (with _pp<N> suffix when pipeline parallelism > 1) unless --output is specified.
Scheduling and resumability
Section titled “Scheduling and resumability”Multi-task benchmark runs (a profile with several pp/tg/depth/concurrency combinations) are driven by a small scheduler that persists progress between tasks. State is checkpointed after each task, so an interrupted run can be resumed without re-executing completed tasks — useful when a long sweep is killed by a transient failure, an OOM, or Ctrl+C mid-sweep.
--exit-on-first-fail (default) stops the sweep on the first task failure so the issue can be addressed before more compute is spent; --no-exit-on-first-fail keeps marching through remaining tasks. State files live under ~/.cache/sparkrun/benchmarks/.
An interactive run prompts before resuming. For non-interactive use, decide up front:
sparkrun benchmark perf my-recipe --resume # continue prior state, no promptsparkrun benchmark perf my-recipe --fresh # discard prior statesparkrun benchmark resume <benchmark-id> # resume a specific run by IDIdentity and image pinning
Section titled “Identity and image pinning”A run’s identity is tied to the content of the recipe, so editing a recipe starts a new run instead of silently resuming state measured against different settings.
Within a run, the container image is pinned by its content-addressable digest on first launch and reused for every subsequent and resumed task. A re-pushed tag or a locally rebuilt image therefore cannot change the bits mid-sweep.
Framework version pinning
Section titled “Framework version pinning”To keep results comparable across long sweeps, sparkrun pins the benchmarking framework version on the first task of a run and reuses that exact version for every subsequent task in the same run (including resumed tasks). For llama-benchy, this means the uvx llama-benchy@<version> reference is locked once and propagated to every task in the schedule — even if a newer release is published mid-sweep. The pinned version is captured in the benchmark output metadata.
llama-benchy framework
Section titled “llama-benchy framework”The default benchmarking framework is llama-benchy, invoked via uvx llama-benchy. It benchmarks OpenAI-compatible inference endpoints with configurable prompt sizes, token generation counts, and concurrency levels.
Common benchmark args:
| Arg | Type | Description |
|---|---|---|
pp | list[int] | Prompt processing token counts (default: [2048]) |
tg | list[int] | Token generation counts (default: [32]) |
depth | list[int] | Context depths (default: [0]) |
concurrency | list[int] | Concurrent requests per test (default: [1]) |
runs | int | Runs per test (default: 3) |
prefix_caching | bool | Enable prefix caching measurement |
# Example: comprehensive benchmarksparkrun benchmark my-recipe --solo \ -b pp=2048 \ -b tg=32,128 \ -b depth=0,4096,16384,65536 \ -b concurrency=1,2,5,10 \ -b prefix_caching=true