Benchmarking Models
This tutorial walks you through benchmarking inference workloads using sparkrun’s integrated benchmarking system. Performance benchmarks are powered by llama-benchy; tool-calling correctness is powered by tool-eval-bench. New benchmark categories appear automatically as plugins register them.
The sparkrun benchmark command automates the full benchmarking flow:
- Launch the inference server
- Run the benchmark against the running server
- Collect results and export to YAML, JSON, and CSV
- Stop the inference server
You can skip any step — benchmark an already-running server, keep the server alive after benchmarking, or both.
1. Browse benchmark profiles
Section titled “1. Browse benchmark profiles”Benchmark profiles define reusable test configurations (prompt sizes, token generation counts, concurrency levels). Browse available profiles from your registries:
sparkrun registry list-benchmark-profilesTo see what a profile includes:
sparkrun registry show-benchmark-profile spark-arena-v12. Pick a benchmark category
Section titled “2. Pick a benchmark category”sparkrun groups benchmarks by category — the kind of question being asked about a deployment:
| Category | Subcommand | Default framework | Measures |
|---|---|---|---|
| performance | sparkrun benchmark performance (alias perf) | llama-benchy | Throughput, latency, TTFT, token-rate |
| tools | sparkrun benchmark tools | tool-eval-bench | Tool-call correctness across a fixed scenario suite |
The simplest benchmark uses a recipe and lets sparkrun handle everything:
sparkrun benchmark perf qwen3-1.7b-sglang --soloThis launches the model, runs the default performance benchmark, saves results, and stops the model.
Bare sparkrun benchmark <recipe> (without a category) is preserved and routes to performance, so existing scripts keep working unchanged.
To use a specific benchmark profile:
sparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1For multi-node benchmarks, add --tp:
sparkrun benchmark perf qwen3-1.7b-sglang --tp 2 --profile spark-arena-v1To run a tool-calling correctness benchmark instead:
sparkrun benchmark tools qwen3-1.7b-sglang --solo3. Output formats
Section titled “3. Output formats”Benchmark results are automatically saved in multiple formats:
- YAML — full metadata including recipe details, cluster info, and raw results
- JSON — same data in JSON format
- CSV — tabular results for easy import into spreadsheets
Output files are named automatically based on the recipe and profile:
benchmark_qwen3-1.7b-sglang_spark-arena-v1_tp1.yamlbenchmark_qwen3-1.7b-sglang_spark-arena-v1_tp1.jsonbenchmark_qwen3-1.7b-sglang_spark-arena-v1_tp1.csvTo specify a custom output path:
sparkrun benchmark qwen3-1.7b-sglang --solo --out my-results.yaml4. Benchmark a running model
Section titled “4. Benchmark a running model”If you already have a model serving, skip the launch step with --skip-run:
# Model is already running from a previous sparkrun runsparkrun benchmark qwen3-1.7b-sglang --skip-run --soloThis connects to the existing server, runs the benchmark, and leaves the server running (since sparkrun didn’t start it).
5. Keep the model alive after benchmarking
Section titled “5. Keep the model alive after benchmarking”Use --no-stop to leave the inference server running after the benchmark completes:
sparkrun benchmark qwen3-1.7b-sglang --solo --no-stopThis is useful when you want to run multiple benchmarks against the same model with different configurations, or when you want to interact with the model after benchmarking.
6. Compare configurations
Section titled “6. Compare configurations”A common workflow is comparing the same model at different TP levels:
# Benchmark at TP=1sparkrun benchmark qwen3-1.7b-sglang --solo --profile spark-arena-v1
# Benchmark at TP=2sparkrun benchmark qwen3-1.7b-sglang --tp 2 --profile spark-arena-v1Each run produces separate output files with the TP level in the filename, making side-by-side comparison straightforward.
7. Custom benchmark parameters
Section titled “7. Custom benchmark parameters”Override individual benchmark parameters with -b:
sparkrun benchmark qwen3-1.7b-sglang --solo \ -b pp=2048 \ -b tg=128 \ -b concurrency=5Common benchmark parameters:
| Parameter | Description | Example |
|---|---|---|
pp | Prompt processing token counts | -b pp=2048 or -b pp=512,2048,4096 |
tg | Token generation counts | -b tg=32,128 |
depth | Context depths | -b depth=0,4096,16384 |
concurrency | Concurrent requests | -b concurrency=1,2,5,10 |
runs | Runs per test point | -b runs=3 |
prefix_caching | Enable prefix caching measurement | -b prefix_caching=true |
Multiple values (comma-separated) create a matrix of test points.
8. Resumable runs
Section titled “8. Resumable runs”Long benchmark sweeps (a profile with many pp × tg × depth × concurrency combinations) can take hours. sparkrun checkpoints progress after every task, so an interrupted sweep — Ctrl+C, OOM, transient SSH hiccup, anything — picks up exactly where it left off on the next invocation.
Re-run the same command and sparkrun prompts to confirm resume:
# First attempt — interrupted partway throughsparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1 --solo
# Re-run — interactive prompt confirms resume (defaults to yes)sparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1 --soloFor automation (CI, agents, scripts) use --resume to skip the prompt:
# Non-interactive: resume if state exists, fresh otherwisesparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1 --solo --resume--resume and --fresh are mutually exclusive. --fresh deletes prior state and starts over. With neither flag, non-TTY invocations (CI, piped stdin) default to resume; TTY invocations prompt.
By default, --exit-on-first-fail halts the sweep on the first failing task so the issue can be fixed before more compute is spent. Pass --no-exit-on-first-fail to keep going through the remaining tasks instead.
To keep results comparable across long, possibly-resumed sweeps, sparkrun pins two things on the first successful launch:
- Framework version — e.g. the
llama-benchyrelease used. Subsequent tasks use the same version even if a newer release ships mid-sweep. - Container image SHA — the content-addressable image ID resolved from a target host. On resume the launch is overridden to that exact SHA, so a re-pushed tag or rebuilt local image cannot silently change the bits under test between sessions.
Both values are recorded in the output metadata.
State files live under ~/.cache/sparkrun/benchmarks/<benchmark_id>/.
9. Preview before running
Section titled “9. Preview before running”Use --dry-run to see what sparkrun would do without executing anything:
sparkrun benchmark perf qwen3-1.7b-sglang --solo --profile spark-arena-v1 --dry-run10. Submit to Spark Arena
Section titled “10. Submit to Spark Arena”Add --arena to any perf invocation to run the opinionated Spark Arena profile (@official/spark-arena-v2) and upload results to the leaderboard:
sparkrun benchmark perf qwen3-1.7b-sglang --solo --arenaThe legacy sparkrun arena benchmark <recipe> continues to work as an alias and uses the same underlying flow (auth check, hardcoded profile, upload). See the Spark Arena tutorial for the full submission walkthrough.
11. Using the Python API
Section titled “11. Using the Python API”Everything in this tutorial is also available as a typed Python API for automation:
from sparkrun.api import benchmark, BenchmarkOptions, ResumeMode
result = benchmark(BenchmarkOptions( recipe="qwen3-1.7b-sglang", category="performance", # or "tools", "perf" alias accepted in CLI only profile="spark-arena-v1", solo=True, bench_args={"pp": [2048], "tg": [128]}, # equivalent to -b pp=2048 -b tg=128 resume=ResumeMode.IF_EXISTS, # or AUTO / FRESH / REQUIRED arena=False, # set True for arena submission flow))
print(result.success, result.benchmark_id)print(result.container_image_sha) # pinned digest used at launchprint(result.outputs["yaml"]) # path to exported result yamlThe API never writes to stdout or calls sys.exit; failures raise typed exceptions (BenchmarkFailed, NoResumableState, FrameworkCategoryMismatch, etc.).
Next steps
Section titled “Next steps”- Submit to the leaderboard: Share your results on Spark Arena with one command
- Serve multiple models: Use the Proxy Gateway to expose models through a unified API
- Explore profiles: Check the benchmark CLI reference for the full option set