Skip to content

Benchmarking Models

This tutorial walks you through benchmarking inference workloads using sparkrun’s integrated benchmarking system. Performance benchmarks are powered by llama-benchy; tool-calling correctness is powered by tool-eval-bench. New benchmark categories appear automatically as plugins register them.

The sparkrun benchmark command automates the full benchmarking flow:

  1. Launch the inference server
  2. Run the benchmark against the running server
  3. Collect results and export to YAML, JSON, and CSV
  4. Stop the inference server

You can skip any step — benchmark an already-running server, keep the server alive after benchmarking, or both.

Benchmark profiles define reusable test configurations (prompt sizes, token generation counts, concurrency levels). Browse available profiles from your registries:

Terminal window
sparkrun registry list-benchmark-profiles

To see what a profile includes:

Terminal window
sparkrun registry show-benchmark-profile spark-arena-v1

sparkrun groups benchmarks by category — the kind of question being asked about a deployment:

CategorySubcommandDefault frameworkMeasures
performancesparkrun benchmark performance (alias perf)llama-benchyThroughput, latency, TTFT, token-rate
toolssparkrun benchmark toolstool-eval-benchTool-call correctness across a fixed scenario suite

The simplest benchmark uses a recipe and lets sparkrun handle everything:

Terminal window
sparkrun benchmark perf qwen3-1.7b-sglang --solo

This launches the model, runs the default performance benchmark, saves results, and stops the model.

Bare sparkrun benchmark <recipe> (without a category) is preserved and routes to performance, so existing scripts keep working unchanged.

To use a specific benchmark profile:

Terminal window
sparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1

For multi-node benchmarks, add --tp:

Terminal window
sparkrun benchmark perf qwen3-1.7b-sglang --tp 2 --profile spark-arena-v1

To run a tool-calling correctness benchmark instead:

Terminal window
sparkrun benchmark tools qwen3-1.7b-sglang --solo

Benchmark results are automatically saved in multiple formats:

  • YAML — full metadata including recipe details, cluster info, and raw results
  • JSON — same data in JSON format
  • CSV — tabular results for easy import into spreadsheets

Output files are named automatically based on the recipe and profile:

benchmark_qwen3-1.7b-sglang_spark-arena-v1_tp1.yaml
benchmark_qwen3-1.7b-sglang_spark-arena-v1_tp1.json
benchmark_qwen3-1.7b-sglang_spark-arena-v1_tp1.csv

To specify a custom output path:

Terminal window
sparkrun benchmark qwen3-1.7b-sglang --solo --out my-results.yaml

If you already have a model serving, skip the launch step with --skip-run:

Terminal window
# Model is already running from a previous sparkrun run
sparkrun benchmark qwen3-1.7b-sglang --skip-run --solo

This connects to the existing server, runs the benchmark, and leaves the server running (since sparkrun didn’t start it).

5. Keep the model alive after benchmarking

Section titled “5. Keep the model alive after benchmarking”

Use --no-stop to leave the inference server running after the benchmark completes:

Terminal window
sparkrun benchmark qwen3-1.7b-sglang --solo --no-stop

This is useful when you want to run multiple benchmarks against the same model with different configurations, or when you want to interact with the model after benchmarking.

A common workflow is comparing the same model at different TP levels:

Terminal window
# Benchmark at TP=1
sparkrun benchmark qwen3-1.7b-sglang --solo --profile spark-arena-v1
# Benchmark at TP=2
sparkrun benchmark qwen3-1.7b-sglang --tp 2 --profile spark-arena-v1

Each run produces separate output files with the TP level in the filename, making side-by-side comparison straightforward.

Override individual benchmark parameters with -b:

Terminal window
sparkrun benchmark qwen3-1.7b-sglang --solo \
-b pp=2048 \
-b tg=128 \
-b concurrency=5

Common benchmark parameters:

ParameterDescriptionExample
ppPrompt processing token counts-b pp=2048 or -b pp=512,2048,4096
tgToken generation counts-b tg=32,128
depthContext depths-b depth=0,4096,16384
concurrencyConcurrent requests-b concurrency=1,2,5,10
runsRuns per test point-b runs=3
prefix_cachingEnable prefix caching measurement-b prefix_caching=true

Multiple values (comma-separated) create a matrix of test points.

Long benchmark sweeps (a profile with many pp × tg × depth × concurrency combinations) can take hours. sparkrun checkpoints progress after every task, so an interrupted sweep — Ctrl+C, OOM, transient SSH hiccup, anything — picks up exactly where it left off on the next invocation.

Re-run the same command and sparkrun prompts to confirm resume:

Terminal window
# First attempt — interrupted partway through
sparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1 --solo
# Re-run — interactive prompt confirms resume (defaults to yes)
sparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1 --solo

For automation (CI, agents, scripts) use --resume to skip the prompt:

Terminal window
# Non-interactive: resume if state exists, fresh otherwise
sparkrun benchmark perf qwen3-1.7b-sglang --profile spark-arena-v1 --solo --resume

--resume and --fresh are mutually exclusive. --fresh deletes prior state and starts over. With neither flag, non-TTY invocations (CI, piped stdin) default to resume; TTY invocations prompt.

By default, --exit-on-first-fail halts the sweep on the first failing task so the issue can be fixed before more compute is spent. Pass --no-exit-on-first-fail to keep going through the remaining tasks instead.

To keep results comparable across long, possibly-resumed sweeps, sparkrun pins two things on the first successful launch:

  • Framework version — e.g. the llama-benchy release used. Subsequent tasks use the same version even if a newer release ships mid-sweep.
  • Container image SHA — the content-addressable image ID resolved from a target host. On resume the launch is overridden to that exact SHA, so a re-pushed tag or rebuilt local image cannot silently change the bits under test between sessions.

Both values are recorded in the output metadata.

State files live under ~/.cache/sparkrun/benchmarks/<benchmark_id>/.

Use --dry-run to see what sparkrun would do without executing anything:

Terminal window
sparkrun benchmark perf qwen3-1.7b-sglang --solo --profile spark-arena-v1 --dry-run

Add --arena to any perf invocation to run the opinionated Spark Arena profile (@official/spark-arena-v2) and upload results to the leaderboard:

Terminal window
sparkrun benchmark perf qwen3-1.7b-sglang --solo --arena

The legacy sparkrun arena benchmark <recipe> continues to work as an alias and uses the same underlying flow (auth check, hardcoded profile, upload). See the Spark Arena tutorial for the full submission walkthrough.

Everything in this tutorial is also available as a typed Python API for automation:

from sparkrun.api import benchmark, BenchmarkOptions, ResumeMode
result = benchmark(BenchmarkOptions(
recipe="qwen3-1.7b-sglang",
category="performance", # or "tools", "perf" alias accepted in CLI only
profile="spark-arena-v1",
solo=True,
bench_args={"pp": [2048], "tg": [128]}, # equivalent to -b pp=2048 -b tg=128
resume=ResumeMode.IF_EXISTS, # or AUTO / FRESH / REQUIRED
arena=False, # set True for arena submission flow
))
print(result.success, result.benchmark_id)
print(result.container_image_sha) # pinned digest used at launch
print(result.outputs["yaml"]) # path to exported result yaml

The API never writes to stdout or calls sys.exit; failures raise typed exceptions (BenchmarkFailed, NoResumableState, FrameworkCategoryMismatch, etc.).

  • Submit to the leaderboard: Share your results on Spark Arena with one command
  • Serve multiple models: Use the Proxy Gateway to expose models through a unified API
  • Explore profiles: Check the benchmark CLI reference for the full option set