Skip to content

benchmark

The sparkrun benchmark command runs end-to-end benchmarking: it launches inference, executes a benchmark framework against the running server, collects results, and stops inference — all in one command.

Terminal window
sparkrun benchmark <recipe> [options]

Benchmarks are grouped into categories, each with its own subcommand. Bare sparkrun benchmark <recipe> falls through to performance:

CommandMeasures
sparkrun benchmark performance <recipe> (alias perf)Throughput and latency. The default.
sparkrun benchmark tools <recipe>Tool-calling / function-calling behavior.
sparkrun benchmark run <recipe>Legacy entry point that imposes no category.

New categories appear automatically when a benchmarking plugin registers them, so sparkrun benchmark --help is the authoritative list for your install.

Terminal window
# Benchmark on localhost with defaults (uses llama-benchy)
sparkrun benchmark qwen3-1.7b-vllm --solo
# Benchmark using a specific profile from a registry
sparkrun benchmark qwen3-1.7b-sglang --profile spark-arena-v1
# Benchmark an already-running inference server (skip launch/stop)
sparkrun benchmark qwen3-1.7b-vllm --skip-run --hosts 192.168.11.13
# Override benchmark args (use -b for benchmark options)
sparkrun benchmark qwen3-1.7b-vllm --solo -b depth=0,4096,16384 -b concurrency=1,2,5
# Override recipe options (use -o for recipe overrides like gpu_memory_utilization)
sparkrun benchmark qwen3-1.7b-vllm --solo -o gpu_memory_utilization=0.9
# Dry run — shows what would be done
sparkrun benchmark qwen3-1.7b-vllm --solo --dry-run

The benchmark command follows a three-step flow:

  1. Launch inference — Starts the inference server using the same logic as sparkrun run. The server’s served_model_name is suppressed so the benchmark framework can address the model by its HuggingFace ID.
  2. Run benchmark — Executes the benchmarking framework (default: llama-benchy) against the running server. Results are streamed live and captured for export.
  3. Stop inference — Stops the inference server after benchmarking completes.

Each step can be controlled independently:

  • --skip-run skips step 1 (benchmark an existing server)
  • --no-stop skips step 3 (leave inference running after benchmarking)
OptionDescription
RECIPE_NAMERecipe to benchmark (required)
--hosts / -HComma-separated host list (first = head)
--hosts-fileFile with hosts (one per line, # comments)
--clusterUse a saved cluster by name
--soloForce single-node mode
--tp / --tensor-parallelOverride tensor parallelism
--pp / --pipeline-parallelOverride pipeline parallelism
--gpu-memOverride GPU memory utilization (0.0–1.0)
--max-model-lenOverride maximum model context length
--portOverride serve port
--imageOverride container image
-o / --optionOverride recipe options: -o key=value (repeatable)
--profileBenchmark profile name or file (@registry/name syntax)
--frameworkOverride benchmarking framework (default: llama-benchy)
--outputOutput file for results YAML
-b / --benchmark-optionOverride benchmark args: -b key=value (repeatable)
--api-key-envName of the env var to read the API key from (e.g. OPENAI_API_KEY)
--exit-on-first-fail / --no-exit-on-first-failAbort on first benchmark failure (default: enabled)
--no-stopDon’t stop inference after benchmarking
--skip-runSkip launching inference (benchmark existing instance)
--sync-tuningSync tuning configs from registries before benchmarking
--rootfulRun with --privileged as root inside container (legacy behavior)
--timeoutBenchmark timeout in seconds (default: 14400, or from the profile)
--freshForce a fresh start, deleting any prior state
--resumeResume prior state without prompting (mutually exclusive with --fresh)
--arenaRun the opinionated Spark Arena flow and submit results
--dry-run / -nShow what would be done without executing

Benchmark arguments are resolved from multiple sources (highest priority first):

  1. CLI overrides (-b key=value)
  2. Benchmark profile (--profile)
  3. Recipe’s benchmark: block (embedded in recipe YAML)
  4. Framework defaults (e.g. llama-benchy defaults)

Profiles are standalone YAML files that define benchmark configurations. They can live in recipe registries alongside recipes.

Terminal window
# List available profiles
sparkrun registry list-benchmark-profiles
# Show profile details
sparkrun registry show-benchmark-profile spark-arena-v1
# Use a profile
sparkrun benchmark my-recipe --profile spark-arena-v1

Profiles from registries can use @registry/name syntax:

Terminal window
sparkrun benchmark my-recipe --profile @sparkrun-testing/spark-arena-v1

Recipes can embed model/recipe-specific default benchmark configuration:

model: Qwen/Qwen3-1.7B
runtime: vllm
container: scitrera/dgx-spark-vllm:0.16.0-t5
benchmark:
framework: llama-benchy
pp: [2048]
depth: [0, 4096]
prefix_caching: true

Results are saved as YAML (and JSON when available) with full metadata:

sparkrun_benchmark:
version: "1"
timestamp: "2026-02-28T..."
recipe:
name: qwen3-1.7b-vllm
model: Qwen/Qwen3-1.7B
runtime: vllm
cluster:
hosts: ["192.168.11.13"]
tp: 1
benchmark:
framework: llama-benchy
args:
pp: [2048]
depth: [0]
results:
rows:
- {pp: 2048, tg: 32, depth: 0, ...}

The output filename is auto-generated as benchmark_<recipe>_<profile>_tp<N>.yaml (with _pp<N> suffix when pipeline parallelism > 1) unless --output is specified.

Multi-task benchmark runs (a profile with several pp/tg/depth/concurrency combinations) are driven by a small scheduler that persists progress between tasks. State is checkpointed after each task, so an interrupted run can be resumed without re-executing completed tasks — useful when a long sweep is killed by a transient failure, an OOM, or Ctrl+C mid-sweep.

--exit-on-first-fail (default) stops the sweep on the first task failure so the issue can be addressed before more compute is spent; --no-exit-on-first-fail keeps marching through remaining tasks. State files live under ~/.cache/sparkrun/benchmarks/.

An interactive run prompts before resuming. For non-interactive use, decide up front:

Terminal window
sparkrun benchmark perf my-recipe --resume # continue prior state, no prompt
sparkrun benchmark perf my-recipe --fresh # discard prior state
sparkrun benchmark resume <benchmark-id> # resume a specific run by ID

A run’s identity is tied to the content of the recipe, so editing a recipe starts a new run instead of silently resuming state measured against different settings.

Within a run, the container image is pinned by its content-addressable digest on first launch and reused for every subsequent and resumed task. A re-pushed tag or a locally rebuilt image therefore cannot change the bits mid-sweep.

To keep results comparable across long sweeps, sparkrun pins the benchmarking framework version on the first task of a run and reuses that exact version for every subsequent task in the same run (including resumed tasks). For llama-benchy, this means the uvx llama-benchy@<version> reference is locked once and propagated to every task in the schedule — even if a newer release is published mid-sweep. The pinned version is captured in the benchmark output metadata.

The default benchmarking framework is llama-benchy, invoked via uvx llama-benchy. It benchmarks OpenAI-compatible inference endpoints with configurable prompt sizes, token generation counts, and concurrency levels.

Common benchmark args:

ArgTypeDescription
pplist[int]Prompt processing token counts (default: [2048])
tglist[int]Token generation counts (default: [32])
depthlist[int]Context depths (default: [0])
concurrencylist[int]Concurrent requests per test (default: [1])
runsintRuns per test (default: 3)
prefix_cachingboolEnable prefix caching measurement
Terminal window
# Example: comprehensive benchmark
sparkrun benchmark my-recipe --solo \
-b pp=2048 \
-b tg=32,128 \
-b depth=0,4096,16384,65536 \
-b concurrency=1,2,5,10 \
-b prefix_caching=true