Skip to content

Benchmark on Spark Arena

Spark Arena is a community leaderboard for DGX Spark inference performance. sparkrun integrates directly with Spark Arena — authenticate, benchmark, and submit results all from the CLI.

After logging in, the sparkrun arena benchmark command handles the rest:

  1. Launches the inference server
  2. Benchmarks using the standardized spark-arena-v1 profile
  3. Collects results (CSV, metadata, effective recipe)
  4. Uploads everything to the Spark Arena leaderboard
  5. Stops the inference server

Before submitting benchmarks, authenticate with your Spark Arena account:

Terminal window
sparkrun arena login

This opens your browser for OAuth authentication. Sign in with your preferred provider (GitHub, Google, etc.) and sparkrun stores a refresh token locally.

Verify your login status anytime:

Terminal window
sparkrun arena status

With authentication in place, run a benchmark and submit in one step:

Terminal window
sparkrun arena benchmark qwen3-1.7b-sglang

This uses your default cluster. To target a specific cluster or hosts:

Terminal window
# Use a saved cluster
sparkrun arena benchmark qwen3-1.7b-sglang --cluster mylab
# Or specify hosts directly
sparkrun arena benchmark qwen3-1.7b-sglang --hosts 10.24.11.13,10.24.11.14

For multi-node benchmarks, set tensor parallelism:

Terminal window
sparkrun arena benchmark qwen3-1.7b-sglang --tp 2

The standardized spark-arena-v1 benchmark profile is used automatically — this ensures all submissions are comparable on the leaderboard.

When the benchmark completes, sparkrun uploads three files to Spark Arena:

FileContents
recipe.yamlThe effective recipe (with all overrides applied)
benchmark.csvRaw benchmark results (prompt sizes, token counts, latencies, throughput)
metadata.jsonSystem metadata — sparkrun version, cluster info, runtime details, container image

All files are cached locally under ~/.cache/sparkrun/benchmarks/<submission-id>/ so you always have a copy.

You can apply the same recipe overrides as sparkrun run:

Terminal window
# Override GPU memory utilization
sparkrun arena benchmark qwen3-1.7b-sglang -o gpu_memory_utilization=0.9
# Override max model length
sparkrun arena benchmark qwen3-1.7b-sglang -o max_model_len=8192
# Override container image
sparkrun arena benchmark qwen3-1.7b-sglang --image my-custom-image:latest

These overrides are captured in the submission metadata, so the leaderboard accurately reflects your exact configuration.

Long spark-arena-v1 sweeps are checkpointed after every benchmark task, so if a submission is interrupted partway through — Ctrl+C, OOM, network blip, anything — just re-run the exact same command and sparkrun picks up where it left off, skipping the tasks that already completed.

Terminal window
# First attempt — interrupted partway through
sparkrun arena benchmark qwen3-1.7b-sglang --cluster mylab --tp 2
# Re-run — sparkrun resumes the sweep and submits when complete
sparkrun arena benchmark qwen3-1.7b-sglang --cluster mylab --tp 2

sparkrun also pins the benchmarking framework version (e.g. the llama-benchy release) on the first task of a run and reuses that exact version for every subsequent task — even after a resume — so the numbers you submit are internally consistent. The pinned version is captured in the uploaded metadata.json.

State files live under ~/.cache/sparkrun/benchmarks/. See Resumable runs in the benchmarking tutorial for more detail.

Your credentials persist across sessions — you only need to log in once. If you want to remove stored credentials (e.g., switching accounts or on a shared machine):

Terminal window
sparkrun arena logout
  • Browse the Spark Arena leaderboard to see community results
  • Try different recipes and runtimes to find the best configuration for your workload
  • See Benchmarking Models for more on sparkrun’s benchmark system (custom profiles, parameters, comparing configurations)