Benchmark on Spark Arena
Spark Arena is a community leaderboard for DGX Spark inference performance. sparkrun integrates directly with Spark Arena — authenticate, benchmark, and submit results all from the CLI.
After logging in, the sparkrun arena benchmark command handles the rest:
- Launches the inference server
- Benchmarks using the standardized
spark-arena-v1profile - Collects results (CSV, metadata, effective recipe)
- Uploads everything to the Spark Arena leaderboard
- Stops the inference server
1. Log in to Spark Arena
Section titled “1. Log in to Spark Arena”Before submitting benchmarks, authenticate with your Spark Arena account:
sparkrun arena loginThis opens your browser for OAuth authentication. Sign in with your preferred provider (GitHub, Google, etc.) and sparkrun stores a refresh token locally.
Verify your login status anytime:
sparkrun arena status2. Run a Spark Arena benchmark
Section titled “2. Run a Spark Arena benchmark”With authentication in place, run a benchmark and submit in one step:
sparkrun arena benchmark qwen3-1.7b-sglangThis uses your default cluster. To target a specific cluster or hosts:
# Use a saved clustersparkrun arena benchmark qwen3-1.7b-sglang --cluster mylab
# Or specify hosts directlysparkrun arena benchmark qwen3-1.7b-sglang --hosts 10.24.11.13,10.24.11.14For multi-node benchmarks, set tensor parallelism:
sparkrun arena benchmark qwen3-1.7b-sglang --tp 2The standardized spark-arena-v1 benchmark profile is used automatically — this ensures all submissions are comparable on the leaderboard.
3. What gets submitted
Section titled “3. What gets submitted”When the benchmark completes, sparkrun uploads three files to Spark Arena:
| File | Contents |
|---|---|
recipe.yaml | The effective recipe (with all overrides applied) |
benchmark.csv | Raw benchmark results (prompt sizes, token counts, latencies, throughput) |
metadata.json | System metadata — sparkrun version, cluster info, runtime details, container image |
All files are cached locally under ~/.cache/sparkrun/benchmarks/<submission-id>/ so you always have a copy.
4. Recipe overrides
Section titled “4. Recipe overrides”You can apply the same recipe overrides as sparkrun run:
# Override GPU memory utilizationsparkrun arena benchmark qwen3-1.7b-sglang -o gpu_memory_utilization=0.9
# Override max model lengthsparkrun arena benchmark qwen3-1.7b-sglang -o max_model_len=8192
# Override container imagesparkrun arena benchmark qwen3-1.7b-sglang --image my-custom-image:latestThese overrides are captured in the submission metadata, so the leaderboard accurately reflects your exact configuration.
5. Resumable submissions
Section titled “5. Resumable submissions”Long spark-arena-v1 sweeps are checkpointed after every benchmark task, so if a submission is interrupted partway through — Ctrl+C, OOM, network blip, anything — just re-run the exact same command and sparkrun picks up where it left off, skipping the tasks that already completed.
# First attempt — interrupted partway throughsparkrun arena benchmark qwen3-1.7b-sglang --cluster mylab --tp 2
# Re-run — sparkrun resumes the sweep and submits when completesparkrun arena benchmark qwen3-1.7b-sglang --cluster mylab --tp 2sparkrun also pins the benchmarking framework version (e.g. the llama-benchy release) on the first task of a run and reuses that exact version for every subsequent task — even after a resume — so the numbers you submit are internally consistent. The pinned version is captured in the uploaded metadata.json.
State files live under ~/.cache/sparkrun/benchmarks/. See Resumable runs in the benchmarking tutorial for more detail.
6. Log out (optional)
Section titled “6. Log out (optional)”Your credentials persist across sessions — you only need to log in once. If you want to remove stored credentials (e.g., switching accounts or on a shared machine):
sparkrun arena logoutNext steps
Section titled “Next steps”- Browse the Spark Arena leaderboard to see community results
- Try different recipes and runtimes to find the best configuration for your workload
- See Benchmarking Models for more on sparkrun’s benchmark system (custom profiles, parameters, comparing configurations)