Contributing
Quick start
Section titled “Quick start”git clone https://github.com/spark-arena/sparkrun.git -b develop-nextcd sparkrunsource dev.shsparkrun --helpsource dev.sh uses uv sync to manage the .venv and install sparkrun with dev dependencies, then installs the pre-commit hooks. After sourcing, the venv is activated in your current shell and sparkrun runs the code from your checkout — edits take effect immediately.
Requires uv (curl -LsSf https://astral.sh/uv/install.sh | sh).
Branch model
Section titled “Branch model”| Branch | Purpose |
|---|---|
main | Stable releases. Protected — no direct pushes |
develop | Release staging; the source of the beta update channel |
develop-next | Active integration; the source of the alpha channel. PRs target here |
feature/* | Feature branches off develop-next |
All PRs should target develop-next, not develop or main. Changes flow
feature/* → develop-next → develop → main.
Running tests
Section titled “Running tests”# Full suite.venv/bin/python -m pytest tests/ -v
# Single file.venv/bin/python -m pytest tests/test_recipe.py -v
# Specific test.venv/bin/python -m pytest tests/test_cli.py::test_run_command_basic -v
# With coverage.venv/bin/python -m pytest tests/ --cov=sparkrun --cov-report=term-missingAll tests are self-contained — no real hosts, SSH, or Docker needed. SSH/Docker operations are mocked via conftest.py fixtures.
Linting
Section titled “Linting”sparkrun uses ruff for linting and formatting:
ruff check src/ tests/ruff format src/ tests/Configuration: line-length 140, target Python 3.12 (in pyproject.toml).
Project layout
Section titled “Project layout”src/sparkrun/├── cli/ # Click CLI (one module per command group) — renders api/├── api/ # Console-free public library API├── core/ # Config, recipe, registry, launcher, scheduler, features├── runtimes/ # Runtime plugins (vllm, sglang, llama-cpp, trtllm, …)├── orchestration/ # SSH, Docker, executors, collectives, InfiniBand, scripts├── transports/ # Cluster connectivity seam├── schedulers/ # Placement schedulers├── platforms/ # Hardware platform registry├── builders/ # Builder plugins (eugr, docker-pull)├── models/ # HuggingFace download, distribution, VRAM estimation├── containers/ # Container image distribution├── tuning/ # Triton kernel tuning (SGLang, vLLM)├── benchmarking/ # Benchmark framework plugins├── diagnostics/ # Host and run diagnostic collection (NDJSON)├── proxy/ # Inference gateway├── telemetry/ # Anonymous usage telemetry├── utils/ # Shared helpers└── scripts/ # Embedded bash scripts (*.sh)tests/ # pytest tests (mirrors src/ structure)New logic belongs in api/ with the CLI rendering it — api/ must never import sparkrun.cli.
Key patterns
Section titled “Key patterns”Plugin system (SAF)
Section titled “Plugin system (SAF)”Runtimes, builders, benchmarking frameworks, executors, schedulers, transports, and telemetry providers are SAF multi-extension plugins. Each is discovered by scanning its module for subclasses of a base class:
cli/__init__.py → core/bootstrap.py → SAF init → find_types_in_modules() → register_plugin() → load_external_plugins()A plugin can gate itself behind a feature flag by declaring required_feature_flag. See Architecture for the full extension-point table.
Config chain (vpd)
Section titled “Config chain (vpd)”sparkrun uses vpd_chain for cascading config resolution throughout the codebase:
from vpd.legacy.yaml_dict import vpd_chainconfig = vpd_chain(cli_overrides, recipe_defaults, runtime_defaults)value = config.get("port") # resolves through the chainPriority: CLI → recipe → runtime defaults. The same pattern is used for executor config, recipe defaults, and benchmark profiles.
Executor abstraction
Section titled “Executor abstraction”Container engine operations go through the Executor ABC (orchestration/executor.py). DockerExecutor is the default and production-supported implementation. Runtimes use self.executor.* instead of importing docker.py directly:
# In a runtime:self.executor.run_cmd(image, command, container_name=name, env=env)self.executor.stop_cmd(container_name)self.executor.generate_launch_script(image, container_name, command, ...)self.executor.container_name(cluster_id, "solo")self.executor.node_container_name(cluster_id, rank)Runtimes never construct an executor directly — RuntimePlugin._resolve_executor() delegates to orchestration.executor:resolve_executor(), the single sanctioned entry point for the layered resolution chain. An explicitly-requested executor that is unknown or gated off raises ExecutorUnavailableError rather than falling back to Docker. See Executors.
SSH execution model
Section titled “SSH execution model”All remote operations use SSH stdin piping — scripts are generated as Python strings and piped to ssh host bash -s. No files are ever copied to remote hosts for execution.
from sparkrun.orchestration.ssh import run_remote_scriptresult = run_remote_script(host, script_string, timeout=120, **ssh_kwargs)Runtime architecture
Section titled “Runtime architecture”All runtimes extend RuntimePlugin (runtimes/base.py):
generate_command()— produce the serve command from recipe + overridesresolve_container()— resolve the container imagerun()/stop()— solo/cluster dispatch (base class handles solo; subclasses implement_run_cluster)cluster_strategy()—"ray"or"native"determines orchestration pathget_extra_docker_opts(),get_extra_volumes(),get_extra_env()— runtime-specific hooks
Test isolation
Section titled “Test isolation”conftest.py provides an isolate_stateful autouse fixture that redirects SAF’s stateful root to tmp_path. Tests never touch ~/.config/sparkrun/. The bootstrap singleton is reset between tests.
Adding a new runtime
Section titled “Adding a new runtime”- Create
src/sparkrun/runtimes/my_runtime.pyextendingRuntimePlugin - Set
runtime_name = "my-runtime"anddefault_image_prefix - Implement
generate_command()and optionallyresolve_container() - For multi-node: implement
_run_cluster()and_stop_cluster() - Register the entry point in
pyproject.toml:[project.entry-points."sparkrun.runtimes"]my_runtime = "sparkrun.runtimes.my_runtime:MyRuntime" - Add tests in
tests/test_my_runtime.py
Adding a new builder
Section titled “Adding a new builder”- Create
src/sparkrun/builders/my_builder.pyextendingBuilderPlugin - Set
builder_name = "my-builder" - Implement
prepare_image()— must return the final image name - Register in
pyproject.tomlundersparkrun.builders - Recipes reference it as
builder: my-builder
Version management
Section titled “Version management”Versions are tracked in versions.yaml at the repo root:
# Sync versions across pyproject.toml and companion packagespython scripts/update-versions.py
# CI check (verify, don't write)python scripts/update-versions.py --checkPR guidelines
Section titled “PR guidelines”- Target
develop-nextbranch for all PRs - Keep commits atomic — one logical change per commit
- Run
pytestandruff checkbefore pushing - Use
--dry-runto verify CLI changes produce correct Docker commands - Include test coverage for new functionality