Skip to content

Contributing

Terminal window
git clone https://github.com/spark-arena/sparkrun.git -b develop-next
cd sparkrun
source dev.sh
sparkrun --help

source dev.sh uses uv sync to manage the .venv and install sparkrun with dev dependencies, then installs the pre-commit hooks. After sourcing, the venv is activated in your current shell and sparkrun runs the code from your checkout — edits take effect immediately.

Requires uv (curl -LsSf https://astral.sh/uv/install.sh | sh).

BranchPurpose
mainStable releases. Protected — no direct pushes
developRelease staging; the source of the beta update channel
develop-nextActive integration; the source of the alpha channel. PRs target here
feature/*Feature branches off develop-next

All PRs should target develop-next, not develop or main. Changes flow feature/* → develop-next → develop → main.

Terminal window
# Full suite
.venv/bin/python -m pytest tests/ -v
# Single file
.venv/bin/python -m pytest tests/test_recipe.py -v
# Specific test
.venv/bin/python -m pytest tests/test_cli.py::test_run_command_basic -v
# With coverage
.venv/bin/python -m pytest tests/ --cov=sparkrun --cov-report=term-missing

All tests are self-contained — no real hosts, SSH, or Docker needed. SSH/Docker operations are mocked via conftest.py fixtures.

sparkrun uses ruff for linting and formatting:

Terminal window
ruff check src/ tests/
ruff format src/ tests/

Configuration: line-length 140, target Python 3.12 (in pyproject.toml).

src/sparkrun/
├── cli/ # Click CLI (one module per command group) — renders api/
├── api/ # Console-free public library API
├── core/ # Config, recipe, registry, launcher, scheduler, features
├── runtimes/ # Runtime plugins (vllm, sglang, llama-cpp, trtllm, …)
├── orchestration/ # SSH, Docker, executors, collectives, InfiniBand, scripts
├── transports/ # Cluster connectivity seam
├── schedulers/ # Placement schedulers
├── platforms/ # Hardware platform registry
├── builders/ # Builder plugins (eugr, docker-pull)
├── models/ # HuggingFace download, distribution, VRAM estimation
├── containers/ # Container image distribution
├── tuning/ # Triton kernel tuning (SGLang, vLLM)
├── benchmarking/ # Benchmark framework plugins
├── diagnostics/ # Host and run diagnostic collection (NDJSON)
├── proxy/ # Inference gateway
├── telemetry/ # Anonymous usage telemetry
├── utils/ # Shared helpers
└── scripts/ # Embedded bash scripts (*.sh)
tests/ # pytest tests (mirrors src/ structure)

New logic belongs in api/ with the CLI rendering it — api/ must never import sparkrun.cli.

Runtimes, builders, benchmarking frameworks, executors, schedulers, transports, and telemetry providers are SAF multi-extension plugins. Each is discovered by scanning its module for subclasses of a base class:

cli/__init__.py → core/bootstrap.py → SAF init → find_types_in_modules() → register_plugin()
→ load_external_plugins()

A plugin can gate itself behind a feature flag by declaring required_feature_flag. See Architecture for the full extension-point table.

sparkrun uses vpd_chain for cascading config resolution throughout the codebase:

from vpd.legacy.yaml_dict import vpd_chain
config = vpd_chain(cli_overrides, recipe_defaults, runtime_defaults)
value = config.get("port") # resolves through the chain

Priority: CLI → recipe → runtime defaults. The same pattern is used for executor config, recipe defaults, and benchmark profiles.

Container engine operations go through the Executor ABC (orchestration/executor.py). DockerExecutor is the default and production-supported implementation. Runtimes use self.executor.* instead of importing docker.py directly:

# In a runtime:
self.executor.run_cmd(image, command, container_name=name, env=env)
self.executor.stop_cmd(container_name)
self.executor.generate_launch_script(image, container_name, command, ...)
self.executor.container_name(cluster_id, "solo")
self.executor.node_container_name(cluster_id, rank)

Runtimes never construct an executor directly — RuntimePlugin._resolve_executor() delegates to orchestration.executor:resolve_executor(), the single sanctioned entry point for the layered resolution chain. An explicitly-requested executor that is unknown or gated off raises ExecutorUnavailableError rather than falling back to Docker. See Executors.

All remote operations use SSH stdin piping — scripts are generated as Python strings and piped to ssh host bash -s. No files are ever copied to remote hosts for execution.

from sparkrun.orchestration.ssh import run_remote_script
result = run_remote_script(host, script_string, timeout=120, **ssh_kwargs)

All runtimes extend RuntimePlugin (runtimes/base.py):

  • generate_command() — produce the serve command from recipe + overrides
  • resolve_container() — resolve the container image
  • run() / stop() — solo/cluster dispatch (base class handles solo; subclasses implement _run_cluster)
  • cluster_strategy() — "ray" or "native" determines orchestration path
  • get_extra_docker_opts(), get_extra_volumes(), get_extra_env() — runtime-specific hooks

conftest.py provides an isolate_stateful autouse fixture that redirects SAF’s stateful root to tmp_path. Tests never touch ~/.config/sparkrun/. The bootstrap singleton is reset between tests.

  1. Create src/sparkrun/runtimes/my_runtime.py extending RuntimePlugin
  2. Set runtime_name = "my-runtime" and default_image_prefix
  3. Implement generate_command() and optionally resolve_container()
  4. For multi-node: implement _run_cluster() and _stop_cluster()
  5. Register the entry point in pyproject.toml:
    [project.entry-points."sparkrun.runtimes"]
    my_runtime = "sparkrun.runtimes.my_runtime:MyRuntime"
  6. Add tests in tests/test_my_runtime.py
  1. Create src/sparkrun/builders/my_builder.py extending BuilderPlugin
  2. Set builder_name = "my-builder"
  3. Implement prepare_image() — must return the final image name
  4. Register in pyproject.toml under sparkrun.builders
  5. Recipes reference it as builder: my-builder

Versions are tracked in versions.yaml at the repo root:

Terminal window
# Sync versions across pyproject.toml and companion packages
python scripts/update-versions.py
# CI check (verify, don't write)
python scripts/update-versions.py --check
  • Target develop-next branch for all PRs
  • Keep commits atomic — one logical change per commit
  • Run pytest and ruff check before pushing
  • Use --dry-run to verify CLI changes produce correct Docker commands
  • Include test coverage for new functionality