Architecture & Source Layout
This page provides an overview of sparkrun’s source code organization, plugin system, and internal data flows for developers working on or integrating with sparkrun.
Source layout
Section titled “Source layout”src/sparkrun/├── cli/ # Click CLI package — a renderer over api/├── api/ # Console-free public library API (see below)├── core/ # Core data models, bootstrap, and business logic├── runtimes/ # Runtime plugins (vllm, sglang, llama-cpp, trtllm, atlas, …)├── orchestration/ # SSH, Docker, InfiniBand, executors, collectives, telemetry├── transports/ # Cluster connectivity seam — how hosts are reached/prepared├── schedulers/ # Placement schedulers (greedy, occupancy-sparse/-dense)├── platforms/ # Hardware platform registry (DGX Spark, generic NVIDIA)├── models/ # HuggingFace model download, distribution, and VRAM estimation├── containers/ # Container image distribution (docker save/load over SSH)├── tuning/ # Triton fused MoE kernel tuning for SGLang and vLLM├── builders/ # Container image builder plugins (docker-pull, eugr)├── diagnostics/ # Host and run diagnostic collection (NDJSON output)├── proxy/ # Inference gateway (LiteLLM engine + gateway selection)├── benchmarking/ # Benchmark framework plugins and result export├── telemetry/ # Anonymous usage telemetry (opt-out)├── utils/ # Shared helpers (coerce_value, suppress_noisy_loggers, etc.)└── scripts/ # Embedded bash scripts (IB detection, container launch, etc.)Layering
Section titled “Layering”cli/ → api/ → {core/, orchestration/, transports/, …}api/ is console-free — it never writes to stdout/stderr and never calls
sys.exit() — and it does not import sparkrun.cli. The CLI is a thin
renderer over it, so anything the CLI does is reachable from Python. See
Python API.
transports/ may import core/ and orchestration/; orchestration/ never
imports transports/.
Plugin system (SAF)
Section titled “Plugin system (SAF)”sparkrun uses scitrera-app-framework (SAF) for plugin discovery and lifecycle management. Each extension point is discovered by scanning a module for subclasses of a base class:
| Plugin type | Base class | Extension point | Module scanned |
|---|---|---|---|
| Runtimes | RuntimePlugin | sparkrun.runtime | sparkrun.runtimes |
| Builders | BuilderPlugin | sparkrun.builder | sparkrun.builders |
| Benchmarking | BenchmarkingPlugin | sparkrun.benchmarking | sparkrun.benchmarking |
| Executors | Executor | sparkrun.executor | sparkrun.orchestration.executors |
| Schedulers | Scheduler | sparkrun.scheduler | sparkrun.schedulers |
| Transports | Transport | sparkrun.transport | sparkrun.transports |
| Telemetry providers | TelemetryProvider | sparkrun.telemetry | sparkrun.orchestration.telemetry |
Hardware platforms are the exception: platforms/ keeps an ordered in-process
registry rather than going through SAF, because resolution is order-sensitive
(most-specific matches() wins). Transports, which select by exact name, moved
onto SAF; platforms did not.
Bootstrap flow
Section titled “Bootstrap flow”cli/__init__.py → core.bootstrap.init_sparkrun() → SAF init_framework_desktop() → find_types_in_modules() over each module above → register_plugin() for each discovered plugin → load_external_plugins(v) # out-of-tree, gated off by defaultSchedulers and transports skip base classes with a blank scheduler_name /
transport_name. A plugin can gate itself behind a
feature flag by declaring required_feature_flag;
SAF then only exposes it when the flag resolves on.
Out-of-tree plugins load last, from plugins.paths in config.yaml — see
External Plugins.
Runtime plugins
Section titled “Runtime plugins”All runtimes extend RuntimePlugin (in runtimes/base.py), which itself extends SAF’s Plugin class. The base class provides solo-mode orchestration; runtimes override run()/stop()/follow_logs() for multi-node support.
Key methods:
generate_command()— produce the serve command from recipe defaultsresolve_container()— determine the container image to usecluster_strategy()— return"ray","native", or"native/rpc"to select orchestration path
Multiplatform architecture
Section titled “Multiplatform architecture”As of 0.3.0 sparkrun has a layered hardware / platform model that future AMD and Intel Gaudi support builds against. Today NVIDIA DGX Spark is the only fully-wired platform; RCCL (AMD) and HCCL (Intel) ship as scaffolds.
launch_inference() | +-----------------+-----------------+ | | v v resolve_per_host_backends resolve_platform (HostHardware → BackendBundle) (HostHardware → HardwarePlatformPlugin) | | v v select_backends(HostHardware) validate_host(hw) (warnings) - accelerator_vendor_for default_image(runtime) - collectives.get_backend(vendor) | v BackendBundle threaded through to runtime.run(..., backends=...) - runtime calls _cluster_ops.resolve_comm_env(ctx, comm_env, backends)Seams contributors plug into:
| Layer | File |
|---|---|
| Per-host hardware model | core/hardware.py (AcceleratorSpec, HostHardware) |
| Combined accelerator + IB probe | core/hardware_probe.py (probe_host, probe_hosts) |
| Per-host backend selection | core/backend_select.py (select_backends) |
| Rank-to-host placement | core/placement.py, core/layout.py |
| Collective backend abstraction | orchestration/collectives/ (nccl, rccl, hccl) |
| Hardware platform plugins | platforms/ (dgx_spark, nvidia_generic) |
| Executor abstraction | orchestration/executor.py, orchestration/executors/ |
For the public-facing summary see Multiplatform;
the contributor-facing deep dive lives in
docs/MULTIPLATFORM.md.
Core modules
Section titled “Core modules”| Module | Purpose |
|---|---|
core/bootstrap.py | SAF plugin initialization and discovery across every extension point |
core/config.py | SparkrunConfig — reads ~/.config/sparkrun/config.yaml, cache dir resolution |
core/registry.py | RegistryManager — git-based recipe registry system |
core/recipe.py | Recipe loading, validation, v1→v2 migration, config chain via SAF Variables |
core/cluster_manager.py | ClusterManager — named cluster CRUD (YAML files) |
core/hosts.py | Host resolution priority chain (CLI → file → cluster → default) |
core/pending_ops.py | PID-based lock files for in-progress operations |
core/benchmark_profiles.py | Benchmark profile discovery and rendering across registries |
core/launcher.py | launch_inference(), per-host backend resolution, recipe trust resolution |
core/scheduler.py | Scheduler ABC, EXT_SCHEDULER, selector resolution chain |
core/cluster_status.py | ClusterStatus — the data-only occupancy snapshot shape |
core/features.py | Channel-aware feature flags |
core/hardware.py, core/hardware_probe.py | AcceleratorSpec / HostHardware + the combined accelerator/IB probe |
core/backend_select.py | select_backends(HostHardware) -> BackendBundle |
core/placement.py, core/layout.py | Rank → (host, local-GPU) placement and the recipe layout: model |
core/limits.py | Usable-memory cap resolution for scheduling and fit |
core/external_plugins.py | Out-of-tree plugin loading from plugins.paths |
The api layer
Section titled “The api layer”api/ is the contract non-CLI callers depend on. Two conventions matter:
- Errors are typed. Everything derives from
api.SparkrunError; internal exceptions (InfeasibleScheduleError,LayoutConflictError, …) are translated at the api boundary. sctxthreads a session. Every entry point takes an optionalSparkrunContextso a chain of calls shares config, registry manager, and cluster manager instead of rebuilding them.
Larger surfaces get their own sub-namespace (api.proxy, api.setup,
api.tailscale), each following the same shape: _ops.py for the console-free
logic, _errors.py for the typed exceptions, __init__.py re-exporting both.
Status discovery
Section titled “Status discovery”All “what’s running where?” questions flow through one source, in two tiers:
| Function | Tier |
|---|---|
api.status() | Lean occupancy snapshot (ClusterStatus) — consumed by schedulers, placement, proxy discovery, teardown |
api.status_report() | Display tier — classifies the snapshot into groups/solo/idle/pending and enriches it with cached job metadata |
Underneath, orchestration/executor.py:query_status_for_cluster() sweeps
every enabled executor sharing the cluster’s substrate and merges the
results. Executor.status_scope (default "host") names that substrate:
executors sharing a scope inspect disjoint state on the same hosts — Docker
containers versus native pidfiles — so both must be asked. A single failing
executor is skipped rather than breaking the report.
Telemetry is the second axis, abstracted the same way: one TelemetryProvider
per scope, composed with the occupancy poll by api.live_monitor into
MonitorFrame snapshots that drive the cluster monitor TUI.
Recipe resolution chain
Section titled “Recipe resolution chain”When you run sparkrun run my-recipe, resolution walks these in order:
@spark-arena/<uuid>shortcut — expands to a Spark Arena URL- A literal HTTP/HTTPS URL — fetched and cached (https-only, host-allowlisted)
@registry/recipe-name— scoped lookup, skipping the search chain- An exact or relative file path
- The current working directory —
.yaml/.ymlfiles that parse as recipes - Registry search — flat name lookup, then a recursive scan
Two rules govern step 6: flat beats nested within a registry (but never
suppresses another registry’s scan), and .yaml beats a same-stem .yml in
the same directory. So an “ambiguous name” error always means genuinely
distinct recipes. The same machinery backs benchmark profiles, tuning configs,
and mods, so listing and lookup can never disagree — see
RegistryAsset / find_asset_in_registries / iter_asset_files.
The catalog side has a single entry point too: api.search_recipes(), which
sparkrun list and sparkrun search both render. They differ only in
unique_names — list shows one row per unqualified name, search shows
every copy.
Registry system internals
Section titled “Registry system internals”The RegistryManager tracks recipe collections from remote git repos using sparse checkouts.
Default registry initialization
Section titled “Default registry initialization”On first run (no registries.yaml):
_load_registries()→ no file →_default_registries()- Clones each URL in
BOOTSTRAP_REGISTRY_URLSand reads its.sparkrun/registry.yamlmanifest - Merges manifest entries with
FALLBACK_DEFAULT_REGISTRIES(manifest entries take priority by name) - Saves combined list to
registries.yaml - If all manifest URLs fail, pure fallback entries are returned (offline safety net)
BOOTSTRAP_REGISTRY_URLS governs manifest discovery only — it does not
control trust. Trust is a per-entry trusted field stored locally in the
user’s registries.yaml, seeded from FALLBACK_DEFAULT_REGISTRIES (every
built-in default ships trusted) and editable via sparkrun registry trust/untrust. A repository’s own manifest cannot grant itself trust.
Manifest format
Section titled “Manifest format”The .sparkrun/registry.yaml file in a git repo supports both canonical and short keys:
| Canonical key | Short key | Purpose |
|---|---|---|
subpath | recipes | Recipe directory within the repo |
tuning_subpath | tuning | Tuning configs directory |
benchmark_subpath | benchmarks | Benchmark profiles directory |
Shared clones
Section titled “Shared clones”When multiple registries point to the same git URL, a single shared clone is used with per-registry symlinks. Sparse checkout paths are the union of all subpaths for that URL.
Reserved name prefixes
Section titled “Reserved name prefixes”Names starting with reserved prefixes (sparkrun, official, arena, etc.) can only be used by repos hosted under allowed GitHub organizations. Enforced via validate_registry_name().
Config & state paths
Section titled “Config & state paths”| Path | Purpose |
|---|---|
~/.config/sparkrun/config.yaml | User configuration, including features: and plugins: |
~/.config/sparkrun/clusters/*.yaml | Named cluster definitions |
~/.config/sparkrun/registries.yaml | Recipe registry list, including per-registry trust |
~/.config/sparkrun/proxy.yaml | Gateway settings and model aliases |
~/.cache/sparkrun/registries/ | Git-cloned recipe registries |
~/.cache/sparkrun/jobs/ | Job metadata (cluster_id → recipe mapping) |
~/.cache/sparkrun/pending/ | PID lock files for in-progress operations |
~/.cache/sparkrun/benchmarks/ | Resumable benchmark run state |
~/.cache/sparkrun/proxy/ | Generated gateway config, process state, logs (0600) |
~/.cache/sparkrun/local/ | PID and log files for local-executor workloads |
~/.cache/huggingface/ | HuggingFace model cache (mounted into containers) |
Orchestration layer
Section titled “Orchestration layer”All remote operations use SSH stdin piping — scripts are generated as Python strings and piped to ssh <host> bash -s. No files are ever copied to remote hosts.
| Module | Purpose |
|---|---|
orchestration/ssh.py | RemoteResult, build_ssh_cmd(), run_remote_script(), parallel execution |
orchestration/sudo.py | run_with_sudo_fallback() — tries non-interactive sudo, then password fallback |
orchestration/docker.py | Pure command-string generators (docker_run_cmd, docker_exec_cmd, etc.) |
orchestration/distribution.py | High-level resource distribution: IB detection, container and model syncing |
orchestration/infiniband.py | IB detection script generation, NCCL env var computation |
orchestration/networking.py | ConnectX-7 NIC detection, IP assignment, CX7 configuration |
orchestration/primitives.py | Higher-level composition: build_ssh_kwargs(), build_volumes(), merge_env() |
orchestration/job_metadata.py | Persistent job metadata stored in ~/.cache/sparkrun/jobs/ |
orchestration/executor.py | Public executor facade — resolve_executor(), query_status_for_cluster() |
orchestration/executors/ | Executor plugins: docker.py (default), local.py (experimental) |
orchestration/collectives/ | CollectiveBackend ABC + nccl (default), rccl/hccl scaffolds |
orchestration/telemetry/ | TelemetryProvider per status scope (host metrics for cluster monitor) |
orchestration/logs.py | Where a runtime’s logs actually live, per executor |
orchestration/hooks.py | pre_exec / post_exec / post_commands runners, with trust gating |
SSH fan-out is bounded by SparkrunConfig.max_parallel_ssh (config key
ssh.max_parallel_ssh, default 20), and launch-path calls carry execution
timeouts so one host that connects and then hangs can’t stall a batch.
Key dependencies
Section titled “Key dependencies”| Package | Purpose |
|---|---|
scitrera-app-framework (SAF) | Plugin system, lifecycle, variables/config management |
vpd | YAML reading and command template placeholder substitution |
click | CLI framework |
huggingface_hub | Model downloading (snapshot_download) |
pyyaml | YAML parsing for recipes, clusters, registries |