Skip to content

Architecture & Source Layout

This page provides an overview of sparkrun’s source code organization, plugin system, and internal data flows for developers working on or integrating with sparkrun.

src/sparkrun/
├── cli/ # Click CLI package — a renderer over api/
├── api/ # Console-free public library API (see below)
├── core/ # Core data models, bootstrap, and business logic
├── runtimes/ # Runtime plugins (vllm, sglang, llama-cpp, trtllm, atlas, …)
├── orchestration/ # SSH, Docker, InfiniBand, executors, collectives, telemetry
├── transports/ # Cluster connectivity seam — how hosts are reached/prepared
├── schedulers/ # Placement schedulers (greedy, occupancy-sparse/-dense)
├── platforms/ # Hardware platform registry (DGX Spark, generic NVIDIA)
├── models/ # HuggingFace model download, distribution, and VRAM estimation
├── containers/ # Container image distribution (docker save/load over SSH)
├── tuning/ # Triton fused MoE kernel tuning for SGLang and vLLM
├── builders/ # Container image builder plugins (docker-pull, eugr)
├── diagnostics/ # Host and run diagnostic collection (NDJSON output)
├── proxy/ # Inference gateway (LiteLLM engine + gateway selection)
├── benchmarking/ # Benchmark framework plugins and result export
├── telemetry/ # Anonymous usage telemetry (opt-out)
├── utils/ # Shared helpers (coerce_value, suppress_noisy_loggers, etc.)
└── scripts/ # Embedded bash scripts (IB detection, container launch, etc.)
cli/ → api/ → {core/, orchestration/, transports/, …}

api/ is console-free — it never writes to stdout/stderr and never calls sys.exit() — and it does not import sparkrun.cli. The CLI is a thin renderer over it, so anything the CLI does is reachable from Python. See Python API.

transports/ may import core/ and orchestration/; orchestration/ never imports transports/.

sparkrun uses scitrera-app-framework (SAF) for plugin discovery and lifecycle management. Each extension point is discovered by scanning a module for subclasses of a base class:

Plugin typeBase classExtension pointModule scanned
RuntimesRuntimePluginsparkrun.runtimesparkrun.runtimes
BuildersBuilderPluginsparkrun.buildersparkrun.builders
BenchmarkingBenchmarkingPluginsparkrun.benchmarkingsparkrun.benchmarking
ExecutorsExecutorsparkrun.executorsparkrun.orchestration.executors
SchedulersSchedulersparkrun.schedulersparkrun.schedulers
TransportsTransportsparkrun.transportsparkrun.transports
Telemetry providersTelemetryProvidersparkrun.telemetrysparkrun.orchestration.telemetry

Hardware platforms are the exception: platforms/ keeps an ordered in-process registry rather than going through SAF, because resolution is order-sensitive (most-specific matches() wins). Transports, which select by exact name, moved onto SAF; platforms did not.

cli/__init__.py → core.bootstrap.init_sparkrun()
→ SAF init_framework_desktop()
→ find_types_in_modules() over each module above
→ register_plugin() for each discovered plugin
→ load_external_plugins(v) # out-of-tree, gated off by default

Schedulers and transports skip base classes with a blank scheduler_name / transport_name. A plugin can gate itself behind a feature flag by declaring required_feature_flag; SAF then only exposes it when the flag resolves on.

Out-of-tree plugins load last, from plugins.paths in config.yaml — see External Plugins.

All runtimes extend RuntimePlugin (in runtimes/base.py), which itself extends SAF’s Plugin class. The base class provides solo-mode orchestration; runtimes override run()/stop()/follow_logs() for multi-node support.

Key methods:

  • generate_command() — produce the serve command from recipe defaults
  • resolve_container() — determine the container image to use
  • cluster_strategy() — return "ray", "native", or "native/rpc" to select orchestration path

As of 0.3.0 sparkrun has a layered hardware / platform model that future AMD and Intel Gaudi support builds against. Today NVIDIA DGX Spark is the only fully-wired platform; RCCL (AMD) and HCCL (Intel) ship as scaffolds.

launch_inference()
|
+-----------------+-----------------+
| |
v v
resolve_per_host_backends resolve_platform
(HostHardware → BackendBundle) (HostHardware → HardwarePlatformPlugin)
| |
v v
select_backends(HostHardware) validate_host(hw) (warnings)
- accelerator_vendor_for default_image(runtime)
- collectives.get_backend(vendor)
|
v
BackendBundle threaded through to runtime.run(..., backends=...)
- runtime calls _cluster_ops.resolve_comm_env(ctx, comm_env, backends)

Seams contributors plug into:

LayerFile
Per-host hardware modelcore/hardware.py (AcceleratorSpec, HostHardware)
Combined accelerator + IB probecore/hardware_probe.py (probe_host, probe_hosts)
Per-host backend selectioncore/backend_select.py (select_backends)
Rank-to-host placementcore/placement.py, core/layout.py
Collective backend abstractionorchestration/collectives/ (nccl, rccl, hccl)
Hardware platform pluginsplatforms/ (dgx_spark, nvidia_generic)
Executor abstractionorchestration/executor.py, orchestration/executors/

For the public-facing summary see Multiplatform; the contributor-facing deep dive lives in docs/MULTIPLATFORM.md.

ModulePurpose
core/bootstrap.pySAF plugin initialization and discovery across every extension point
core/config.pySparkrunConfig — reads ~/.config/sparkrun/config.yaml, cache dir resolution
core/registry.pyRegistryManager — git-based recipe registry system
core/recipe.pyRecipe loading, validation, v1→v2 migration, config chain via SAF Variables
core/cluster_manager.pyClusterManager — named cluster CRUD (YAML files)
core/hosts.pyHost resolution priority chain (CLI → file → cluster → default)
core/pending_ops.pyPID-based lock files for in-progress operations
core/benchmark_profiles.pyBenchmark profile discovery and rendering across registries
core/launcher.pylaunch_inference(), per-host backend resolution, recipe trust resolution
core/scheduler.pyScheduler ABC, EXT_SCHEDULER, selector resolution chain
core/cluster_status.pyClusterStatus — the data-only occupancy snapshot shape
core/features.pyChannel-aware feature flags
core/hardware.py, core/hardware_probe.pyAcceleratorSpec / HostHardware + the combined accelerator/IB probe
core/backend_select.pyselect_backends(HostHardware) -> BackendBundle
core/placement.py, core/layout.pyRank → (host, local-GPU) placement and the recipe layout: model
core/limits.pyUsable-memory cap resolution for scheduling and fit
core/external_plugins.pyOut-of-tree plugin loading from plugins.paths

api/ is the contract non-CLI callers depend on. Two conventions matter:

  • Errors are typed. Everything derives from api.SparkrunError; internal exceptions (InfeasibleScheduleError, LayoutConflictError, …) are translated at the api boundary.
  • sctx threads a session. Every entry point takes an optional SparkrunContext so a chain of calls shares config, registry manager, and cluster manager instead of rebuilding them.

Larger surfaces get their own sub-namespace (api.proxy, api.setup, api.tailscale), each following the same shape: _ops.py for the console-free logic, _errors.py for the typed exceptions, __init__.py re-exporting both.

All “what’s running where?” questions flow through one source, in two tiers:

FunctionTier
api.status()Lean occupancy snapshot (ClusterStatus) — consumed by schedulers, placement, proxy discovery, teardown
api.status_report()Display tier — classifies the snapshot into groups/solo/idle/pending and enriches it with cached job metadata

Underneath, orchestration/executor.py:query_status_for_cluster() sweeps every enabled executor sharing the cluster’s substrate and merges the results. Executor.status_scope (default "host") names that substrate: executors sharing a scope inspect disjoint state on the same hosts — Docker containers versus native pidfiles — so both must be asked. A single failing executor is skipped rather than breaking the report.

Telemetry is the second axis, abstracted the same way: one TelemetryProvider per scope, composed with the occupancy poll by api.live_monitor into MonitorFrame snapshots that drive the cluster monitor TUI.

When you run sparkrun run my-recipe, resolution walks these in order:

  1. @spark-arena/<uuid> shortcut — expands to a Spark Arena URL
  2. A literal HTTP/HTTPS URL — fetched and cached (https-only, host-allowlisted)
  3. @registry/recipe-name — scoped lookup, skipping the search chain
  4. An exact or relative file path
  5. The current working directory — .yaml/.yml files that parse as recipes
  6. Registry search — flat name lookup, then a recursive scan

Two rules govern step 6: flat beats nested within a registry (but never suppresses another registry’s scan), and .yaml beats a same-stem .yml in the same directory. So an “ambiguous name” error always means genuinely distinct recipes. The same machinery backs benchmark profiles, tuning configs, and mods, so listing and lookup can never disagree — see RegistryAsset / find_asset_in_registries / iter_asset_files.

The catalog side has a single entry point too: api.search_recipes(), which sparkrun list and sparkrun search both render. They differ only in unique_names — list shows one row per unqualified name, search shows every copy.

The RegistryManager tracks recipe collections from remote git repos using sparse checkouts.

On first run (no registries.yaml):

  1. _load_registries() → no file → _default_registries()
  2. Clones each URL in BOOTSTRAP_REGISTRY_URLS and reads its .sparkrun/registry.yaml manifest
  3. Merges manifest entries with FALLBACK_DEFAULT_REGISTRIES (manifest entries take priority by name)
  4. Saves combined list to registries.yaml
  5. If all manifest URLs fail, pure fallback entries are returned (offline safety net)

BOOTSTRAP_REGISTRY_URLS governs manifest discovery only — it does not control trust. Trust is a per-entry trusted field stored locally in the user’s registries.yaml, seeded from FALLBACK_DEFAULT_REGISTRIES (every built-in default ships trusted) and editable via sparkrun registry trust/untrust. A repository’s own manifest cannot grant itself trust.

The .sparkrun/registry.yaml file in a git repo supports both canonical and short keys:

Canonical keyShort keyPurpose
subpathrecipesRecipe directory within the repo
tuning_subpathtuningTuning configs directory
benchmark_subpathbenchmarksBenchmark profiles directory

When multiple registries point to the same git URL, a single shared clone is used with per-registry symlinks. Sparse checkout paths are the union of all subpaths for that URL.

Names starting with reserved prefixes (sparkrun, official, arena, etc.) can only be used by repos hosted under allowed GitHub organizations. Enforced via validate_registry_name().

PathPurpose
~/.config/sparkrun/config.yamlUser configuration, including features: and plugins:
~/.config/sparkrun/clusters/*.yamlNamed cluster definitions
~/.config/sparkrun/registries.yamlRecipe registry list, including per-registry trust
~/.config/sparkrun/proxy.yamlGateway settings and model aliases
~/.cache/sparkrun/registries/Git-cloned recipe registries
~/.cache/sparkrun/jobs/Job metadata (cluster_id → recipe mapping)
~/.cache/sparkrun/pending/PID lock files for in-progress operations
~/.cache/sparkrun/benchmarks/Resumable benchmark run state
~/.cache/sparkrun/proxy/Generated gateway config, process state, logs (0600)
~/.cache/sparkrun/local/PID and log files for local-executor workloads
~/.cache/huggingface/HuggingFace model cache (mounted into containers)

All remote operations use SSH stdin piping — scripts are generated as Python strings and piped to ssh <host> bash -s. No files are ever copied to remote hosts.

ModulePurpose
orchestration/ssh.pyRemoteResult, build_ssh_cmd(), run_remote_script(), parallel execution
orchestration/sudo.pyrun_with_sudo_fallback() — tries non-interactive sudo, then password fallback
orchestration/docker.pyPure command-string generators (docker_run_cmd, docker_exec_cmd, etc.)
orchestration/distribution.pyHigh-level resource distribution: IB detection, container and model syncing
orchestration/infiniband.pyIB detection script generation, NCCL env var computation
orchestration/networking.pyConnectX-7 NIC detection, IP assignment, CX7 configuration
orchestration/primitives.pyHigher-level composition: build_ssh_kwargs(), build_volumes(), merge_env()
orchestration/job_metadata.pyPersistent job metadata stored in ~/.cache/sparkrun/jobs/
orchestration/executor.pyPublic executor facade — resolve_executor(), query_status_for_cluster()
orchestration/executors/Executor plugins: docker.py (default), local.py (experimental)
orchestration/collectives/CollectiveBackend ABC + nccl (default), rccl/hccl scaffolds
orchestration/telemetry/TelemetryProvider per status scope (host metrics for cluster monitor)
orchestration/logs.pyWhere a runtime’s logs actually live, per executor
orchestration/hooks.pypre_exec / post_exec / post_commands runners, with trust gating

SSH fan-out is bounded by SparkrunConfig.max_parallel_ssh (config key ssh.max_parallel_ssh, default 20), and launch-path calls carry execution timeouts so one host that connects and then hangs can’t stall a batch.

PackagePurpose
scitrera-app-framework (SAF)Plugin system, lifecycle, variables/config management
vpdYAML reading and command template placeholder substitution
clickCLI framework
huggingface_hubModel downloading (snapshot_download)
pyyamlYAML parsing for recipes, clusters, registries