Execution Flow & Internals
This page documents the internal execution flow of sparkrun run — from recipe resolution through container launch. This is reference material for developers building custom runtimes, debugging orchestration issues, or understanding sparkrun’s internals.
Launch lifecycle
Section titled “Launch lifecycle”Every sparkrun run invocation follows a structured 6-phase pipeline, tracked by the LaunchProgress system. Each phase emits progress output at the default verbosity level, with sub-steps visible at -v:
[1/6] Preparing[2/6] Building image — skipped (no builder)[3/6] Distributing resources[4/6] Syncing tuning configs[5/6] Launching vLLM-distributed runtime Step 1/4: Cleanup existing containers Step 2/4: Detect head node IP Step 3/4: Launch head node (rank 0) Step 4/4: Launch worker nodes[6/6] Post-launch hooksPhase 1: Preparing
Section titled “Phase 1: Preparing”Resolves recipe configuration, port availability, container image, and generates a deterministic cluster ID from the recipe + host list + overrides. Also resolves the transfer mode (auto/local/push/delegated) and builds SSH kwargs.
This phase also resolves the recipe’s mods: field. For each mods entry, sparkrun looks up the named mod in the appropriate registry’s mods_subpath (see Manifest fields), syncs the supporting files locally if needed, and injects the corresponding shell commands into the recipe’s pre_exec list. The conversion is builder-agnostic — it runs before Phase 2 so resolution failures surface early, and so any builder (or no builder at all) sees a uniform pre_exec once mods are expanded.
Placement and backend selection
Section titled “Placement and backend selection”Before Phase 2 runs, sparkrun resolves three things from the cluster’s per-host hardware metadata:
- Per-host
BackendBundle—core/launcher.py:resolve_per_host_backends()walkshost_listand callscore/backend_select.py:select_backends(hw)on each host’sHostHardware. The result (accelerator_vendor+CollectiveBackend) is threaded through toruntime.run(..., backends=...)and persisted in job metadata. Hosts that fail to resolve a bundle (unknown vendor, multi-vendor host) are dropped from the map; per-host env then flows throughruntimes/_cluster_ops.py:resolve_comm_env(ctx, comm_env, backends), which falls back to the legacy NCCL generator for those hosts — byte-identical output for NVIDIA. (The oldresolve_ib_envwrapper has been removed;resolve_comm_envis the only path.) - Compatibility check — for runtimes that declare
requires_capability: {"rdma:roce-v2"}(or similar), the central check (runtimes/compatibility.py:check_runtime_host_compatibility) verifies every placed host satisfies the set. Failures raiseIncompatibleHardwareErrorbefore any side effects (no container pull, no model sync, no docker run). - Platform validation —
platforms/resolve_platform(hw)picks the first matchingHardwarePlatformPlugin(DGX Spark before Generic NVIDIA), and itsvalidate_host(hw)returns warning strings. sparkrun logs each warning; it does not raise. Hosts without explicit metadata default to DGX Spark so the check always runs against something sensible.
For multi-rank workloads on heterogeneous-vendor clusters, recipes must also
declare an explicit layout: block (core/layout.py:RecipeLayout) — an
experimental surface, see
Schedulers & Placement.
The auto-packer (core/placement.py:compute_placement) refuses to guess which
ranks land on which accelerator vendor.
Phase 2: Building image
Section titled “Phase 2: Building image”Builders prepare container images before distribution. Most recipes skip this phase — the default docker-pull builder is a no-op since the distribution phase handles image pulling. The eugr builder is the main “real” builder; it’s invoked when a recipe sets builder: eugr, declares build_args, or references an eugr-nightly sentinel image at :latest. Since eugr’s July 2026 move to published images the builder is pull-first — it resolves the sentinel to the authoritative nightly and pulls it, and only runs build-and-copy.sh when build_args request a wheels or custom build. See Builders.
For the full builder catalog, sentinel-image semantics, the defaults.builders.eugr.use_sentinel_image opt-out, build hygiene/observability (per-build logs, phase-aware errors, post-build flashinfer smoke test), and the BuilderPlugin interface, see Builders.
Phase 3: Distributing resources
Section titled “Phase 3: Distributing resources”sparkrun distributes container images and model files to all target hosts before launching. The transfer mode (set per-cluster or per-command) controls how resources flow.
Before this phase runs, sparkrun calls the runtime’s prepare() hook (see Recipe hooks), which can mutate the recipe’s distribution_config — for example, vLLM/SGLang/Atlas all use prepare() to add the speculative-decoding draft model to the distribution list so it ships alongside the primary model.
SSH stdin piping
Section titled “SSH stdin piping”All remote operations use SSH stdin piping — scripts are generated as Python strings and piped to ssh <host> bash -s. No files are ever copied to remote hosts for orchestration purposes.
Container sync
Section titled “Container sync”Container images are distributed via docker save | ssh docker load. sparkrun matches images using both the local docker image ID and RepoDigests (the storage-driver-agnostic content hash) — a host is skipped if either matches, so the same <registry>/<image>@sha256:... is treated as “already present” even when local image IDs differ across drivers.
Model sync
Section titled “Model sync”Models are downloaded locally via HuggingFace Hub’s snapshot_download, then rsynced to target hosts over the high-speed CX-7 network. GGUF models use colon syntax (repo:quant) for selective quant-file download. If HF_TOKEN (or HUGGING_FACE_HUB_TOKEN) is set in the control-machine environment, sparkrun forwards it into the remote sync scripts so gated repos download successfully on workers too.
InfiniBand / NCCL auto-configuration
Section titled “InfiniBand / NCCL auto-configuration”For multi-node workloads, sparkrun detects InfiniBand interfaces on all hosts and computes NCCL environment variables (NCCL_SOCKET_IFNAME, UCX_NET_DEVICES, etc.) automatically.
Phase 4: Syncing tuning configs
Section titled “Phase 4: Syncing tuning configs”Syncs Triton fused MoE kernel tuning configs from registries to the local cache, then distributes them to all target hosts. Skipped when --no-sync-tuning is passed or for delegating runtimes.
Phase 5: Launching runtime
Section titled “Phase 5: Launching runtime”Generates the serve command, clears the page cache (best-effort, requires sudo), configures the executor (rootless mode, user/group mapping, security options), and calls runtime.run(). Each runtime reports sub-steps (cleanup, IP detection, head launch, worker launch, etc.) via the progress tracker.
Phase 6: Post-launch hooks
Section titled “Phase 6: Post-launch hooks”After a successful detached launch, handles post-serve lifecycle:
- Waits for the serve port to become ready
- Runs HTTP health check (if configured)
- Executes
post_execcommands inside the container - Executes
post_commands(registry-defined hooks — requires--trustfor third-party registries) - Handles
stop_after_post(auto-stop after hooks complete, used by benchmarks)
Runtime orchestration internals
Section titled “Runtime orchestration internals”vllm-distributed multi-node
Section titled “vllm-distributed multi-node”vllm-distributed uses vLLM’s built-in multi-node support. Each node runs the full vllm serve command with node-specific arguments:
- Cleanup — remove existing containers for this cluster ID on all hosts
- InfiniBand detection — detect IB interfaces and compute NCCL environment variables
- Head node IP detection — determine the head node’s management IP
- Launch head node (rank 0) with
--nnodes,--node-rank 0, and--master-addr - Wait for master port to become ready
- Launch worker nodes in parallel, each with
--node-rank Nand--headless
No Ray cluster is involved — vLLM coordinates the nodes directly.
vllm-ray multi-node
Section titled “vllm-ray multi-node”vllm-ray forms a Ray cluster first, then runs the serve command on the head node:
- Cleanup — remove existing containers for this cluster ID on all hosts
- InfiniBand detection — detect IB interfaces and compute NCCL environment variables
- Start Ray head node on the first host
- Join workers to the Ray cluster
- Launch vLLM serve on the head container
SGLang multi-node
Section titled “SGLang multi-node”SGLang uses its native distributed backend. sparkrun launches the head node with --dist-init-addr, --nnodes, and --node-rank 0, then launches workers in parallel with their respective --node-rank values.
TensorRT-LLM multi-node (MPI)
Section titled “TensorRT-LLM multi-node (MPI)”TRT-LLM uses a custom MPI-based approach. Instead of installing openssh-server inside containers, sparkrun generates a custom mpirun rsh agent that routes through host-level SSH and docker exec:
- Cleanup — remove existing containers for this cluster ID on all hosts
- InfiniBand detection — detect IB interfaces and compute NCCL environment variables
- IP detection — detect management IPs on all hosts for MPI host list
- Container launch — launch containers with
sleep infinityon all hosts in parallel - Verify — confirm all containers are running
- Write rsh wrapper + config — generate the MPI rsh wrapper script and extra LLM API config YAML, write both into the head container
- Exec mpirun — execute
mpirunon the head container withtrtllm-llmapi-launchwrapping thetrtllm-servecommand
The rsh wrapper script maps each host IP to its container name, then uses ssh <host> docker exec <container> <command> to route MPI commands into the correct worker container.
TRT-LLM environment variables
Section titled “TRT-LLM environment variables”sparkrun sets these environment variables automatically for TRT-LLM clusters:
| Variable | Value | Purpose |
|---|---|---|
OMPI_ALLOW_RUN_AS_ROOT | 1 | Allow mpirun as root |
OMPI_ALLOW_RUN_AS_ROOT_CONFIRM | 1 | Confirm root execution |
NCCL_CUMEM_ENABLE | 0 | Disable NCCL cuMem (compatibility) |
OMPI_MCA_rmaps_ppr_n_pernode | 1 | One process per node |
NCCL-related variables (NCCL_SOCKET_IFNAME, UCX_NET_DEVICES, etc.) and HF_TOKEN are also propagated to all nodes via mpirun -x.
Recipe hooks
Section titled “Recipe hooks”Recipes (and runtimes) can mutate or inject behavior at specific lifecycle points:
prepare()— runtime-side hook called between Phase 2 (Building image) and Phase 3 (Distributing resources). Runtimes use this to mutate the recipe’sdistribution_configso additional models or containers ship to all hosts. All current vLLM, SGLang, and Atlas runtimes use this to pre-sync the speculative-decoding draft model when configured. Not user-facing — implemented in runtime plugins.pre_exec— recipe-level shell commands. Run after container launch but before the serve command starts. Recipemods:entries are expanded intopre_execduring Phase 1.post_exec— recipe-level shell commands. Run after a successful detached launch (Phase 6), once the serve port is ready. Executes inside the head container.post_commands— registry-defined hooks that run on the host (not inside the container) after a successful launch. Requires--trustfor third-party registries.
pre_exec: - "echo 'Preparing environment...'" - "/opt/setup-script.sh"