Skip to content

Execution Flow & Internals

This page documents the internal execution flow of sparkrun run — from recipe resolution through container launch. This is reference material for developers building custom runtimes, debugging orchestration issues, or understanding sparkrun’s internals.

Every sparkrun run invocation follows a structured 6-phase pipeline, tracked by the LaunchProgress system. Each phase emits progress output at the default verbosity level, with sub-steps visible at -v:

[1/6] Preparing
[2/6] Building image — skipped (no builder)
[3/6] Distributing resources
[4/6] Syncing tuning configs
[5/6] Launching vLLM-distributed runtime
Step 1/4: Cleanup existing containers
Step 2/4: Detect head node IP
Step 3/4: Launch head node (rank 0)
Step 4/4: Launch worker nodes
[6/6] Post-launch hooks

Resolves recipe configuration, port availability, container image, and generates a deterministic cluster ID from the recipe + host list + overrides. Also resolves the transfer mode (auto/local/push/delegated) and builds SSH kwargs.

This phase also resolves the recipe’s mods: field. For each mods entry, sparkrun looks up the named mod in the appropriate registry’s mods_subpath (see Manifest fields), syncs the supporting files locally if needed, and injects the corresponding shell commands into the recipe’s pre_exec list. The conversion is builder-agnostic — it runs before Phase 2 so resolution failures surface early, and so any builder (or no builder at all) sees a uniform pre_exec once mods are expanded.

Before Phase 2 runs, sparkrun resolves three things from the cluster’s per-host hardware metadata:

  1. Per-host BackendBundle — core/launcher.py:resolve_per_host_backends() walks host_list and calls core/backend_select.py:select_backends(hw) on each host’s HostHardware. The result (accelerator_vendor + CollectiveBackend) is threaded through to runtime.run(..., backends=...) and persisted in job metadata. Hosts that fail to resolve a bundle (unknown vendor, multi-vendor host) are dropped from the map; per-host env then flows through runtimes/_cluster_ops.py:resolve_comm_env(ctx, comm_env, backends), which falls back to the legacy NCCL generator for those hosts — byte-identical output for NVIDIA. (The old resolve_ib_env wrapper has been removed; resolve_comm_env is the only path.)
  2. Compatibility check — for runtimes that declare requires_capability: {"rdma:roce-v2"} (or similar), the central check (runtimes/compatibility.py:check_runtime_host_compatibility) verifies every placed host satisfies the set. Failures raise IncompatibleHardwareError before any side effects (no container pull, no model sync, no docker run).
  3. Platform validation — platforms/resolve_platform(hw) picks the first matching HardwarePlatformPlugin (DGX Spark before Generic NVIDIA), and its validate_host(hw) returns warning strings. sparkrun logs each warning; it does not raise. Hosts without explicit metadata default to DGX Spark so the check always runs against something sensible.

For multi-rank workloads on heterogeneous-vendor clusters, recipes must also declare an explicit layout: block (core/layout.py:RecipeLayout) — an experimental surface, see Schedulers & Placement. The auto-packer (core/placement.py:compute_placement) refuses to guess which ranks land on which accelerator vendor.

Builders prepare container images before distribution. Most recipes skip this phase — the default docker-pull builder is a no-op since the distribution phase handles image pulling. The eugr builder is the main “real” builder; it’s invoked when a recipe sets builder: eugr, declares build_args, or references an eugr-nightly sentinel image at :latest. Since eugr’s July 2026 move to published images the builder is pull-first — it resolves the sentinel to the authoritative nightly and pulls it, and only runs build-and-copy.sh when build_args request a wheels or custom build. See Builders.

For the full builder catalog, sentinel-image semantics, the defaults.builders.eugr.use_sentinel_image opt-out, build hygiene/observability (per-build logs, phase-aware errors, post-build flashinfer smoke test), and the BuilderPlugin interface, see Builders.

sparkrun distributes container images and model files to all target hosts before launching. The transfer mode (set per-cluster or per-command) controls how resources flow.

Before this phase runs, sparkrun calls the runtime’s prepare() hook (see Recipe hooks), which can mutate the recipe’s distribution_config — for example, vLLM/SGLang/Atlas all use prepare() to add the speculative-decoding draft model to the distribution list so it ships alongside the primary model.

All remote operations use SSH stdin piping — scripts are generated as Python strings and piped to ssh <host> bash -s. No files are ever copied to remote hosts for orchestration purposes.

Container images are distributed via docker save | ssh docker load. sparkrun matches images using both the local docker image ID and RepoDigests (the storage-driver-agnostic content hash) — a host is skipped if either matches, so the same <registry>/<image>@sha256:... is treated as “already present” even when local image IDs differ across drivers.

Models are downloaded locally via HuggingFace Hub’s snapshot_download, then rsynced to target hosts over the high-speed CX-7 network. GGUF models use colon syntax (repo:quant) for selective quant-file download. If HF_TOKEN (or HUGGING_FACE_HUB_TOKEN) is set in the control-machine environment, sparkrun forwards it into the remote sync scripts so gated repos download successfully on workers too.

For multi-node workloads, sparkrun detects InfiniBand interfaces on all hosts and computes NCCL environment variables (NCCL_SOCKET_IFNAME, UCX_NET_DEVICES, etc.) automatically.

Syncs Triton fused MoE kernel tuning configs from registries to the local cache, then distributes them to all target hosts. Skipped when --no-sync-tuning is passed or for delegating runtimes.

Generates the serve command, clears the page cache (best-effort, requires sudo), configures the executor (rootless mode, user/group mapping, security options), and calls runtime.run(). Each runtime reports sub-steps (cleanup, IP detection, head launch, worker launch, etc.) via the progress tracker.

After a successful detached launch, handles post-serve lifecycle:

  1. Waits for the serve port to become ready
  2. Runs HTTP health check (if configured)
  3. Executes post_exec commands inside the container
  4. Executes post_commands (registry-defined hooks — requires --trust for third-party registries)
  5. Handles stop_after_post (auto-stop after hooks complete, used by benchmarks)

vllm-distributed uses vLLM’s built-in multi-node support. Each node runs the full vllm serve command with node-specific arguments:

  1. Cleanup — remove existing containers for this cluster ID on all hosts
  2. InfiniBand detection — detect IB interfaces and compute NCCL environment variables
  3. Head node IP detection — determine the head node’s management IP
  4. Launch head node (rank 0) with --nnodes, --node-rank 0, and --master-addr
  5. Wait for master port to become ready
  6. Launch worker nodes in parallel, each with --node-rank N and --headless

No Ray cluster is involved — vLLM coordinates the nodes directly.

vllm-ray forms a Ray cluster first, then runs the serve command on the head node:

  1. Cleanup — remove existing containers for this cluster ID on all hosts
  2. InfiniBand detection — detect IB interfaces and compute NCCL environment variables
  3. Start Ray head node on the first host
  4. Join workers to the Ray cluster
  5. Launch vLLM serve on the head container

SGLang uses its native distributed backend. sparkrun launches the head node with --dist-init-addr, --nnodes, and --node-rank 0, then launches workers in parallel with their respective --node-rank values.

TRT-LLM uses a custom MPI-based approach. Instead of installing openssh-server inside containers, sparkrun generates a custom mpirun rsh agent that routes through host-level SSH and docker exec:

  1. Cleanup — remove existing containers for this cluster ID on all hosts
  2. InfiniBand detection — detect IB interfaces and compute NCCL environment variables
  3. IP detection — detect management IPs on all hosts for MPI host list
  4. Container launch — launch containers with sleep infinity on all hosts in parallel
  5. Verify — confirm all containers are running
  6. Write rsh wrapper + config — generate the MPI rsh wrapper script and extra LLM API config YAML, write both into the head container
  7. Exec mpirun — execute mpirun on the head container with trtllm-llmapi-launch wrapping the trtllm-serve command

The rsh wrapper script maps each host IP to its container name, then uses ssh <host> docker exec <container> <command> to route MPI commands into the correct worker container.

sparkrun sets these environment variables automatically for TRT-LLM clusters:

VariableValuePurpose
OMPI_ALLOW_RUN_AS_ROOT1Allow mpirun as root
OMPI_ALLOW_RUN_AS_ROOT_CONFIRM1Confirm root execution
NCCL_CUMEM_ENABLE0Disable NCCL cuMem (compatibility)
OMPI_MCA_rmaps_ppr_n_pernode1One process per node

NCCL-related variables (NCCL_SOCKET_IFNAME, UCX_NET_DEVICES, etc.) and HF_TOKEN are also propagated to all nodes via mpirun -x.

Recipes (and runtimes) can mutate or inject behavior at specific lifecycle points:

  • prepare() — runtime-side hook called between Phase 2 (Building image) and Phase 3 (Distributing resources). Runtimes use this to mutate the recipe’s distribution_config so additional models or containers ship to all hosts. All current vLLM, SGLang, and Atlas runtimes use this to pre-sync the speculative-decoding draft model when configured. Not user-facing — implemented in runtime plugins.
  • pre_exec — recipe-level shell commands. Run after container launch but before the serve command starts. Recipe mods: entries are expanded into pre_exec during Phase 1.
  • post_exec — recipe-level shell commands. Run after a successful detached launch (Phase 6), once the serve port is ready. Executes inside the head container.
  • post_commands — registry-defined hooks that run on the host (not inside the container) after a successful launch. Requires --trust for third-party registries.
pre_exec:
- "echo 'Preparing environment...'"
- "/opt/setup-script.sh"