Skip to content

Atlas

Atlas is a pure-Rust LLM inference server with a native NCCL-based distribution layer. sparkrun’s atlas runtime launches spark serve and orchestrates the container lifecycle, NCCL bootstrap, and OpenAI-compatible HTTP endpoint.

  • Pure-Rust inference server with low overhead
  • Native NCCL bootstrap via --rank, --world-size, --master-addr, and --master-port
  • Tensor parallelism (--tp-size) and expert parallelism (--ep-size) — both can share the same physical ranks (overlapping groups) or form an orthogonal mesh
  • Optional speculative decoding (--speculative, --ngram-speculative, --self-speculative), DFlash attention, and prefix caching
  • --high-speed-swap path that uses io_uring for fast KV/cache spill (containers run with seccomp=unconfined and IPC_LOCK/SYS_NICE capabilities)

Atlas derives its world size from the two parallelism dimensions, so you set tensor_parallel and ep_size and sparkrun maps the resulting ranks onto physical hosts:

  • world_size == tensor_parallel × ep_size — orthogonal mesh
  • world_size == tensor_parallel == ep_size — overlapping groups

Atlas requires a GB10 host: the runtime declares a gb10 capability requirement, and a launch placed on a host without it fails the compatibility check before any side effects.

The Atlas maintainers publish an official sparkrun recipe registry at Avarok-Cybersecurity/atlas-recipes. Treat this as the authoritative source for official Atlas recipes — recipe formats, supported flags, and container tags there are kept in sync with upstream Atlas releases.

The @atlas registry is bundled as a default registry in sparkrun v0.2.31 and later, so it’s discoverable out of the box without manual sparkrun registry add:

Terminal window
# List recipes from the official Atlas registry
sparkrun recipe list --registry atlas
# Run a recipe from the @atlas registry
sparkrun run @atlas/<recipe-name>

If you’re on an older sparkrun release, you’ll need to update to use the atlas runtime and registry.

Terminal window
sparkrun run @atlas/qwen3.5-35b-a3b-nvfp4

Default image prefix: avarok/atlas-gb10

container: avarok/atlas-gb10:latest

See the Atlas GitHub repository for image tags and release notes.

This is the official qwen3.5-35b-a3b-nvfp4 recipe from the @atlas registry:

recipe_version: "2"
model: Sehyo/Qwen3.5-35B-A3B-NVFP4
runtime: atlas
container: avarok/atlas-gb10:latest
max_nodes: 1
metadata:
description: |
Qwen3.5-35B (A3B MoE) NVFP4 + MTP — Atlas's fastest catalogued model
on a single GB10 (~133 tok/s decode at concurrency=1, ISL=128, OSL=128
per the avarok/atlas-gb10 QUICKSTART).
maintainer: avarok
category: agent
model_params: 35B
model_dtype: nvfp4
quantization: nvfp4
kv_dtype: nvfp4
defaults:
port: 8888
host: 0.0.0.0
max_model_len: 8192
kv_cache_dtype: nvfp4
gpu_memory_utilization: 0.88
scheduling_policy: slai
speculative: true
mtp_quantization: nvfp4
enable_prefix_caching: true

When no command template is provided, sparkrun generates spark serve {model} with flags derived from defaults via the Atlas flag map.

These recipe defaults keys map to spark serve CLI flags:

KeyFlagDescription
port--portServing port
host--hostBind address
tensor_parallel--tp-sizeTensor parallelism degree (distribution inference coming soon)
gpu_memory_utilization--gpu-memory-utilizationFraction of GPU memory for the engine
max_model_len--max-seq-lenMaximum sequence length
max_num_seqs--max-num-seqsMaximum concurrent sequences
max_num_batched_tokens--max-prefill-tokensMaximum prefill token budget per step
served_model_name--model-namePublic model name exposed by the OpenAI API
kv_cache_dtype--kv-cache-dtypeKV cache datatype
KeyFlagDescription
ep_size--ep-sizeExpert parallelism degree (MoE models)
max_batch_size--max-batch-sizeMaximum batch size
block_size--block-sizeKV cache block size
kv_high_precision_layers--kv-high-precision-layersLayers kept in higher KV precision
tool_call_parser--tool-call-parserTool-call parser name
tool_max_tokens--tool-max-tokensMaximum tokens per tool call
scheduling_policy--scheduling-policyRequest scheduling policy
tbt_deadline_ms--tbt-deadline-msToken-by-token latency deadline (ms)
max_prefill_tokens--max-prefill-tokensMaximum prefill tokens per step
oom_guard_mb--oom-guard-mbHeadroom reserved to avoid OOM (MB)
ssm_cache_slots--ssm-cache-slotsSSM cache slot count
ssm_checkpoint_interval--ssm-checkpoint-intervalSSM checkpoint cadence
mtp_quantization--mtp-quantizationMulti-token-prediction quantization
mtp_vocab--mtp-vocabMulti-token-prediction vocabulary
num_drafts--num-draftsNumber of speculative draft tokens
draft_model--draft-modelDraft model for speculative decoding
dflash_gamma--dflash-gammaDFlash gamma parameter
dflash_window_size--dflash-window-sizeDFlash attention window size
max_thinking_budget--max-thinking-budgetMaximum reasoning tokens
model_from_path--model-from-pathLoad model from a local path
cache_dir--cache-dirOverride cache directory
gpu_ordinal--gpu-ordinalGPU ordinal override

Present-when-truthy flags (no value emitted):

KeyFlag
enable_prefix_caching--enable-prefix-caching
speculative--speculative
self_speculative--self-speculative
ngram_speculative--ngram-speculative
dflash--dflash
disable_thinking--disable-thinking
high_speed_swap--high-speed-swap
require_auth--require-auth

When a recipe sets draft_model in defaults, sparkrun automatically pre-syncs that model to all target hosts alongside the primary model — the same way it pre-syncs draft models for vLLM and SGLang speculative recipes. This happens during the distribution phase so the container never has to fetch the draft from HuggingFace at launch time (which would fail under sparkrun’s offline-hub default).

defaults:
speculative: true
draft_model: avarok/qwen3.5-35b-a3b-nvfp4-draft
num_drafts: 2

This works with any of the speculative modes (speculative, self_speculative, ngram_speculative, dflash); the only requirement is that the draft model name is set in defaults.draft_model.

Atlas derives world_size from tensor_parallel and ep_size:

  • world_size == tp_size * ep_size — orthogonal mesh (TP and EP on independent ranks)
  • world_size == tp_size == ep_size — overlapping groups (both groups share the same physical ranks)

sparkrun maps the resulting world size to physical hosts automatically. While the multi-node path is gated, set tensor_parallel: 1 and omit ep_size (or set it to 1).

Atlas containers run with extra Docker options required by its NCCL and io_uring paths:

  • --cap-add=IPC_LOCK — required for ibv_reg_mr (RDMA memory registration)
  • --cap-add=SYS_NICE — required by the SQPOLL kernel thread used with --high-speed-swap
  • --security-opt seccomp=unconfined — Docker’s default seccomp profile blocks the io_uring_* syscalls used by --high-speed-swap