Writing Recipes
Start from an existing recipe
Section titled “Start from an existing recipe”The easiest way to create a new recipe is to copy one that uses the same runtime and modify it. Sparkrun will see any recipes in the current directory and can check and run them.
For more complex recipes involving mods, tunings, or benchmarks, it is recommended to first review the Sparkrun security model so you can be aware of important limitations and caveats, and then set up your own recipe registry.
# See what's availablesparkrun list
# Show a recipe's full configurationsparkrun show qwen3-1.7b-vllmStep by step
Section titled “Step by step”- Set
modelto the HuggingFace model ID you want to serve - Set
container— see Runtimes Overview for maintained images and latest tags. - Set
runtimeto the engine you’re targeting (vllm,sglang,llama-cpp, etc.). While sparkrun can auto-detect the runtime from the command prefix, explicitly setting it is preferred for clarity and to avoid surprises if the heuristics evolve. - Tune
defaults— start conservative withgpu_memory_utilization: 0.5and a lowmax_model_len, then increase after confirming the model loads - Add
metadatawithmodel_paramsandmodel_dtypeso VRAM estimation works - Validate:
sparkrun recipe validate your-recipe.yaml - Test:
sparkrun run your-recipe.yaml --solo --dry-run
Example: vLLM recipe
Section titled “Example: vLLM recipe”Using runtime: vllm here resolves to vllm-distributed by default (see vLLM runtime for details on the two variants). While runtime could be omitted since sparkrun would auto-detect it from the vllm serve command, explicitly setting it is recommended.
model: my-org/my-modelruntime: vllmcontainer: scitrera/dgx-spark-vllm:0.16.0-t5
metadata: description: My custom model maintainer: me@example.com model_params: 7B model_dtype: bf16
defaults: port: 8000 host: 0.0.0.0 tensor_parallel: 1 gpu_memory_utilization: 0.8 max_model_len: 32768
command: | vllm serve {model} \ --max-model-len {max_model_len} \ --gpu-memory-utilization {gpu_memory_utilization} \ -tp {tensor_parallel} \ --host {host} --port {port}- Check the Runtimes Overview to verify you are using the correct or latest container images and runtime options for DGX Spark.
- Use
--dry-runliberally. It shows the exact Docker commands without executing anything. sparkrun show <recipe>displays rendered defaults plus VRAM estimation — use it to sanity-check before launching.- For models that barely fit in VRAM, lower
max_model_lenfirst. KV cache scales with sequence length. tensor_parallelon DGX Spark maps 1:1 to node count.--tp 2means 2 hosts, each contributing 1 GPU.- GGUF recipes should set
max_nodes: 1unless using the experimental RPC multi-node backend.
Sharing recipes
Section titled “Sharing recipes”Create a git repository with your YAML files and a .sparkrun/registry.yaml manifest:
registries: - name: my-team description: My team's recipes for DGX Spark recipes: recipesThen anyone can add your registry:
sparkrun registry add https://github.com/myorg/spark-recipes.gitSee Registries — Creating a registry for the full manifest format, including optional tuning and benchmark paths.