Skip to content

Writing Recipes

The easiest way to create a new recipe is to copy one that uses the same runtime and modify it. Sparkrun will see any recipes in the current directory and can check and run them.

For more complex recipes involving mods, tunings, or benchmarks, it is recommended to first review the Sparkrun security model so you can be aware of important limitations and caveats, and then set up your own recipe registry.

Terminal window
# See what's available
sparkrun list
# Show a recipe's full configuration
sparkrun show qwen3-1.7b-vllm
  1. Set model to the HuggingFace model ID you want to serve
  2. Set container — see Runtimes Overview for maintained images and latest tags.
  3. Set runtime to the engine you’re targeting (vllm, sglang, llama-cpp, etc.). While sparkrun can auto-detect the runtime from the command prefix, explicitly setting it is preferred for clarity and to avoid surprises if the heuristics evolve.
  4. Tune defaults — start conservative with gpu_memory_utilization: 0.5 and a low max_model_len, then increase after confirming the model loads
  5. Add metadata with model_params and model_dtype so VRAM estimation works
  6. Validate: sparkrun recipe validate your-recipe.yaml
  7. Test: sparkrun run your-recipe.yaml --solo --dry-run

Using runtime: vllm here resolves to vllm-distributed by default (see vLLM runtime for details on the two variants). While runtime could be omitted since sparkrun would auto-detect it from the vllm serve command, explicitly setting it is recommended.

model: my-org/my-model
runtime: vllm
container: scitrera/dgx-spark-vllm:0.16.0-t5
metadata:
description: My custom model
maintainer: me@example.com
model_params: 7B
model_dtype: bf16
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
gpu_memory_utilization: 0.8
max_model_len: 32768
command: |
vllm serve {model} \
--max-model-len {max_model_len} \
--gpu-memory-utilization {gpu_memory_utilization} \
-tp {tensor_parallel} \
--host {host} --port {port}
  • Check the Runtimes Overview to verify you are using the correct or latest container images and runtime options for DGX Spark.
  • Use --dry-run liberally. It shows the exact Docker commands without executing anything.
  • sparkrun show <recipe> displays rendered defaults plus VRAM estimation — use it to sanity-check before launching.
  • For models that barely fit in VRAM, lower max_model_len first. KV cache scales with sequence length.
  • tensor_parallel on DGX Spark maps 1:1 to node count. --tp 2 means 2 hosts, each contributing 1 GPU.
  • GGUF recipes should set max_nodes: 1 unless using the experimental RPC multi-node backend.

Create a git repository with your YAML files and a .sparkrun/registry.yaml manifest:

.sparkrun/registry.yaml
registries:
- name: my-team
description: My team's recipes for DGX Spark
recipes: recipes

Then anyone can add your registry:

Terminal window
sparkrun registry add https://github.com/myorg/spark-recipes.git

See Registries — Creating a registry for the full manifest format, including optional tuning and benchmark paths.