llama.cpp
llama.cpp support in sparkrun provides a lightweight runtime for GGUF quantized models via llama-server.
Features
Section titled “Features”- Solo mode with GGUF quantized models
- Loads models directly from HuggingFace (e.g.
Qwen/Qwen3-1.7B-GGUF:Q8_0) - Automatic model pre-sync with selective quant file download
- Lightweight alternative to vLLM/SGLang
Container images
Section titled “Container images”sparkrun works with ready-built images including:
- Spark Arena Official Image:
ghcr.io/spark-arena/dgx-llama-cpp:latest - scitrera/dgx-spark-llama-cpp
- Build your own
We recommend you use the Spark Arena image for sparkrun recipes.
GGUF model syntax
Section titled “GGUF model syntax”GGUF models use colon syntax to select a quantization variant:
model: Qwen/Qwen3-1.7B-GGUF:Q8_0sparkrun pre-downloads only the matching quant files and resolves the local cache path so the container doesn’t need to re-download at serve time.
Example recipe
Section titled “Example recipe”model: Qwen/Qwen3-1.7B-GGUF:Q8_0runtime: llama-cppmax_nodes: 1container: ghcr.io/spark-arena/dgx-llama-cpp:latest
metadata: description: Qwen3 1.7B (Q8_0 GGUF) model_params: 1.7B model_dtype: q8_0
defaults: port: 8000 host: 0.0.0.0 n_gpu_layers: 99 ctx_size: 8192
command: | llama-server \ -hf {model} \ --host {host} \ --port {port} \ --n-gpu-layers {n_gpu_layers} \ --ctx-size {ctx_size} \ --flash-attn on \ --jinja \ --no-webuiKey parameters
Section titled “Key parameters”| Parameter | Description |
|---|---|
n_gpu_layers | Number of layers offloaded to GPU (99 = all) |
ctx_size | Context window size in tokens |
Running
Section titled “Running”sparkrun run qwen3-1.7b-llama-cppExperimental: Multi-node via RPC
Section titled “Experimental: Multi-node via RPC”llama.cpp has an experimental RPC backend for multi-node inference. Worker nodes run rpc-server and the head node connects via --rpc.