Skip to content

llama.cpp

llama.cpp support in sparkrun provides a lightweight runtime for GGUF quantized models via llama-server.

  • Solo mode with GGUF quantized models
  • Loads models directly from HuggingFace (e.g. Qwen/Qwen3-1.7B-GGUF:Q8_0)
  • Automatic model pre-sync with selective quant file download
  • Lightweight alternative to vLLM/SGLang

sparkrun works with ready-built images including:

We recommend you use the Spark Arena image for sparkrun recipes.

GGUF models use colon syntax to select a quantization variant:

model: Qwen/Qwen3-1.7B-GGUF:Q8_0

sparkrun pre-downloads only the matching quant files and resolves the local cache path so the container doesn’t need to re-download at serve time.

model: Qwen/Qwen3-1.7B-GGUF:Q8_0
runtime: llama-cpp
max_nodes: 1
container: ghcr.io/spark-arena/dgx-llama-cpp:latest
metadata:
description: Qwen3 1.7B (Q8_0 GGUF)
model_params: 1.7B
model_dtype: q8_0
defaults:
port: 8000
host: 0.0.0.0
n_gpu_layers: 99
ctx_size: 8192
command: |
llama-server \
-hf {model} \
--host {host} \
--port {port} \
--n-gpu-layers {n_gpu_layers} \
--ctx-size {ctx_size} \
--flash-attn on \
--jinja \
--no-webui
ParameterDescription
n_gpu_layersNumber of layers offloaded to GPU (99 = all)
ctx_sizeContext window size in tokens
Terminal window
sparkrun run qwen3-1.7b-llama-cpp

llama.cpp has an experimental RPC backend for multi-node inference. Worker nodes run rpc-server and the head node connects via --rpc.