Your First Model
This tutorial walks you through launching your first LLM on a DGX Spark using sparkrun — from browsing available recipes to testing the API endpoint and stopping the workload.
Prerequisites
Section titled “Prerequisites”- sparkrun installed (Installation guide)
- SSH access to a DGX Spark (password or key-based)
- Docker running on the target Spark
If you haven’t set up SSH yet, see SSH Setup.
1. Browse available recipes
Section titled “1. Browse available recipes”Start by seeing what models are available:
sparkrun listThis searches all configured registries and displays available recipes. Each recipe name follows a pattern: <model>-<runtime> (e.g., qwen3-1.7b-vllm runs the Qwen3 1.7B model on the vLLM runtime).
You can filter results with sparkrun search:
sparkrun search qwen32. Inspect a recipe
Section titled “2. Inspect a recipe”Before launching, inspect a recipe to understand what it will do and whether it fits in memory:
sparkrun show qwen3-1.7b-vllmThe output includes:
- Model — the HuggingFace model ID (e.g.,
Qwen/Qwen3-1.7B) - Runtime — the inference engine (e.g.,
vllm) - Container — the Docker image that will be pulled
- VRAM estimate — whether the model fits DGX Spark’s 128 GB unified memory at the current tensor parallelism
3. Launch the model
Section titled “3. Launch the model”Run the recipe, pointing at your DGX Spark:
sparkrun run qwen3-1.7b-vllm --hosts <your-spark-ip>If you’ve already created a default cluster, you can omit --hosts:
sparkrun run qwen3-1.7b-vllmsparkrun will:
- Pull the container image to the target host (skipped if already present)
- Download and sync the model from HuggingFace (skipped if already cached)
- Launch the inference container in detached mode
- Stream container logs so you can watch the server start up
You’ll see log output from the inference engine as it loads the model and starts serving.
4. Ctrl+C is safe
Section titled “4. Ctrl+C is safe”Once you see the server is ready (look for a message like Uvicorn running on ...), you can press Ctrl+C at any time.
5. Test the endpoint
Section titled “5. Test the endpoint”The model is now serving an OpenAI-compatible API. Test it with curl:
# List available modelscurl http://<your-spark-ip>:8000/v1/modelsSend a completion request:
curl http://<your-spark-ip>:8000/v1/completions \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen/Qwen3-1.7B", "prompt": "The DGX Spark is", "max_tokens": 50 }'Or use the chat completions endpoint:
curl http://<your-spark-ip>:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen/Qwen3-1.7B", "messages": [{"role": "user", "content": "Hello!"}], "max_tokens": 50 }'6. Re-attach to logs
Section titled “6. Re-attach to logs”If you detached with Ctrl+C and want to see logs again:
sparkrun logs qwen3-1.7b-vllmPass --hosts or --cluster if you used them during launch.
7. Check status
Section titled “7. Check status”See all running sparkrun workloads:
sparkrun statusThis shows running containers grouped by job, with their host, role, status, and ready-to-use stop and logs commands.
8. Stop the workload
Section titled “8. Stop the workload”When you’re done, stop the inference server:
sparkrun stop qwen3-1.7b-vllmPass --hosts or --cluster if you used them during launch. sparkrun stops the containers on all target hosts and cleans up.
Next steps
Section titled “Next steps”- Scale up: Run larger models across multiple Sparks with Multi-Node Tensor Parallelism
- Measure performance: Benchmark your model with Benchmarking Models
- Serve multiple models: Use the Proxy Gateway to expose multiple models through a single API