Skip to content

Your First Model

This tutorial walks you through launching your first LLM on a DGX Spark using sparkrun — from browsing available recipes to testing the API endpoint and stopping the workload.

  • sparkrun installed (Installation guide)
  • SSH access to a DGX Spark (password or key-based)
  • Docker running on the target Spark

If you haven’t set up SSH yet, see SSH Setup.

Start by seeing what models are available:

Terminal window
sparkrun list

This searches all configured registries and displays available recipes. Each recipe name follows a pattern: <model>-<runtime> (e.g., qwen3-1.7b-vllm runs the Qwen3 1.7B model on the vLLM runtime).

You can filter results with sparkrun search:

Terminal window
sparkrun search qwen3

Before launching, inspect a recipe to understand what it will do and whether it fits in memory:

Terminal window
sparkrun show qwen3-1.7b-vllm

The output includes:

  • Model — the HuggingFace model ID (e.g., Qwen/Qwen3-1.7B)
  • Runtime — the inference engine (e.g., vllm)
  • Container — the Docker image that will be pulled
  • VRAM estimate — whether the model fits DGX Spark’s 128 GB unified memory at the current tensor parallelism

Run the recipe, pointing at your DGX Spark:

Terminal window
sparkrun run qwen3-1.7b-vllm --hosts <your-spark-ip>

If you’ve already created a default cluster, you can omit --hosts:

Terminal window
sparkrun run qwen3-1.7b-vllm

sparkrun will:

  1. Pull the container image to the target host (skipped if already present)
  2. Download and sync the model from HuggingFace (skipped if already cached)
  3. Launch the inference container in detached mode
  4. Stream container logs so you can watch the server start up

You’ll see log output from the inference engine as it loads the model and starts serving.

Once you see the server is ready (look for a message like Uvicorn running on ...), you can press Ctrl+C at any time.

The model is now serving an OpenAI-compatible API. Test it with curl:

Terminal window
# List available models
curl http://<your-spark-ip>:8000/v1/models

Send a completion request:

Terminal window
curl http://<your-spark-ip>:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-1.7B",
"prompt": "The DGX Spark is",
"max_tokens": 50
}'

Or use the chat completions endpoint:

Terminal window
curl http://<your-spark-ip>:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-1.7B",
"messages": [{"role": "user", "content": "Hello!"}],
"max_tokens": 50
}'

If you detached with Ctrl+C and want to see logs again:

Terminal window
sparkrun logs qwen3-1.7b-vllm

Pass --hosts or --cluster if you used them during launch.

See all running sparkrun workloads:

Terminal window
sparkrun status

This shows running containers grouped by job, with their host, role, status, and ready-to-use stop and logs commands.

When you’re done, stop the inference server:

Terminal window
sparkrun stop qwen3-1.7b-vllm

Pass --hosts or --cluster if you used them during launch. sparkrun stops the containers on all target hosts and cleans up.