Claude Code Plugin
The sparkrun plugin for Claude Code teaches Claude how to manage LLM inference workloads on your DGX Spark systems.
What it does
Section titled “What it does”- Slash commands — Quick actions for running, stopping, and managing inference jobs
- Skills — Detailed reference that Claude uses automatically when working with sparkrun
Installation
Section titled “Installation”From the marketplace
Section titled “From the marketplace”# Add the marketplace (one-time setup)claude plugin marketplace add spark-arena/sparkrun
# Install the pluginclaude plugin install sparkrun@sparkrunPrerequisites
Section titled “Prerequisites”sparkrun CLI
Section titled “sparkrun CLI”# Install via uvx (recommended)uvx sparkrun setup installDGX Spark cluster
Section titled “DGX Spark cluster”You need SSH access to one or more DGX Spark systems:
sparkrun cluster create mylab --hosts 192.168.11.13,192.168.11.14 -d "My DGX Spark lab"sparkrun cluster set-default mylabsparkrun setup ssh --cluster mylabClaude Code
Section titled “Claude Code”And you’ll obviously need to be using Claude Code.
Natural language usage
Section titled “Natural language usage”Just describe what you want — Claude uses the skills automatically:
- “Run the Qwen3 1.7B model on my cluster”
- “What inference jobs are running?”
- “Stop the nemotron model”
- “Show me available recipes for llama models”
- “Set up sparkrun on my DGX Spark cluster”
Key concepts
Section titled “Key concepts”- Recipes are YAML files describing an inference workload (model, runtime, container, defaults)
- Runtimes are inference engines: vLLM, SGLang, llama.cpp
- Clusters are named groups of DGX Spark hosts
- Each DGX Spark has 1 GPU, so
--tp N(tensor parallelism) = N hosts - sparkrun launches detached containers — Ctrl+C detaches from logs, never kills the job