Skip to content

TensorRT-LLM

TensorRT-LLM is NVIDIA’s high-performance inference engine with TensorRT optimizations. sparkrun’s trtllm runtime uses MPI (OpenMPI) for multi-node tensor parallelism. Solo mode (tp=1) uses the standard sleep-infinity + exec pattern.

  • You want NVIDIA’s TensorRT-LLM inference engine with TensorRT optimizations
  • You need high-throughput serving with features like FP8 KV cache, CUDA graphs, and MoE backends
  • Your model is supported by TensorRT-LLM’s PyTorch or TensorRT backend

Default image prefix: nvcr.io/nvidia/tensorrt-llm/release

sparkrun resolves container tags from the recipe. For example, container: nvcr.io/nvidia/tensorrt-llm/release:0.18.0 pulls the 0.18.0 release image from NVIDIA’s NGC registry.

model: nvidia/Llama-3.1-8B-Instruct-FP8
runtime: trtllm
container: nvcr.io/nvidia/tensorrt-llm/release:0.18.0
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 1
backend: pytorch

These recipe defaults keys map to trtllm-serve CLI flags:

KeyFlagDescription
port--portServing port
host--hostBind address
tensor_parallel--tp_sizeTensor parallelism degree
pipeline_parallel--pp_sizePipeline parallelism degree
expert_parallel--ep_sizeExpert parallelism degree (MoE models)
max_num_tokens--max_num_tokensMaximum number of tokens per batch
max_batch_size--max_batch_sizeMaximum batch size
max_model_len--max_seq_lenMaximum sequence length
backend--backendInference backend (pytorch or tensorrt)
tokenizer--tokenizerTokenizer path or name
kv_cache_free_gpu_memory_fraction--kv_cache_free_gpu_memory_fractionFraction of free GPU memory for KV cache
trust_remote_code--trust_remote_codeTrust remote code from HuggingFace (boolean flag)

If no backend is specified, sparkrun defaults to pytorch.

Certain advanced options are passed via a generated extra-llm-api-config.yml file rather than CLI flags. When any of these keys appear in recipe defaults, sparkrun generates the YAML file, writes it into the container, and passes it via --extra_llm_api_options:

Recipe keyConfig sectionYAML field
free_gpu_memory_fractionkv_cache_configfree_gpu_memory_fraction
kv_cache_dtypekv_cache_configdtype
kv_cache_enable_block_reusekv_cache_configenable_block_reuse
cuda_graph_paddingcuda_graph_configenable_padding
cuda_graph_max_batch_sizecuda_graph_configmax_batch_size
moe_backendmoe_configbackend
print_iter_log(top-level)print_iter_log

Example recipe using extra config keys:

model: some/moe-model
runtime: trtllm
container: nvcr.io/nvidia/tensorrt-llm/release:0.18.0
defaults:
port: 8000
host: 0.0.0.0
tensor_parallel: 2
backend: pytorch
free_gpu_memory_fraction: 0.85
kv_cache_dtype: fp8_e5m2
cuda_graph_padding: true
moe_backend: cutlass

This generates an extra-llm-api-config.yml like:

kv_cache_config:
free_gpu_memory_fraction: 0.85
dtype: fp8_e5m2
cuda_graph_config:
enable_padding: true
moe_config:
backend: cutlass

TRT-LLM uses MPI for multi-node inference. sparkrun handles the full orchestration automatically — container lifecycle, InfiniBand detection, MPI configuration, and environment setup:

Terminal window
sparkrun run my-trtllm-recipe --tp 2
  • SSH keys at ~/.ssh on the control machine (mounted into containers)
  • All hosts must be SSH-accessible from each other (use sparkrun setup ssh to configure)

TRT-LLM containers automatically get these extra Docker options for large memory allocations:

  • --ulimit memlock=-1
  • --ulimit stack=67108864

See Execution Flow for the full MPI orchestration details, including the rsh wrapper mechanism and environment variables.

sparkrun auto-detects TRT-LLM recipes by checking the command template. If the command starts with trtllm-serve or mpirun...trtllm, the trtllm runtime is selected automatically.