Install any skill in seconds. Free to start, no credit card required.
Get Started Free →High-throughput LLM inference on NVIDIA GPUs.
.claude/skills/nousresearch-tensorrt-llm/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 26% | 0% |
NVIDIA's open-source library for optimizing LLM inference with high performance on NVIDIA GPUs.
Use TensorRT-LLM when:
Use vLLM instead when:
Use llama.cpp instead when:
bash# Docker (recommended) — images are on NGC (nvcr.io), not Docker Hub. # Replace x.y.z with the desired version (e.g. 1.2.1). Browse tags on NGC: # https://catalog.ngc.nvidia.com/orgs/nvidia/teams/tensorrt-llm/containers/release/tags docker pull nvcr.io/nvidia/tensorrt-llm/release:x.y.z # pip install (current stable GA) pip install tensorrt_llm # Requires CUDA 13.2.1, TensorRT 10.x, Python 3.10-3.12
pythonfrom tensorrt_llm import LLM, SamplingParams # Initialize model llm = LLM(model="meta-llama/Meta-Llama-3-8B") # Configure sampling sampling_params = SamplingParams( max_tokens=100, temperature=0.7, top_p=0.9 ) # Generate prompts = ["Explain quantum computing"] outputs = llm.generate(prompts, sampling_params) for output in outputs: print(output.text)
bash# Start server (automatic model download and compilation) trtllm-serve meta-llama/Meta-Llama-3-8B \ --tp_size 4 \ # Tensor parallelism (4 GPUs) --max_batch_size 256 \ --max_num_tokens 4096 # Client request curl -X POST http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "meta-llama/Meta-Llama-3-8B", "messages": [{"role": "user", "content": "Hello!"}], "temperature": 0.7, "max_tokens": 100 }'
pythonfrom tensorrt_llm import LLM # Load FP8 quantized model (2× faster, 50% memory) llm = LLM( model="meta-llama/Meta-Llama-3-70B", dtype="fp8", max_num_tokens=8192 ) # Inference same as before outputs = llm.generate(["Summarize this article..."])
python# Tensor parallelism across 8 GPUs llm = LLM( model="meta-llama/Meta-Llama-3-405B", tensor_parallel_size=8, dtype="fp8" )
python# Process 100 prompts efficiently prompts = [f"Question {i}: ..." for i in range(100)] outputs = llm.generate( prompts, sampling_params=SamplingParams(max_tokens=200) ) # Automatic in-flight batching for maximum throughput
Meta Llama 3-8B (H100 GPU):
Llama 3-70B (8× A100 80GB):
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 11,835 | 5,663 | -52% | 1 | 1 | 0% | 2,178 | 2,615 | +20% | 0 | 0 | — |
case-01 | fail→fail | 17,184 | 9,443 | -45% | 1 | 1 | 0% | 3,475 | 3,387 | -3% | 0 | 0 | — |
case-02 | fail→pass | 10,382 | 7,645 | -26% | 1 | 1 | 0% | 2,138 | 3,065 | +43% | 0 | 0 | — |
case-03 | pass→fail | 14,650 | 7,273 | -50% | 1 | 1 | 0% | 2,546 | 2,826 | +11% | 0 | 0 | — |
case-05 | pass→pass | 13,197 | 6,041 | -54% | 1 | 1 | 0% | 2,279 | 2,592 | +14% | 0 | 0 | — |
case-06 | fail→pass | 12,174 | 3,163 | -74% | 1 | 1 | 0% | 2,304 | 2,161 | -6% | 0 | 0 | — |
case-07 | pass→pass | 7,856 | 3,514 | -55% | 1 | 1 | 0% | 1,651 | 2,291 | +39% | 0 | 0 | — |
case-08 | pass→pass | 8,370 | 2,953 | -65% | 1 | 1 | 0% | 1,764 | 2,222 | +26% | 0 | 0 | — |
case-09 | fail→fail | 11,851 | 7,390 | -38% | 1 | 1 | 0% | 2,408 | 3,153 | +31% | 0 | 0 | — |
case-10 | fail→fail | 16,074 | 5,717 | -64% | 1 | 1 | 0% | 3,393 | 2,949 | -13% | 0 | 0 | — |
case-11 | pass→pass | 12,303 | 5,945 | -52% | 1 | 1 | 0% | 2,371 | 2,746 | +16% | 0 | 0 | — |
case-12 | pass→pass | 8,138 | 2,764 | -66% | 1 | 1 | 0% | 2,125 | 2,094 | -1% | 0 | 0 | — |
case-13 | pass→pass | 3,747 | 2,372 | -37% | 1 | 1 | 0% | 822 | 2,030 | +147% | 0 | 0 | — |
case-14 | fail→pass | 8,790 | 1,289 | -85% | 1 | 1 | 0% | 1,675 | 1,761 | +5% | 0 | 0 | — |
case-15 | fail→pass | 10,437 | 4,682 | -55% | 1 | 1 | 0% | 2,200 | 2,405 | +9% | 0 | 0 | — |
case-16 | fail→pass | 7,903 | 2,162 | -73% | 1 | 1 | 0% | 1,549 | 1,950 | +26% | 0 | 0 | — |
case-17 | fail→pass | 11,148 | 1,984 | -82% | 1 | 1 | 0% | 2,136 | 1,874 | -12% | 0 | 0 | — |
case-18 | pass→pass | 9,061 | 2,116 | -77% | 1 | 1 | 0% | 1,770 | 1,853 | +5% | 0 | 0 | — |
case-19 | pass→pass | 5,912 | 4,111 | -30% | 1 | 1 | 0% | 1,016 | 2,204 | +117% | 0 | 0 | — |
case-20 | pass→pass | 12,805 | 11,180 | -13% | 1 | 1 | 0% | 2,434 | 3,455 | +42% | 0 | 0 | — |
case-21 | fail→pass | 6,029 | 2,481 | -59% | 1 | 1 | 0% | 1,120 | 1,937 | +73% | 0 | 0 | — |
case-22 | fail→pass | 9,520 | 1,851 | -81% | 1 | 1 | 0% | 2,109 | 1,837 | -13% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/8/2026 | +9% |
Other measured skills in the registry, with their headline benchmark lift.