Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optimizes LLM inference with NVIDIA TensorRT for maximum throughput and lowest latency. Use for production deployment on NVIDIA GPUs (A100/H100), when you need 10-100x faster inference than PyTorch, or for serving models with quantization (FP8/INT4), in-flight batching, and multi-GPU scaling.
.claude/skills/openlair-tensorrt-llm/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 91% | 0% |
NVIDIA's open-source library for optimizing LLM inference with state-of-the-art performance on NVIDIA GPUs.
Use TensorRT-LLM when:
Use vLLM instead when:
Use llama.cpp instead when:
bash# Docker (recommended) docker pull nvidia/tensorrt_llm:latest # pip install pip install tensorrt_llm==1.2.0rc3 # Requires CUDA 13.0.0, TensorRT 10.13.2, Python 3.10-3.12
pythonfrom tensorrt_llm import LLM, SamplingParams # Initialize model llm = LLM(model="meta-llama/Meta-Llama-3-8B") # Configure sampling sampling_params = SamplingParams( max_tokens=100, temperature=0.7, top_p=0.9 ) # Generate prompts = ["Explain quantum computing"] outputs = llm.generate(prompts, sampling_params) for output in outputs: print(output.text)
bash# Start server (automatic model download and compilation) trtllm-serve meta-llama/Meta-Llama-3-8B \ --tp_size 4 \ # Tensor parallelism (4 GPUs) --max_batch_size 256 \ --max_num_tokens 4096 # Client request curl -X POST http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "meta-llama/Meta-Llama-3-8B", "messages": [{"role": "user", "content": "Hello!"}], "temperature": 0.7, "max_tokens": 100 }'
pythonfrom tensorrt_llm import LLM # Load FP8 quantized model (2× faster, 50% memory) llm = LLM( model="meta-llama/Meta-Llama-3-70B", dtype="fp8", max_num_tokens=8192 ) # Inference same as before outputs = llm.generate(["Summarize this article..."])
python# Tensor parallelism across 8 GPUs llm = LLM( model="meta-llama/Meta-Llama-3-405B", tensor_parallel_size=8, dtype="fp8" )
python# Process 100 prompts efficiently prompts = [f"Question {i}: ..." for i in range(100)] outputs = llm.generate( prompts, sampling_params=SamplingParams(max_tokens=200) ) # Automatic in-flight batching for maximum throughput
Meta Llama 3-8B (H100 GPU):
Llama 3-70B (8× A100 80GB):
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 10,710 | 5,183 | -52% | 1 | 1 | 0% | 2,403 | 2,504 | +4% | 0 | 0 | — |
case-02 | pass→pass | 6,931 | 4,476 | -35% | 1 | 1 | 0% | 1,172 | 2,239 | +91% | 0 | 0 | — |
case-03 | pass→pass | 12,357 | 6,062 | -51% | 1 | 1 | 0% | 2,183 | 2,585 | +18% | 0 | 0 | — |
case-04 | pass→pass | 12,375 | 6,117 | -51% | 1 | 1 | 0% | 2,217 | 2,591 | +17% | 0 | 0 | — |
case-05 | pass→pass | 6,963 | 4,188 | -40% | 1 | 1 | 0% | 1,535 | 2,301 | +50% | 0 | 0 | — |
case-06 | pass→pass | 5,482 | 2,552 | -53% | 1 | 1 | 0% | 1,141 | 1,955 | +71% | 0 | 0 | — |
case-07 | fail→fail | 9,546 | 6,235 | -35% | 1 | 1 | 0% | 1,915 | 2,756 | +44% | 0 | 0 | — |
case-08 | fail→fail | 11,034 | 5,295 | -52% | 1 | 1 | 0% | 2,228 | 2,550 | +14% | 0 | 0 | — |
case-09 | fail→pass | 11,015 | 2,530 | -77% | 1 | 1 | 0% | 2,074 | 1,885 | -9% | 0 | 0 | — |
case-10 | fail→pass | 8,044 | 1,459 | -82% | 1 | 1 | 0% | 1,994 | 1,725 | -13% | 0 | 0 | — |
case-11 | fail→pass | 6,051 | 1,768 | -71% | 1 | 1 | 0% | 1,178 | 1,721 | +46% | 0 | 0 | — |
case-12 | pass→pass | 10,039 | 2,735 | -73% | 1 | 1 | 0% | 2,256 | 2,061 | -9% | 0 | 0 | — |
case-13 | pass→pass | 5,098 | 2,686 | -47% | 1 | 1 | 0% | 1,044 | 1,921 | +84% | 0 | 0 | — |
case-14 | pass→pass | 5,059 | 2,524 | -50% | 1 | 1 | 0% | 961 | 1,871 | +95% | 0 | 0 | — |
case-15 | pass→pass | 5,033 | 5,855 | +16% | 1 | 1 | 0% | 910 | 2,627 | +189% | 0 | 0 | — |
case-16 | pass→pass | 3,400 | 2,209 | -35% | 1 | 1 | 0% | 619 | 1,804 | +191% | 0 | 0 | — |
case-17 | pass→pass | 4,366 | 3,983 | -9% | 1 | 1 | 0% | 830 | 2,190 | +164% | 0 | 0 | — |
case-18 | pass→pass | 8,466 | 1,874 | -78% | 1 | 1 | 0% | 1,674 | 1,769 | +6% | 0 | 0 | — |
case-19 | pass→pass | 14,522 | 8,528 | -41% | 1 | 1 | 0% | 2,620 | 3,032 | +16% | 0 | 0 | — |
case-20 | pass→pass | 12,512 | 5,502 | -56% | 1 | 1 | 0% | 2,491 | 2,592 | +4% | 0 | 0 | — |
case-21 | pass→pass | 13,141 | 6,358 | -52% | 1 | 1 | 0% | 2,426 | 2,586 | +7% | 0 | 0 | — |
case-22 | pass→pass | 4,579 | 2,268 | -50% | 1 | 1 | 0% | 1,016 | 1,896 | +87% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.