Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
.claude/skills/openlair-nemo-evaluator-sdk/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 90% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 246% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 120% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 143% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 137% | 0% |
NeMo Evaluator SDK evaluates LLMs across 100+ benchmarks from 18+ harnesses using containerized, reproducible evaluation with multi-backend execution (local Docker, Slurm HPC, Lepton cloud).
Installation:
bashpip install nemo-evaluator-launcher
Set API key and run evaluation:
bashexport NGC_API_KEY=nvapi-your-key-here # Create minimal config cat > config.yaml << 'EOF' defaults: - execution: local - deployment: none - _self_ execution: output_dir: ./results target: api_endpoint: model_id: meta/llama-3.1-8b-instruct url: https://integrate.api.nvidia.com/v1/chat/completions api_key_name: NGC_API_KEY evaluation: tasks: - name: ifeval EOF # Run evaluation nemo-evaluator-launcher run --config-dir . --config-name config
View available tasks:
bashnemo-evaluator-launcher ls tasks
Run core academic benchmarks (MMLU, GSM8K, IFEval) on any OpenAI-compatible endpoint.
Checklist:
Standard Evaluation:
- [ ] Step 1: Configure API endpoint
- [ ] Step 2: Select benchmarks
- [ ] Step 3: Run evaluation
- [ ] Step 4: Check resultsStep 1: Configure API endpoint
yaml# config.yaml defaults: - execution: local - deployment: none - _self_ execution: output_dir: ./results target: api_endpoint: model_id: meta/llama-3.1-8b-instruct url: https://integrate.api.nvidia.com/v1/chat/completions api_key_name: NGC_API_KEY
For self-hosted endpoints (vLLM, TRT-LLM):
yamltarget: api_endpoint: model_id: my-model url: http://localhost:8000/v1/chat/completions api_key_name: "" # No key needed for local
Step 2: Select benchmarks
Add tasks to your config:
yamlevaluation: tasks: - name: ifeval # Instruction following - name: gpqa_diamond # Graduate-level QA env_vars: HF_TOKEN: HF_TOKEN # Some tasks need HF token - name: gsm8k_cot_instruct # Math reasoning - name: humaneval # Code generation
Step 3: Run evaluation
bash# Run with config file nemo-evaluator-launcher run \ --config-dir . \ --config-name config # Override output directory nemo-evaluator-launcher run \ --config-dir . \ --config-name config \ -o execution.output_dir=./my_results # Limit samples for quick testing nemo-evaluator-launcher run \ --config-dir . \ --config-name config \ -o +evaluation.nemo_evaluator_config.config.params.limit_samples=10
Step 4: Check results
bash# Check job status nemo-evaluator-launcher status <invocation_id> # List all runs nemo-evaluator-launcher ls runs # View results cat results/<invocation_id>/<task>/artifacts/results.yml
Execute large-scale evaluation on HPC infrastructure.
Checklist:
Slurm Evaluation:
- [ ] Step 1: Configure Slurm settings
- [ ] Step 2: Set up model deployment
- [ ] Step 3: Launch evaluation
- [ ] Step 4: Monitor job statusStep 1: Configure Slurm settings
yaml# slurm_config.yaml defaults: - execution: slurm - deployment: vllm - _self_ execution: hostname: cluster.example.com account: my_slurm_account partition: gpu output_dir: /shared/results walltime: "04:00:00" nodes: 1 gpus_per_node: 8
Step 2: Set up model deployment
yamldeployment: checkpoint_path: /shared/models/llama-3.1-8b tensor_parallel_size: 2 data_parallel_size: 4 max_model_len: 4096 target: api_endpoint: model_id: llama-3.1-8b # URL auto-generated by deployment
Step 3: Launch evaluation
bashnemo-evaluator-launcher run \ --config-dir . \ --config-name slurm_config
Step 4: Monitor job status
bash# Check status (queries sacct) nemo-evaluator-launcher status <invocation_id> # View detailed info nemo-evaluator-launcher info <invocation_id> # Kill if needed nemo-evaluator-launcher kill <invocation_id>
Benchmark multiple models on the same tasks for comparison.
Checklist:
Model Comparison:
- [ ] Step 1: Create base config
- [ ] Step 2: Run evaluations with overrides
- [ ] Step 3: Export and compare resultsStep 1: Create base config
yaml# base_eval.yaml defaults: - execution: local - deployment: none - _self_ execution: output_dir: ./comparison_results evaluation: nemo_evaluator_config: config: params: temperature: 0.01 parallelism: 4 tasks: - name: mmlu_pro - name: gsm8k_cot_instruct - name: ifeval
Step 2: Run evaluations with model overrides
bash# Evaluate Llama 3.1 8B nemo-evaluator-launcher run \ --config-dir . \ --config-name base_eval \ -o target.api_endpoint.model_id=meta/llama-3.1-8b-instruct \ -o target.api_endpoint.url=https://integrate.api.nvidia.com/v1/chat/completions # Evaluate Mistral 7B nemo-evaluator-launcher run \ --config-dir . \ --config-name base_eval \ -o target.api_endpoint.model_id=mistralai/mistral-7b-instruct-v0.3 \ -o target.api_endpoint.url=https://integrate.api.nvidia.com/v1/chat/completions
Step 3: Export and compare
bash# Export to MLflow nemo-evaluator-launcher export <invocation_id_1> --dest mlflow nemo-evaluator-launcher export <invocation_id_2> --dest mlflow # Export to local JSON nemo-evaluator-launcher export <invocation_id> --dest local --format json # Export to Weights & Biases nemo-evaluator-launcher export <invocation_id> --dest wandb
Evaluate models on safety benchmarks and VLM tasks.
Checklist:
Safety/VLM Evaluation:
- [ ] Step 1: Configure safety tasks
- [ ] Step 2: Set up VLM tasks (if applicable)
- [ ] Step 3: Run evaluationStep 1: Configure safety tasks
yamlevaluation: tasks: - name: aegis # Safety harness - name: wildguard # Safety classification - name: garak # Security probing
Step 2: Configure VLM tasks
yaml# For vision-language models target: api_endpoint: type: vlm # Vision-language endpoint model_id: nvidia/llama-3.2-90b-vision-instruct url: https://integrate.api.nvidia.com/v1/chat/completions evaluation: tasks: - name: ocrbench # OCR evaluation - name: chartqa # Chart understanding - name: mmmu # Multimodal understanding
Use NeMo Evaluator when:
Use alternatives instead:
| Harness | Task Count | Categories | |---------|-----------|------------| | lm-evaluation-harness | 60+ | MMLU, GSM8K, HellaSwag, ARC | | simple-evals | 20+ | GPQA, MATH, AIME | | bigcode-evaluation-harness | 25+ | HumanEval, MBPP, MultiPL-E | | safety-harness | 3 | Aegis, WildGuard | | garak | 1 | Security probing | | vlmevalkit | 6+ | OCRBench, ChartQA, MMMU | | bfcl | 6 | Function calling v2/v3 | | mtbench | 2 | Multi-turn conversation | | livecodebench | 10+ | Live coding evaluation | | helm | 15 | Medical domain | | nemo-skills | 8 | Math, science, agentic |
Issue: Container pull fails
Ensure NGC credentials are configured:
bashdocker login nvcr.io -u '$oauthtoken' -p $NGC_API_KEY
Issue: Task requires environment variable
Some tasks need HF_TOKEN or JUDGE_API_KEY:
yamlevaluation: tasks: - name: gpqa_diamond env_vars: HF_TOKEN: HF_TOKEN # Maps env var name to env var
Issue: Evaluation timeout
Increase parallelism or reduce samples:
bash-o +evaluation.nemo_evaluator_config.config.params.parallelism=8 -o +evaluation.nemo_evaluator_config.config.params.limit_samples=100
Issue: Slurm job not starting
Check Slurm account and partition:
yamlexecution: account: correct_account partition: gpu qos: normal # May need specific QOS
Issue: Different results than expected
Verify configuration matches reported settings:
yamlevaluation: nemo_evaluator_config: config: params: temperature: 0.0 # Deterministic num_fewshot: 5 # Check paper's fewshot count
| Command | Description | |---------|-------------| | run | Execute evaluation with config | | status <id> | Check job status | | info <id> | View detailed job info | | ls tasks | List available benchmarks | | ls runs | List all invocations | | export <id> | Export results (mlflow/wandb/local) | | kill <id> | Terminate running job |
bash# Override model endpoint -o target.api_endpoint.model_id=my-model -o target.api_endpoint.url=http://localhost:8000/v1/chat/completions # Add evaluation parameters -o +evaluation.nemo_evaluator_config.config.params.temperature=0.5 -o +evaluation.nemo_evaluator_config.config.params.parallelism=8 -o +evaluation.nemo_evaluator_config.config.params.limit_samples=50 # Change execution settings -o execution.output_dir=/custom/path -o execution.mode=parallel # Dynamically set tasks -o 'evaluation.tasks=[{name: ifeval}, {name: gsm8k}]'
For programmatic evaluation without the CLI:
pythonfrom nemo_evaluator.core.evaluate import evaluate from nemo_evaluator.api.api_dataclasses import ( EvaluationConfig, EvaluationTarget, ApiEndpoint, EndpointType, ConfigParams ) # Configure evaluation eval_config = EvaluationConfig( type="mmlu_pro", output_dir="./results", params=ConfigParams( limit_samples=10, temperature=0.0, max_new_tokens=1024, parallelism=4 ) ) # Configure target endpoint target_config = EvaluationTarget( api_endpoint=ApiEndpoint( model_id="meta/llama-3.1-8b-instruct", url="https://integrate.api.nvidia.com/v1/chat/completions", type=EndpointType.CHAT, api_key="nvapi-your-key-here" ) ) # Run evaluation result = evaluate(eval_cfg=eval_config, target_cfg=target_config)
Multi-backend execution: See references/execution-backends.md Configuration deep-dive: See references/configuration.md Adapter and interceptor system: See references/adapter-system.md Custom benchmark integration: See references/custom-benchmarks.md
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→pass | 13,009 | 7,796 | -40% | 1 | 1 | 0% | 2,883 | 5,476 | +90% | 0 | 0 | — |
case-01 | fail→pass | 7,173 | 10,400 | +45% | 1 | 1 | 0% | 1,749 | 6,047 | +246% | 0 | 0 | — |
case-02 | fail→pass | 10,842 | 6,792 | -37% | 1 | 1 | 0% | 2,368 | 5,203 | +120% | 0 | 0 | — |
case-04 | fail→fail | 11,053 | 4,309 | -61% | 1 | 1 | 0% | 2,101 | 4,393 | +109% | 0 | 0 | — |
case-05 | fail→pass | 10,732 | 4,974 | -54% | 1 | 1 | 0% | 1,867 | 4,541 | +143% | 0 | 0 | — |
case-06 | fail→pass | 10,130 | 6,044 | -40% | 1 | 1 | 0% | 2,097 | 4,976 | +137% | 0 | 0 | — |
case-07 | fail→pass | 11,789 | 6,985 | -41% | 1 | 1 | 0% | 2,624 | 5,194 | +98% | 0 | 0 | — |
case-08 | fail→pass | 17,190 | 3,387 | -80% | 1 | 1 | 0% | 3,398 | 4,351 | +28% | 0 | 0 | — |
case-09 | fail→pass | 12,909 | 2,790 | -78% | 1 | 1 | 0% | 2,485 | 4,198 | +69% | 0 | 0 | — |
case-10 | pass→pass | 3,338 | 3,173 | -5% | 1 | 1 | 0% | 702 | 4,000 | +470% | 0 | 0 | — |
case-11 | fail→pass | 7,568 | 4,072 | -46% | 1 | 1 | 0% | 1,702 | 4,453 | +162% | 0 | 0 | — |
case-12 | fail→pass | 8,620 | 2,761 | -68% | 1 | 1 | 0% | 1,899 | 4,206 | +121% | 0 | 0 | — |
case-13 | fail→pass | 7,524 | 2,444 | -68% | 1 | 1 | 0% | 1,562 | 4,142 | +165% | 0 | 0 | — |
case-14 | fail→pass | 8,390 | 1,403 | -83% | 1 | 1 | 0% | 1,738 | 3,851 | +122% | 0 | 0 | — |
case-15 | fail→pass | 10,047 | 4,140 | -59% | 1 | 1 | 0% | 1,873 | 4,456 | +138% | 0 | 0 | — |
case-16 | fail→pass | 6,835 | 2,751 | -60% | 1 | 1 | 0% | 1,421 | 4,145 | +192% | 0 | 0 | — |
case-17 | fail→pass | 8,219 | 3,333 | -59% | 1 | 1 | 0% | 1,545 | 4,275 | +177% | 0 | 0 | — |
case-18 | fail→pass | 11,870 | 2,682 | -77% | 1 | 1 | 0% | 2,432 | 4,185 | +72% | 0 | 0 | — |
case-19 | fail→pass | 8,623 | 3,173 | -63% | 1 | 1 | 0% | 1,650 | 4,230 | +156% | 0 | 0 | — |
case-20 | pass→pass | 9,050 | 8,517 | -6% | 1 | 1 | 0% | 1,986 | 5,718 | +188% | 0 | 0 | — |
case-21 | pass→pass | 6,711 | 5,715 | -15% | 1 | 1 | 0% | 1,444 | 4,768 | +230% | 0 | 0 | — |
case-22 | pass→pass | 8,369 | 7,084 | -15% | 1 | 1 | 0% | 1,795 | 5,131 | +186% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +77 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.