Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluation Methodology Archaeology Campaign — 5 strategies for systematic analysis of AI/ML benchmarks, metrics, and leaderboards. Reveals construct validity issues, saturation, data contamination, and evaluation protocol inconsistencies.
.claude/skills/yogsoth-ai-benchmark-archaeology/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -17% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -42% | 0% |
Systematic excavation and critical analysis of AI/ML evaluation methodology. Treats benchmarks as historical artifacts requiring forensic examination — uncovering hidden assumptions, methodological drift, validity decay, and coverage gaps that accumulate over time.
| Signal | Route To | |--------|----------| | "audit this benchmark", "benchmark quality", "BetterBench" | benchmark-audit | | "saturation", "ceiling", "score plateau", "when will X be solved" | saturation-analysis | | "does it actually measure", "construct validity", "what does score mean" | validity-probing | | "what's not tested", "coverage gaps", "missing capabilities" | coverage-mapping | | "different papers get different scores", "protocol differences" | protocol-forensics |
| Strategy | Purpose | |----------|---------| | benchmark-audit | Systematic quality assessment using BetterBench 46-criterion framework | | saturation-analysis | Track score trajectories, detect saturation and failure points | | validity-probing | Challenge construct validity — does benchmark measure claimed capability? | | coverage-mapping | Map evaluation coverage, identify untested capability dimensions | | protocol-forensics | Analyze evaluation protocol differences across papers for same benchmark |
| Tactic | Purpose | |--------|---------| | score-trajectory-analysis | Collect historical scores, fit saturation curves, detect inflection points | | artifact-detection | Detect annotation artifacts and shortcuts in benchmarks | | evaluation-protocol-comparison | Compare implementation differences of same benchmark across papers |
| SOP | Purpose | |-----|---------| | benchmark-inventory | Identify and catalog all relevant benchmarks in target domain | | metric-decomposition | Decompose composite metrics into constituent signals | | contamination-audit | Detect train-test data leakage and memorization artifacts | | construct-validity-assessment | Evaluate whether benchmark measures its claimed capability | | documentation-audit | Assess documentation completeness against BetterBench/Datasheets standards | | capability-taxonomy-mapping | Build capability taxonomy, map existing benchmark coverage | | leaderboard-dynamics-analysis | Analyze leaderboard score distributions, compression, selective reporting | | protocol-element-extraction | Extract evaluation protocol parameters from papers | | benchmark-synthesis | Produce final structured audit report | | saturation-detection (shared) | Detect saturation signals in score trajectories (from literature-survey) |
| Strategy | Benchmarks | Papers | Web Searches | |----------|-----------|--------|--------------| | benchmark-audit | 5 | 30 | 40 | | saturation-analysis | 15 | 50 | 60 | | validity-probing | 3 | 40 | 30 | | coverage-mapping | 20 | 30 | 50 | | protocol-forensics | 5 | 60 | 30 | | Total | 48 | 210 | 210 |
| MCP Server | Tools | |------------|-------| | brave-search | brave_web_search, brave_llm_context | | apify | rag-web-browser, google-scholar-scraper | | alphaxiv | get_paper_content, answer_pdf_queries | | semantic-scholar | ss_paper, ss_relevance_search, ss_citations, ss_references |
All outputs write to context/benchmark-archaeology/:
context/benchmark-archaeology/
audit/ # benchmark-audit outputs
saturation/ # saturation-analysis outputs
validity/ # validity-probing outputs
coverage/ # coverage-mapping outputs
forensics/ # protocol-forensics outputs
synthesis/ # Final cross-strategy synthesisEach strategy maintains its own state ledger within its output directory.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | benchmark-audit | Systematic quality assessment using BetterBench 46-criterion framework — 5 benchmarks, 30 papers, 40 web searches | | coverage-mapping | Map evaluation coverage, identify untested capability dimensions — 20 benchmarks, 30 papers, 50 web searches | | protocol-forensics | Analyze evaluation protocol differences across papers for same benchmark — 5 benchmarks, 60 papers, 30 web searches | | saturation-analysis | Track score trajectories, detect saturation/failure points — 15 benchmarks, 50 papers, 60 web searches | | validity-probing | Challenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,247 | 47,614 | +96% | 1 | 1 | 0% | 3,660 | 7,610 | +108% | 0 | 0 | — |
case-02 | fail→fail | 19,253 | 5,651 | -71% | 1 | 1 | 0% | 3,068 | 1,610 | -48% | 0 | 0 | — |
case-03 | fail→fail | 28,754 | 8,954 | -69% | 1 | 1 | 0% | 4,281 | 1,589 | -63% | 0 | 0 | — |
case-04 | pass→fail | 15,418 | 8,677 | -44% | 1 | 1 | 0% | 2,220 | 1,852 | -17% | 0 | 0 | — |
case-05 | pass→fail | 19,668 | 6,589 | -66% | 1 | 1 | 0% | 2,788 | 1,612 | -42% | 0 | 0 | — |
case-06 | fail→fail | 9,042 | 9,092 | +1% | 1 | 1 | 0% | 1,480 | 1,819 | +23% | 0 | 0 | — |
case-07 | pass→fail | 15,682 | 8,024 | -49% | 1 | 1 | 0% | 2,336 | 1,845 | -21% | 0 | 0 | — |
case-08 | pass→fail | 18,875 | 10,982 | -42% | 1 | 1 | 0% | 2,819 | 1,786 | -37% | 0 | 0 | — |
case-09 | pass→fail | 20,020 | 8,001 | -60% | 1 | 1 | 0% | 3,245 | 1,755 | -46% | 0 | 0 | — |
case-10 | fail→fail | 25,213 | 6,370 | -75% | 1 | 1 | 0% | 3,954 | 1,566 | -60% | 0 | 0 | — |
case-11 | pass→fail | 23,442 | 5,805 | -75% | 1 | 1 | 0% | 4,059 | 1,541 | -62% | 0 | 0 | — |
case-12 | pass→fail | 24,250 | 7,328 | -70% | 1 | 1 | 0% | 3,861 | 1,530 | -60% | 0 | 0 | — |
case-13 | fail→fail | 21,673 | 11,706 | -46% | 1 | 1 | 0% | 3,078 | 1,942 | -37% | 0 | 0 | — |
case-14 | fail→pass | 16,219 | 4,058 | -75% | 1 | 1 | 0% | 2,620 | 1,867 | -29% | 0 | 0 | — |
case-15 | fail→pass | 12,680 | 2,027 | -84% | 1 | 1 | 0% | 1,976 | 1,503 | -24% | 0 | 0 | — |
case-16 | fail→pass | 17,006 | 2,983 | -82% | 1 | 1 | 0% | 2,369 | 1,694 | -28% | 0 | 0 | — |
case-17 | fail→fail | 28,786 | 12,259 | -57% | 1 | 1 | 0% | 4,686 | 1,794 | -62% | 0 | 0 | — |
case-18 | fail→fail | 11,610 | 6,519 | -44% | 1 | 1 | 0% | 1,650 | 1,684 | +2% | 0 | 0 | — |
case-19 | fail→fail | 10,809 | 28,557 | +164% | 1 | 1 | 0% | 1,473 | 3,719 | +152% | 0 | 0 | — |
case-20 | pass→pass | 10,475 | 9,302 | -11% | 1 | 1 | 0% | 1,832 | 2,868 | +57% | 0 | 0 | — |
case-21 | pass→pass | 15,901 | 23,872 | +50% | 1 | 1 | 0% | 2,864 | 5,343 | +87% | 0 | 0 | — |
case-22 | pass→pass | 19,430 | 21,252 | +9% | 1 | 1 | 0% | 3,711 | 5,845 | +58% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 7 counted toward the lift figure. The other 15 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -18 percentage points is the difference between those two pass rates over the 7 comparable cases. 9 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.