Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Analyze evaluation protocol differences across papers for same benchmark — 5 benchmarks, 60 papers, 30 web searches
.claude/skills/yogsoth-ai-protocol-forensics/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 190% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 12% | 0% |
Forensic analysis of how the same benchmark is implemented differently across papers. Reveals hidden variance in evaluation protocols that makes cross-paper score comparisons unreliable.
Expose the "reproducibility gap" in benchmark evaluation by documenting how papers differ in their implementation of supposedly standardized evaluation protocols. Quantify the score variance attributable to protocol differences rather than model improvements.
| Resource | Floor | Target | |----------|-------|--------| | Benchmarks forensically analyzed | 3 | 5 | | Papers read | 45 | 60 | | Web searches | 20 | 30 |
<HARD-GATE>
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| Benchmarks analyzed | 0 | 5 | PENDING |
| Papers fetched | 0 | 60 | PENDING |
| Papers read | 0 | 45 | PENDING |
| Web searches | 0 | 30 | PENDING |
| Protocol extractions complete | 0 | 60 | PENDING |
| Difference matrices built | 0 | 5 | PENDING |
| Variance attributions done | 0 | 5 | PENDING |
| Impact assessments complete | 0 | 5 | PENDING |
</HARD-GATE>Cannot exit until 80% of all targets met.
a. Collect 10-15 papers that report results on the same benchmark b. Prioritize papers from different labs, time periods, and model families
a. Run protocol-element-extraction on each paper b. Extract: prompt format, few-shot examples, decoding parameters, evaluation script version, data split, preprocessing
a. Run evaluation-protocol-comparison tactic to build difference matrix b. Identify which protocol elements vary most across papers c. Estimate score impact of each protocol difference
yamlprotocol_forensics: benchmark_name: string papers_analyzed: int protocol_elements: - element: string # e.g., "few-shot examples", "decoding temperature" variation_level: none|low|medium|high|extreme values_observed: list[string] score_impact_estimate: string variance_attribution: genuine_improvement: float # proportion of score gains protocol_optimization: float implementation_differences: float unexplained: float most_impactful_differences: - element: string score_range: string # e.g., "+3.2 to +7.8 points" evidence: string reproducibility_grade: A|B|C|D|F recommendations: - recommendation: string priority: high|medium|low
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | evaluation-protocol-comparison | Compare implementation differences of same benchmark across papers |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | benchmark-synthesis | Produce final structured audit report | | knowledge-acquisition-benchmark-inventory | Identify and catalog all relevant benchmarks in target domain | | protocol-element-extraction | Extract evaluation protocol parameters from papers |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 26,173 | 17,241 | -34% | 1 | 1 | 0% | 4,049 | 1,627 | -60% | 0 | 0 | — |
case-02 | fail→fail | 28,620 | 25,379 | -11% | 1 | 1 | 0% | 4,754 | 2,582 | -46% | 0 | 0 | — |
case-19 | fail→fail | 41,465 | 59,891 | +44% | 1 | 1 | 0% | 5,071 | 9,215 | +82% | 0 | 0 | — |
case-20 | fail→fail | 48,564 | 48,873 | +1% | 1 | 1 | 0% | 8,237 | 9,215 | +12% | 0 | 0 | — |
case-03 | fail→fail | 35,030 | 52,247 | +49% | 1 | 1 | 0% | 6,342 | 9,297 | +47% | 0 | 0 | — |
case-04 | fail→fail | 48,091 | 20,034 | -58% | 1 | 1 | 0% | 8,252 | 2,053 | -75% | 0 | 0 | — |
case-05 | fail→fail | 49,092 | 54,092 | +10% | 1 | 1 | 0% | 8,243 | 9,221 | +12% | 0 | 0 | — |
case-06 | fail→fail | 43,663 | 21,433 | -51% | 1 | 1 | 0% | 8,251 | 1,764 | -79% | 0 | 0 | — |
case-07 | fail→fail | 49,597 | 56,969 | +15% | 1 | 1 | 0% | 8,256 | 9,234 | +12% | 0 | 0 | — |
case-08 | fail→fail | 33,360 | 51,765 | +55% | 1 | 1 | 0% | 5,766 | 9,239 | +60% | 0 | 0 | — |
case-09 | fail→pass | 39,467 | 48,534 | +23% | 1 | 1 | 0% | 8,238 | 9,216 | +12% | 0 | 0 | — |
case-10 | fail→fail | 37,808 | 101,853 | +169% | 1 | 1 | 0% | 5,841 | 9,225 | +58% | 0 | 0 | — |
case-11 | fail→fail | 40,242 | 14,280 | -65% | 1 | 1 | 0% | 8,248 | 1,707 | -79% | 0 | 0 | — |
case-12 | fail→fail | 47,445 | 61,956 | +31% | 1 | 1 | 0% | 8,234 | 9,212 | +12% | 0 | 0 | — |
case-13 | fail→pass | 27,816 | 24,307 | -13% | 1 | 1 | 0% | 4,008 | 4,576 | +14% | 0 | 0 | — |
case-14 | fail→pass | 19,293 | 19,484 | +1% | 1 | 1 | 0% | 2,291 | 3,551 | +55% | 0 | 0 | — |
case-15 | fail→fail | 19,516 | 18,958 | -3% | 1 | 1 | 0% | 2,332 | 1,858 | -20% | 0 | 0 | — |
case-16 | fail→pass | 24,318 | 48,476 | +99% | 1 | 1 | 0% | 3,180 | 9,213 | +190% | 0 | 0 | — |
case-17 | pass→pass | 24,486 | 58,947 | +141% | 1 | 1 | 0% | 3,772 | 8,686 | +130% | 0 | 0 | — |
case-18 | fail→fail | 51,329 | 22,466 | -56% | 1 | 1 | 0% | 8,244 | 1,892 | -77% | 0 | 0 | — |
case-21 | fail→pass | 45,271 | 49,791 | +10% | 1 | 1 | 0% | 8,239 | 9,217 | +12% | 0 | 0 | — |
case-22 | fail→pass | 42,404 | 22,345 | -47% | 1 | 1 | 0% | 2,629 | 4,523 | +72% | 0 | 0 | — |
case-23 | pass→pass | 19,857 | 25,890 | +30% | 1 | 1 | 0% | 4,218 | 6,114 | +45% | 0 | 0 | — |
case-24 | pass→pass | 16,770 | 23,988 | +43% | 1 | 1 | 0% | 2,531 | 5,548 | +119% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 16 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +25 percentage points is the difference between those two pass rates over the 16 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.