Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Map evaluation coverage, identify untested capability dimensions — 20 benchmarks, 30 papers, 50 web searches
.claude/skills/yogsoth-ai-coverage-mapping/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 89% | 0% |
| case-01 | ✓→✗ | ▼ Worse | -74% | 0% |
| case-03 | ✓→✗ | ▼ Worse | -71% | 0% |
| case-07 | ✓→✗ | ▼ Worse | -39% | 0% |
| case-12 | ✓→✗ | ▼ Worse | -16% | 0% |
Map the evaluation landscape for a domain to identify which capabilities are well-tested, which are undertested, and which have no evaluation coverage at all. Produces a capability taxonomy with benchmark coverage annotations.
Build a comprehensive map of "what we can and cannot measure" for a given AI capability domain. Identify white spaces where important capabilities lack rigorous evaluation, and redundancies where multiple benchmarks test the same narrow skill.
| Resource | Floor | Target | |----------|-------|--------| | Benchmarks mapped | 15 | 20 | | Papers read | 20 | 30 | | Web searches | 35 | 50 |
<HARD-GATE>
| Metric | Current | Target | Status |
|--------|---------|--------|--------|
| Benchmarks mapped | 0 | 20 | PENDING |
| Capability taxonomy nodes | 0 | 30 | PENDING |
| Papers fetched | 0 | 30 | PENDING |
| Papers read | 0 | 20 | PENDING |
| Web searches | 0 | 50 | PENDING |
| Coverage annotations complete | 0 | 20 | PENDING |
| White spaces identified | 0 | 5 | PENDING |
| Redundancy clusters found | 0 | 3 | PENDING |
</HARD-GATE>Cannot exit until 80% of all targets met.
a. Run capability-taxonomy-mapping to build hierarchical capability tree b. Use papers and web searches to refine taxonomy with community consensus
a. Run benchmark-inventory to collect all known benchmarks in domain b. For each benchmark, run metric-decomposition to identify tested capabilities
a. Map each benchmark to taxonomy nodes it covers b. Identify coverage density per node (over-tested vs under-tested) c. Mark white spaces (zero coverage nodes)
yamlcoverage_map: domain: string taxonomy: - node: string level: int children: list coverage_status: well-covered|partial|minimal|none benchmarks: list[string] white_spaces: - capability: string importance: high|medium|low reason_untested: string proposed_evaluation: string redundancy_clusters: - capability: string benchmarks: list[string] differentiation: string coverage_statistics: total_capabilities: int well_covered: int partial: int minimal: int none: int coverage_ratio: float
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | score-trajectory-analysis | Collect historical scores, fit saturation curves, detect inflection points |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | benchmark-synthesis | Produce final structured audit report | | capability-taxonomy-mapping | Build capability taxonomy, map existing benchmark coverage | | knowledge-acquisition-benchmark-inventory | Identify and catalog all relevant benchmarks in target domain | | metric-decomposition | Decompose composite metrics into constituent signals, analyze polarity and ceiling effects |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 13,287 | 24,988 | +88% | 1 | 1 | 0% | 1,997 | 4,867 | +144% | 0 | 0 | — |
case-15 | pass→pass | 12,512 | 21,499 | +72% | 1 | 1 | 0% | 2,000 | 4,268 | +113% | 0 | 0 | — |
case-21 | pass→pass | 9,909 | 17,956 | +81% | 1 | 1 | 0% | 1,399 | 3,429 | +145% | 0 | 0 | — |
case-01 | pass→fail | 33,956 | 9,576 | -72% | 1 | 1 | 0% | 6,226 | 1,638 | -74% | 0 | 0 | — |
case-02 | fail→fail | 35,773 | 33,800 | -6% | 1 | 1 | 0% | 6,210 | 7,157 | +15% | 0 | 0 | — |
case-03 | pass→fail | 31,107 | 10,019 | -68% | 1 | 1 | 0% | 6,207 | 1,786 | -71% | 0 | 0 | — |
case-09 | pass→pass | 19,873 | 40,319 | +103% | 1 | 1 | 0% | 2,980 | 7,123 | +139% | 0 | 0 | — |
case-04 | pass→pass | 4,037 | 6,550 | +62% | 1 | 1 | 0% | 805 | 2,364 | +194% | 0 | 0 | — |
case-05 | pass→pass | 4,804 | 4,283 | -11% | 1 | 1 | 0% | 1,052 | 1,841 | +75% | 0 | 0 | — |
case-06 | pass→pass | 5,936 | 7,470 | +26% | 1 | 1 | 0% | 1,442 | 2,756 | +91% | 0 | 0 | — |
case-07 | pass→fail | 18,515 | 12,448 | -33% | 1 | 1 | 0% | 2,906 | 1,782 | -39% | 0 | 0 | — |
case-08 | fail→pass | 13,535 | 18,176 | +34% | 1 | 1 | 0% | 2,148 | 4,062 | +89% | 0 | 0 | — |
case-10 | pass→pass | 12,374 | 29,155 | +136% | 1 | 1 | 0% | 1,842 | 4,718 | +156% | 0 | 0 | — |
case-11 | pass→pass | 9,535 | 11,375 | +19% | 1 | 1 | 0% | 2,033 | 3,152 | +55% | 0 | 0 | — |
case-12 | pass→fail | 14,976 | 11,804 | -21% | 1 | 1 | 0% | 2,208 | 1,862 | -16% | 0 | 0 | — |
case-13 | pass→fail | 15,318 | 11,989 | -22% | 1 | 1 | 0% | 2,425 | 1,667 | -31% | 0 | 0 | — |
case-14 | pass→pass | 14,512 | 30,756 | +112% | 1 | 1 | 0% | 2,256 | 6,428 | +185% | 0 | 0 | — |
case-16 | pass→fail | 13,832 | 7,442 | -46% | 1 | 1 | 0% | 2,279 | 2,030 | -11% | 0 | 0 | — |
case-17 | pass→pass | 12,987 | 21,875 | +68% | 1 | 1 | 0% | 2,033 | 4,902 | +141% | 0 | 0 | — |
case-18 | pass→fail | 15,252 | 5,866 | -62% | 1 | 1 | 0% | 2,447 | 1,581 | -35% | 0 | 0 | — |
case-19 | fail→fail | 18,002 | 37,839 | +110% | 1 | 1 | 0% | 2,677 | 7,119 | +166% | 0 | 0 | — |
case-20 | pass→fail | 11,592 | 65,703 | +467% | 1 | 1 | 0% | 1,597 | 7,421 | +365% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -60 percentage points is the difference between those two pass rates over the 17 comparable cases. 9 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.