Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Systematic evidence map construction — search, classify, locate gaps, visualize. Combines concept-matrix-construction, gap-keyword-extraction, evidence-grading, and egm-construction SOPs.
.claude/skills/yogsoth-ai-evidence-mapping/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 86% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -63% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 72% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 57% | 0% |
Systematically build evidence maps to reveal structural gaps.
Subagent: concept-matrix-construction, gap-keyword-extraction, evidence-grading, egm-construction Import: paper-overview, paper-search
Build concept matrix first (articles × concepts), identify empty cells as gap candidates, verify with keyword extraction from full texts, grade evidence quality in populated cells, construct final EGM for structural view.
<HARD-GATE>
- concept-matrix: >= 1 constructed
- gap-keyword passes: >= 3
- evidence grading: >= 5 items graded
- EGM: >= 1 constructed
</HARD-GATE><!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | concept-matrix-construction | Build articles × concepts coverage matrix to visualize research landscape and identify empty cells as gap candidates. | | deep-insight-paper-overview | Paper metadata and abstract-level overview. Import of literature-engine/literature-overview skill. Abstracts only — no substantive claims without deeper reading. | | deep-insight-paper-search | AI-powered paper summary and search. Import of literature-engine/literature-search skill. AI summary level — cite as "AI-extracted" not "paper states". | | egm-construction | Build structured Evidence Gap Maps — define axes (intervention × outcome or method × domain), place gaps in cells, annotate with evidence density and quality. | | evidence-grading | Assess evidence quality using GRADE/SOE framework. Rates certainty level and identifies downgrade reasons. | | gap-keyword-extraction | Extract gap-indicating sentences and phrases from papers/reviews. Identifies linguistic markers of research gaps (e.g., "remains unclear", "has not been explored", "limited understanding"). |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 27,803 | 6,020 | -78% | 1 | 1 | 0% | 4,890 | 885 | -82% | 0 | 0 | — |
case-02 | fail→pass | 15,540 | 24,632 | +59% | 1 | 1 | 0% | 2,360 | 4,378 | +86% | 0 | 0 | — |
case-03 | pass→pass | 26,159 | 33,566 | +28% | 1 | 1 | 0% | 4,236 | 6,651 | +57% | 0 | 0 | — |
case-04 | fail→fail | 11,834 | 7,310 | -38% | 1 | 1 | 0% | 1,793 | 869 | -52% | 0 | 0 | — |
case-05 | pass→fail | 16,316 | 6,204 | -62% | 1 | 1 | 0% | 2,472 | 906 | -63% | 0 | 0 | — |
case-06 | pass→pass | 20,065 | 25,399 | +27% | 1 | 1 | 0% | 3,332 | 5,266 | +58% | 0 | 0 | — |
case-07 | pass→pass | 10,436 | 6,212 | -40% | 1 | 1 | 0% | 1,770 | 1,554 | -12% | 0 | 0 | — |
case-08 | pass→pass | 12,411 | 8,213 | -34% | 1 | 1 | 0% | 1,999 | 1,793 | -10% | 0 | 0 | — |
case-17 | pass→pass | 7,022 | 2,056 | -71% | 1 | 1 | 0% | 1,050 | 808 | -23% | 0 | 0 | — |
case-09 | pass→pass | 11,464 | 9,048 | -21% | 1 | 1 | 0% | 1,808 | 1,873 | +4% | 0 | 0 | — |
case-10 | pass→pass | 16,858 | 11,545 | -32% | 1 | 1 | 0% | 2,597 | 2,268 | -13% | 0 | 0 | — |
case-11 | pass→pass | 16,144 | 27,358 | +69% | 1 | 1 | 0% | 2,635 | 5,337 | +103% | 0 | 0 | — |
case-12 | pass→pass | 12,395 | 9,684 | -22% | 1 | 1 | 0% | 1,871 | 1,870 | -0% | 0 | 0 | — |
case-13 | fail→pass | 9,022 | 5,012 | -44% | 1 | 1 | 0% | 1,448 | 1,195 | -17% | 0 | 0 | — |
case-14 | pass→pass | 14,690 | 33,564 | +128% | 1 | 1 | 0% | 2,388 | 6,567 | +175% | 0 | 0 | — |
case-15 | pass→pass | 10,111 | 2,359 | -77% | 1 | 1 | 0% | 1,517 | 846 | -44% | 0 | 0 | — |
case-16 | pass→pass | 8,404 | 2,192 | -74% | 1 | 1 | 0% | 1,320 | 810 | -39% | 0 | 0 | — |
case-18 | pass→pass | 6,968 | 2,295 | -67% | 1 | 1 | 0% | 1,009 | 810 | -20% | 0 | 0 | — |
case-19 | pass→pass | 12,198 | 9,918 | -19% | 1 | 1 | 0% | 2,008 | 1,934 | -4% | 0 | 0 | — |
case-20 | fail→fail | 12,605 | 29,739 | +136% | 1 | 1 | 0% | 2,433 | 6,676 | +174% | 0 | 0 | — |
case-21 | pass→fail | 20,744 | 34,863 | +68% | 1 | 1 | 0% | 3,885 | 6,665 | +72% | 0 | 0 | — |
case-22 | pass→pass | 15,126 | 31,257 | +107% | 1 | 1 | 0% | 2,894 | 6,071 | +110% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 19 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.