Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluate exploration status of each cell in a method×problem matrix, annotating as explored, partial, or unexplored.
.claude/skills/yogsoth-ai-intersection-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✓→✓ | = Same ✓ | 3% | 0% |
| case-11 | ✓→✓ | = Same ✓ | -28% | 0% |
| case-12 | ✓→✓ | = Same ✓ | 11% | 0% |
| case-13 | ✓→✓ | = Same ✓ | 5% | 0% |
| case-21 | ✗→✗ | = Same ✗ | 2% | 0% |
Evaluate the exploration status of each cell in a method×problem matrix and prioritize unexplored intersections.
Subagent — spawned via subagent-spawning/spawn-agent skill.
Evaluation requires careful assessment of each intersection's exploration depth, consulting literature and benchmarks. Benefits from focused context to maintain consistent evaluation criteria.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | fail→fail | 52,721 | 50,923 | -3% | 1 | 1 | 0% | 8,225 | 8,385 | +2% | 0 | 0 | — |
case-04 | fail→fail | 21,366 | 31,621 | +48% | 1 | 1 | 0% | 2,940 | 4,790 | +63% | 0 | 0 | — |
case-01 | fail→fail | 20,045 | 104,904 | +423% | 1 | 1 | 0% | 3,339 | 1,256 | -62% | 0 | 0 | — |
case-02 | fail→fail | 35,579 | 52,914 | +49% | 1 | 1 | 0% | 5,882 | 6,950 | +18% | 0 | 0 | — |
case-03 | fail→fail | 51,822 | 49,516 | -4% | 1 | 1 | 0% | 8,138 | 8,408 | +3% | 0 | 0 | — |
case-05 | fail→fail | 53,425 | 40,779 | -24% | 1 | 1 | 0% | 8,238 | 8,398 | +2% | 0 | 0 | — |
case-06 | fail→fail | 92,955 | 64,454 | -31% | 1 | 1 | 0% | 7,851 | 8,390 | +7% | 0 | 0 | — |
case-07 | fail→fail | 23,592 | 58,192 | +147% | 1 | 1 | 0% | 3,834 | 8,591 | +124% | 0 | 0 | — |
case-08 | fail→fail | 23,422 | 23,494 | +0% | 1 | 1 | 0% | 3,773 | 906 | -76% | 0 | 0 | — |
case-09 | fail→fail | 39,820 | 57,910 | +45% | 1 | 1 | 0% | 3,667 | 8,390 | +129% | 0 | 0 | — |
case-10 | pass→pass | 20,500 | 24,916 | +22% | 1 | 1 | 0% | 4,652 | 4,798 | +3% | 0 | 0 | — |
case-11 | pass→pass | 26,695 | 7,683 | -71% | 1 | 1 | 0% | 2,420 | 1,737 | -28% | 0 | 0 | — |
case-12 | pass→pass | 13,810 | 30,011 | +117% | 1 | 1 | 0% | 2,746 | 3,038 | +11% | 0 | 0 | — |
case-13 | pass→pass | 20,816 | 16,987 | -18% | 1 | 1 | 0% | 2,145 | 2,258 | +5% | 0 | 0 | — |
case-14 | fail→fail | 26,698 | 73,445 | +175% | 1 | 1 | 0% | 3,704 | 8,388 | +126% | 0 | 0 | — |
case-15 | fail→fail | 39,551 | 48,120 | +22% | 1 | 1 | 0% | 6,313 | 8,388 | +33% | 0 | 0 | — |
case-16 | fail→fail | 25,799 | 42,980 | +67% | 1 | 1 | 0% | 3,411 | 473 | -86% | 0 | 0 | — |
case-17 | fail→fail | 14,635 | 43,356 | +196% | 1 | 1 | 0% | 1,435 | 7,010 | +389% | 0 | 0 | — |
case-18 | fail→fail | 44,650 | 19,274 | -57% | 1 | 1 | 0% | 8,225 | 1,864 | -77% | 0 | 0 | — |
case-19 | fail→fail | 21,061 | 50,008 | +137% | 1 | 1 | 0% | 3,315 | 8,388 | +153% | 0 | 0 | — |
case-20 | fail→fail | 19,389 | 52,702 | +172% | 1 | 1 | 0% | 3,417 | 8,384 | +145% | 0 | 0 | — |
case-22 | fail→fail | 29,901 | 6,520 | -78% | 1 | 1 | 0% | 4,010 | 594 | -85% | 0 | 0 | — |
case-23 | fail→fail | 195,951 | 78,349 | -60% | 1 | 1 | 0% | 5,604 | 8,379 | +50% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 18 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -33 percentage points is the difference between those two pass rates over the 18 comparable cases. 7 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.