Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Systematically generate counterexamples (monsters) to a given claim using diverse heuristic strategies.
.claude/skills/yogsoth-ai-counterexample-generation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-20 | ✓→✗ | ▼ Worse | -11% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 16% | 0% |
| case-04 | ✓→✓ | = Same ✓ | -34% | 0% |
| case-09 | ✓→✓ | = Same ✓ | -13% | 0% |
| case-11 | ✓→✓ | = Same ✓ | -38% | 0% |
Subagent that produces counterexamples to a claim using boundary cases, degenerate cases, and domain-specific heuristics.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | pass→fail | 7,135 | 5,905 | -17% | 1 | 1 | 0% | 1,436 | 1,271 | -11% | 0 | 0 | — |
case-01 | pass→pass | 10,356 | 12,182 | +18% | 1 | 1 | 0% | 1,955 | 2,262 | +16% | 0 | 0 | — |
case-02 | fail→fail | 10,967 | 12,573 | +15% | 1 | 1 | 0% | 2,197 | 2,839 | +29% | 0 | 0 | — |
case-03 | fail→fail | 8,487 | 9,086 | +7% | 1 | 1 | 0% | 1,667 | 2,042 | +22% | 0 | 0 | — |
case-04 | pass→pass | 11,754 | 8,972 | -24% | 1 | 1 | 0% | 1,884 | 1,236 | -34% | 0 | 0 | — |
case-05 | fail→fail | 12,780 | 10,806 | -15% | 1 | 1 | 0% | 2,127 | 2,166 | +2% | 0 | 0 | — |
case-06 | fail→fail | 8,277 | 9,020 | +9% | 1 | 1 | 0% | 1,636 | 2,069 | +26% | 0 | 0 | — |
case-07 | fail→fail | 12,023 | 9,555 | -21% | 1 | 1 | 0% | 2,505 | 2,031 | -19% | 0 | 0 | — |
case-08 | fail→fail | 12,161 | 8,213 | -32% | 1 | 1 | 0% | 2,544 | 1,980 | -22% | 0 | 0 | — |
case-09 | pass→pass | 8,479 | 7,150 | -16% | 1 | 1 | 0% | 1,789 | 1,557 | -13% | 0 | 0 | — |
case-10 | fail→fail | 5,870 | 5,618 | -4% | 1 | 1 | 0% | 1,014 | 1,041 | +3% | 0 | 0 | — |
case-11 | pass→pass | 7,515 | 3,838 | -49% | 1 | 1 | 0% | 1,518 | 944 | -38% | 0 | 0 | — |
case-12 | pass→pass | 9,021 | 10,634 | +18% | 1 | 1 | 0% | 1,521 | 1,957 | +29% | 0 | 0 | — |
case-13 | fail→fail | 6,789 | 4,314 | -36% | 1 | 1 | 0% | 1,214 | 764 | -37% | 0 | 0 | — |
case-14 | fail→fail | 7,945 | 8,142 | +2% | 1 | 1 | 0% | 1,463 | 1,563 | +7% | 0 | 0 | — |
case-15 | fail→fail | 10,561 | 7,132 | -32% | 1 | 1 | 0% | 1,894 | 1,739 | -8% | 0 | 0 | — |
case-16 | fail→fail | 8,962 | 7,241 | -19% | 1 | 1 | 0% | 1,836 | 1,727 | -6% | 0 | 0 | — |
case-17 | fail→fail | 19,827 | 16,249 | -18% | 1 | 1 | 0% | 3,851 | 3,473 | -10% | 0 | 0 | — |
case-18 | fail→fail | 11,593 | 15,403 | +33% | 1 | 1 | 0% | 2,419 | 3,565 | +47% | 0 | 0 | — |
case-19 | pass→pass | 8,152 | 5,318 | -35% | 1 | 1 | 0% | 1,610 | 1,129 | -30% | 0 | 0 | — |
case-21 | pass→pass | 7,081 | 6,097 | -14% | 1 | 1 | 0% | 1,676 | 1,528 | -9% | 0 | 0 | — |
case-22 | pass→pass | 5,043 | 6,162 | +22% | 1 | 1 | 0% | 963 | 1,333 | +38% | 0 | 0 | — |
case-23 | fail→fail | 8,024 | 8,737 | +9% | 1 | 1 | 0% | 1,463 | 1,764 | +21% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of -100 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Other measured skills in the registry, with their headline benchmark lift.