Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Construct the strongest counter-argument against a specific assumption and propose alternatives.
.claude/skills/yogsoth-ai-convergence-assumption-challenge/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✓→✓ | = Same ✓ | 6% | 0% |
| case-07 | ✓→✓ | = Same ✓ | 93% | 0% |
| case-08 | ✓→✓ | = Same ✓ | 66% | 0% |
| case-01 | ✗→✗ | = Same ✗ | -58% | 0% |
| case-02 | ✗→✗ | = Same ✗ | 19% | 0% |
Attacks a specific assumption adversarially — constructing the strongest argument for why it might be wrong, proposing an alternative assumption, and assessing the impact on the overall decision if the assumption fails.
Spawns a subagent that takes a single assumption and builds the strongest possible case against it.
Output must include:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,916 | 17,089 | -18% | 1 | 1 | 0% | 3,203 | 1,354 | -58% | 0 | 0 | — |
case-02 | fail→fail | 14,902 | 16,554 | +11% | 1 | 1 | 0% | 2,254 | 2,671 | +19% | 0 | 0 | — |
case-03 | fail→fail | 14,615 | 14,391 | -2% | 1 | 1 | 0% | 2,115 | 2,369 | +12% | 0 | 0 | — |
case-04 | fail→fail | 18,947 | 18,711 | -1% | 1 | 1 | 0% | 2,598 | 2,959 | +14% | 0 | 0 | — |
case-05 | fail→fail | 15,207 | 17,234 | +13% | 1 | 1 | 0% | 2,377 | 3,009 | +27% | 0 | 0 | — |
case-06 | pass→pass | 10,992 | 10,948 | -0% | 1 | 1 | 0% | 1,834 | 1,948 | +6% | 0 | 0 | — |
case-07 | pass→pass | 8,958 | 17,628 | +97% | 1 | 1 | 0% | 1,891 | 3,642 | +93% | 0 | 0 | — |
case-08 | pass→pass | 3,280 | 3,876 | +18% | 1 | 1 | 0% | 487 | 806 | +66% | 0 | 0 | — |
case-09 | fail→fail | 19,925 | 21,135 | +6% | 1 | 1 | 0% | 3,101 | 3,440 | +11% | 0 | 0 | — |
case-10 | fail→fail | 15,899 | 15,049 | -5% | 1 | 1 | 0% | 2,357 | 2,415 | +2% | 0 | 0 | — |
case-11 | fail→fail | 17,953 | 18,326 | +2% | 1 | 1 | 0% | 2,786 | 2,910 | +4% | 0 | 0 | — |
case-12 | fail→fail | 14,170 | 17,205 | +21% | 1 | 1 | 0% | 2,022 | 2,720 | +35% | 0 | 0 | — |
case-13 | fail→fail | 15,357 | 14,834 | -3% | 1 | 1 | 0% | 2,331 | 2,416 | +4% | 0 | 0 | — |
case-14 | fail→fail | 18,423 | 22,084 | +20% | 1 | 1 | 0% | 2,545 | 3,141 | +23% | 0 | 0 | — |
case-15 | fail→fail | 15,022 | 17,297 | +15% | 1 | 1 | 0% | 2,185 | 2,842 | +30% | 0 | 0 | — |
case-16 | fail→fail | 19,976 | 25,021 | +25% | 1 | 1 | 0% | 3,026 | 4,238 | +40% | 0 | 0 | — |
case-17 | fail→fail | 17,243 | 18,748 | +9% | 1 | 1 | 0% | 2,432 | 2,925 | +20% | 0 | 0 | — |
case-18 | fail→fail | 22,713 | 18,578 | -18% | 1 | 1 | 0% | 3,148 | 2,902 | -8% | 0 | 0 | — |
case-19 | fail→fail | 21,010 | 20,795 | -1% | 1 | 1 | 0% | 2,888 | 3,425 | +19% | 0 | 0 | — |
case-20 | fail→fail | 16,253 | 15,132 | -7% | 1 | 1 | 0% | 2,508 | 2,543 | +1% | 0 | 0 | — |
case-21 | fail→fail | 19,222 | 25,993 | +35% | 1 | 1 | 0% | 2,839 | 4,137 | +46% | 0 | 0 | — |
case-22 | fail→fail | 19,642 | 13,556 | -31% | 1 | 1 | 0% | 2,741 | 1,093 | -60% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 20 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.