Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Construct extreme but plausible worst-case scenarios for stress testing
.claude/skills/yogsoth-ai-worst-case-construction/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✓→✗ | ▼ Worse | 71% | 0% |
| case-09 | ✓→✗ | ▼ Worse | 22% | 0% |
| case-14 | ✓→✗ | ▼ Worse | 28% | 0% |
| case-16 | ✓→✓ | = Same ✓ | 21% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 19% | 0% |
Construct extreme but plausible worst-case scenarios that maximally stress the research approach. Identify breaking points, failure cascades, and recovery possibilities.
Subagent — spawned via subagent-spawning/spawn-agent skill.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-16 | pass→pass | 16,998 | 18,256 | +7% | 1 | 1 | 0% | 1,912 | 2,311 | +21% | 0 | 0 | — |
case-01 | pass→pass | 21,451 | 23,521 | +10% | 1 | 1 | 0% | 2,329 | 2,760 | +19% | 0 | 0 | — |
case-02 | pass→pass | 24,698 | 39,070 | +58% | 1 | 1 | 0% | 3,047 | 2,879 | -6% | 0 | 0 | — |
case-03 | fail→fail | 25,751 | 22,915 | -11% | 1 | 1 | 0% | 3,025 | 2,898 | -4% | 0 | 0 | — |
case-04 | pass→pass | 19,216 | 15,141 | -21% | 1 | 1 | 0% | 2,083 | 1,515 | -27% | 0 | 0 | — |
case-05 | pass→pass | 21,512 | 27,682 | +29% | 1 | 1 | 0% | 2,380 | 3,383 | +42% | 0 | 0 | — |
case-06 | pass→pass | 17,932 | 20,453 | +14% | 1 | 1 | 0% | 1,846 | 2,415 | +31% | 0 | 0 | — |
case-07 | pass→fail | 21,366 | 34,017 | +59% | 1 | 1 | 0% | 2,847 | 4,856 | +71% | 0 | 0 | — |
case-08 | pass→pass | 19,285 | 33,662 | +75% | 1 | 1 | 0% | 2,138 | 4,550 | +113% | 0 | 0 | — |
case-09 | pass→fail | 19,033 | 20,748 | +9% | 1 | 1 | 0% | 1,999 | 2,445 | +22% | 0 | 0 | — |
case-10 | fail→fail | 7,564 | 26,109 | +245% | 1 | 1 | 0% | 343 | 3,458 | +908% | 0 | 0 | — |
case-11 | pass→pass | 21,181 | 32,036 | +51% | 1 | 1 | 0% | 2,255 | 4,131 | +83% | 0 | 0 | — |
case-12 | pass→pass | 21,651 | 21,728 | +0% | 1 | 1 | 0% | 2,472 | 2,516 | +2% | 0 | 0 | — |
case-13 | pass→pass | 18,900 | 34,514 | +83% | 1 | 1 | 0% | 2,154 | 4,749 | +120% | 0 | 0 | — |
case-14 | pass→fail | 26,476 | 33,536 | +27% | 1 | 1 | 0% | 3,743 | 4,777 | +28% | 0 | 0 | — |
case-15 | pass→pass | 21,764 | 25,602 | +18% | 1 | 1 | 0% | 2,535 | 3,294 | +30% | 0 | 0 | — |
case-23 | pass→pass | 21,127 | 45,777 | +117% | 1 | 1 | 0% | 2,274 | 3,985 | +75% | 0 | 0 | — |
case-17 | pass→pass | 21,232 | 29,167 | +37% | 1 | 1 | 0% | 2,371 | 3,779 | +59% | 0 | 0 | — |
case-18 | pass→pass | 23,397 | 32,201 | +38% | 1 | 1 | 0% | 2,757 | 4,643 | +68% | 0 | 0 | — |
case-19 | pass→pass | 24,044 | 23,735 | -1% | 1 | 1 | 0% | 2,781 | 2,925 | +5% | 0 | 0 | — |
case-20 | pass→pass | 22,960 | 27,435 | +19% | 1 | 1 | 0% | 2,752 | 3,522 | +28% | 0 | 0 | — |
case-21 | fail→fail | 25,578 | 23,662 | -7% | 1 | 1 | 0% | 3,314 | 3,158 | -5% | 0 | 0 | — |
case-22 | pass→pass | 22,838 | 38,372 | +68% | 1 | 1 | 0% | 2,662 | 5,388 | +102% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of -100 percentage points is the difference between those two pass rates over the 23 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.