Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Non-threatening 'Why?' questioning of current practices to reveal historical accidents vs. genuine constraints.
.claude/skills/yogsoth-ai-challenge-questioning/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✓→✗ | ▼ Worse | 89% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 48% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 14% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 122% | 0% |
| case-01 | ✗→✗ | = Same ✗ | 56% | 0% |
Non-threatening 'Why?' questioning of current practices.
Subagent — spawned via subagent-spawning/spawn-agent skill.
Challenge questioning requires maintaining a non-judgmental stance while systematically probing every aspect of current practice. Benefits from dedicated context that can track which practices have been challenged and which remain.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 19,079 | 28,958 | +52% | 1 | 1 | 0% | 2,880 | 4,479 | +56% | 0 | 0 | — |
case-02 | fail→fail | 10,069 | 25,681 | +155% | 1 | 1 | 0% | 1,579 | 3,299 | +109% | 0 | 0 | — |
case-03 | fail→fail | 19,548 | 17,899 | -8% | 1 | 1 | 0% | 3,053 | 3,059 | +0% | 0 | 0 | — |
case-04 | pass→pass | 13,064 | 38,992 | +198% | 1 | 1 | 0% | 2,158 | 3,204 | +48% | 0 | 0 | — |
case-05 | pass→pass | 18,220 | 20,757 | +14% | 1 | 1 | 0% | 3,186 | 3,628 | +14% | 0 | 0 | — |
case-06 | pass→pass | 8,797 | 30,189 | +243% | 1 | 1 | 0% | 1,364 | 3,032 | +122% | 0 | 0 | — |
case-07 | pass→fail | 9,422 | 16,788 | +78% | 1 | 1 | 0% | 1,536 | 2,909 | +89% | 0 | 0 | — |
case-08 | fail→fail | 11,527 | 22,549 | +96% | 1 | 1 | 0% | 1,545 | 3,783 | +145% | 0 | 0 | — |
case-09 | fail→fail | 18,126 | 17,845 | -2% | 1 | 1 | 0% | 839 | 2,771 | +230% | 0 | 0 | — |
case-10 | fail→fail | 6,995 | 11,253 | +61% | 1 | 1 | 0% | 1,015 | 1,605 | +58% | 0 | 0 | — |
case-11 | fail→fail | 10,610 | 29,665 | +180% | 1 | 1 | 0% | 1,607 | 2,501 | +56% | 0 | 0 | — |
case-12 | fail→fail | 19,645 | 42,445 | +116% | 1 | 1 | 0% | 2,671 | 5,681 | +113% | 0 | 0 | — |
case-13 | fail→fail | 13,761 | 16,740 | +22% | 1 | 1 | 0% | 2,057 | 2,659 | +29% | 0 | 0 | — |
case-14 | fail→fail | 13,572 | 30,150 | +122% | 1 | 1 | 0% | 1,818 | 3,100 | +71% | 0 | 0 | — |
case-15 | fail→fail | 13,090 | 24,947 | +91% | 1 | 1 | 0% | 1,866 | 3,776 | +102% | 0 | 0 | — |
case-16 | fail→fail | 14,593 | 32,794 | +125% | 1 | 1 | 0% | 1,929 | 4,487 | +133% | 0 | 0 | — |
case-17 | fail→fail | 15,001 | 16,036 | +7% | 1 | 1 | 0% | 1,988 | 2,638 | +33% | 0 | 0 | — |
case-18 | fail→fail | 13,416 | 40,016 | +198% | 1 | 1 | 0% | 1,964 | 3,580 | +82% | 0 | 0 | — |
case-19 | fail→fail | 14,536 | 15,904 | +9% | 1 | 1 | 0% | 1,894 | 2,260 | +19% | 0 | 0 | — |
case-20 | fail→fail | 14,120 | 22,740 | +61% | 1 | 1 | 0% | 2,086 | 1,851 | -11% | 0 | 0 | — |
case-21 | fail→fail | 14,750 | 16,206 | +10% | 1 | 1 | 0% | 2,012 | 2,183 | +8% | 0 | 0 | — |
case-22 | fail→fail | 7,957 | 17,229 | +117% | 1 | 1 | 0% | 1,054 | 2,670 | +153% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Other measured skills in the registry, with their headline benchmark lift.