Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions
.claude/skills/mkurman-using-superpowers/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 362% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 322% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -43% | 0% |
---|---------| | "This is just a simple question" | Questions are tasks. Check for skills. | | "I need more context first" | Skill check comes BEFORE clarifying questions. | | "Let me explore the codebase first" | Skills tell you HOW to explore. Check first. | | "I can check git/files quickly" | Files lack conversation context. Check for skills. | | "Let me gather information first" | Skills tell you HOW to gather information. | | "This doesn't need a formal skill" | If a skill exists, use it. | | "I remember this skill" | Skills evolve. Read current version. | | "This doesn't count as a task" | Action = task. Check for skills. | | "The skill is overkill" | Simple things become complex. Use it. | | "I'll just do this one thing first" | Check BEFORE doing anything. | | "This feels productive" | Undisciplined action wastes time. Skills prevent this. | | "I know what that means" | Knowing the concept ≠ using the skill. Invoke it. |
When multiple skills could apply, use this order:
"Let's build X" → brainstorming first, then implementation skills. "Fix this bug" → debugging first, then domain-specific skills.
Rigid (TDD, debugging): Follow exactly. Don't adapt away discipline.
Flexible (patterns): Adapt principles to context.
The skill itself tells you which.
Instructions say WHAT, not HOW. "Add X" or "Fix Y" doesn't mean skip workflows.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 33,761 | 30,715 | -9% | 1 | 1 | 0% | 6,173 | 6,585 | +7% | 0 | 0 | — |
case-02 | fail→fail | 19,863 | 4,077 | -79% | 1 | 1 | 0% | 3,414 | 811 | -76% | 0 | 0 | — |
case-03 | fail→fail | 14,258 | 12,474 | -13% | 1 | 1 | 0% | 2,523 | 2,622 | +4% | 0 | 0 | — |
case-04 | fail→pass | 4,247 | 15,492 | +265% | 1 | 1 | 0% | 601 | 2,776 | +362% | 0 | 0 | — |
case-05 | fail→fail | 6,733 | 6,358 | -6% | 1 | 1 | 0% | 1,198 | 1,255 | +5% | 0 | 0 | — |
case-06 | pass→pass | 12,111 | 13,290 | +10% | 1 | 1 | 0% | 2,013 | 2,403 | +19% | 0 | 0 | — |
case-07 | fail→fail | 2,908 | 10,925 | +276% | 1 | 1 | 0% | 468 | 1,976 | +322% | 0 | 0 | — |
case-08 | fail→pass | 3,500 | 10,342 | +195% | 1 | 1 | 0% | 534 | 2,252 | +322% | 0 | 0 | — |
case-09 | fail→fail | 6,332 | 9,111 | +44% | 1 | 1 | 0% | 1,116 | 2,066 | +85% | 0 | 0 | — |
case-10 | pass→pass | 10,306 | 8,688 | -16% | 1 | 1 | 0% | 1,765 | 1,125 | -36% | 0 | 0 | — |
case-11 | pass→pass | 15,084 | 10,438 | -31% | 1 | 1 | 0% | 2,166 | 1,740 | -20% | 0 | 0 | — |
case-12 | fail→fail | 15,070 | 32,563 | +116% | 1 | 1 | 0% | 3,254 | 4,225 | +30% | 0 | 0 | — |
case-13 | fail→fail | 28,833 | 33,209 | +15% | 1 | 1 | 0% | 6,170 | 6,855 | +11% | 0 | 0 | — |
case-14 | fail→fail | 13,276 | 11,226 | -15% | 1 | 1 | 0% | 2,273 | 2,077 | -9% | 0 | 0 | — |
case-15 | fail→pass | 5,783 | 8,248 | +43% | 1 | 1 | 0% | 908 | 1,507 | +66% | 0 | 0 | — |
case-16 | fail→pass | 4,921 | 5,843 | +19% | 1 | 1 | 0% | 863 | 1,227 | +42% | 0 | 0 | — |
case-17 | pass→pass | 5,847 | 2,367 | -60% | 1 | 1 | 0% | 879 | 737 | -16% | 0 | 0 | — |
case-18 | fail→pass | 12,172 | 4,297 | -65% | 1 | 1 | 0% | 1,839 | 1,056 | -43% | 0 | 0 | — |
case-19 | pass→pass | 7,727 | 6,612 | -14% | 1 | 1 | 0% | 1,045 | 1,351 | +29% | 0 | 0 | — |
case-20 | pass→pass | 8,503 | 4,967 | -42% | 1 | 1 | 0% | 1,608 | 1,201 | -25% | 0 | 0 | — |
case-21 | pass→pass | 8,725 | 4,513 | -48% | 1 | 1 | 0% | 1,541 | 1,168 | -24% | 0 | 0 | — |
case-22 | pass→pass | 2,044 | 2,282 | +12% | 1 | 1 | 0% | 322 | 739 | +130% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.