Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Catalog all known solutions/methods in a domain with performance, applicability, and limitations.
.claude/skills/yogsoth-ai-creative-ideation-benchmark-inventory/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-20 | ✓→✓ | = Same ✓ | 23% | 0% |
| case-21 | ✓→✓ | = Same ✓ | -7% | 0% |
| case-22 | ✓→✓ | = Same ✓ | 16% | 0% |
| case-23 | ✓→✓ | = Same ✓ | 62% | 0% |
| case-18 | ✗→✗ | = Same ✗ | 41% | 0% |
Catalog all known solutions/methods in a domain. Produces a structured inventory with performance metrics, applicability scope, and known limitations.
Subagent — spawned via subagent-spawning/spawn-agent skill.
Comprehensive benchmarking requires deep, focused research across multiple sources (papers, benchmarks, surveys). Benefits from dedicated context that can accumulate findings without polluting the orchestrator's state.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | fail→fail | 26,140 | 32,914 | +26% | 1 | 1 | 0% | 3,901 | 5,483 | +41% | 0 | 0 | — |
case-05 | fail→fail | 27,928 | 36,965 | +32% | 1 | 1 | 0% | 4,535 | 6,365 | +40% | 0 | 0 | — |
case-12 | fail→fail | 29,920 | 34,650 | +16% | 1 | 1 | 0% | 5,117 | 5,707 | +12% | 0 | 0 | — |
case-01 | fail→fail | 30,907 | 12,289 | -60% | 1 | 1 | 0% | 4,923 | 525 | -89% | 0 | 0 | — |
case-02 | fail→fail | 27,705 | 28,745 | +4% | 1 | 1 | 0% | 5,289 | 6,362 | +20% | 0 | 0 | — |
case-03 | fail→fail | 33,833 | 32,523 | -4% | 1 | 1 | 0% | 5,791 | 6,359 | +10% | 0 | 0 | — |
case-04 | fail→fail | 27,094 | 29,349 | +8% | 1 | 1 | 0% | 4,577 | 5,425 | +19% | 0 | 0 | — |
case-06 | fail→fail | 27,170 | 34,863 | +28% | 1 | 1 | 0% | 4,483 | 5,274 | +18% | 0 | 0 | — |
case-07 | fail→fail | 23,221 | 34,031 | +47% | 1 | 1 | 0% | 4,163 | 6,360 | +53% | 0 | 0 | — |
case-08 | fail→fail | 26,412 | 37,824 | +43% | 1 | 1 | 0% | 4,440 | 6,352 | +43% | 0 | 0 | — |
case-09 | fail→fail | 31,806 | 33,218 | +4% | 1 | 1 | 0% | 5,572 | 5,879 | +6% | 0 | 0 | — |
case-10 | fail→fail | 25,474 | 35,921 | +41% | 1 | 1 | 0% | 4,224 | 6,354 | +50% | 0 | 0 | — |
case-11 | fail→fail | 30,937 | 38,492 | +24% | 1 | 1 | 0% | 4,731 | 6,244 | +32% | 0 | 0 | — |
case-13 | fail→fail | 23,861 | 15,388 | -36% | 1 | 1 | 0% | 3,911 | 694 | -82% | 0 | 0 | — |
case-14 | fail→fail | 24,332 | 30,850 | +27% | 1 | 1 | 0% | 4,063 | 5,264 | +30% | 0 | 0 | — |
case-15 | fail→fail | 26,790 | 99,458 | +271% | 1 | 1 | 0% | 4,656 | 6,341 | +36% | 0 | 0 | — |
case-16 | fail→fail | 30,455 | 41,982 | +38% | 1 | 1 | 0% | 4,270 | 6,178 | +45% | 0 | 0 | — |
case-17 | fail→fail | 28,886 | 38,541 | +33% | 1 | 1 | 0% | 4,698 | 6,346 | +35% | 0 | 0 | — |
case-19 | fail→fail | 38,127 | 38,063 | -0% | 1 | 1 | 0% | 6,063 | 6,117 | +1% | 0 | 0 | — |
case-20 | pass→pass | 24,259 | 29,306 | +21% | 1 | 1 | 0% | 3,545 | 4,362 | +23% | 0 | 0 | — |
case-21 | pass→pass | 16,354 | 13,966 | -15% | 1 | 1 | 0% | 3,159 | 2,929 | -7% | 0 | 0 | — |
case-22 | pass→pass | 22,211 | 23,946 | +8% | 1 | 1 | 0% | 4,321 | 4,998 | +16% | 0 | 0 | — |
case-23 | pass→pass | 9,839 | 14,721 | +50% | 1 | 1 | 0% | 1,701 | 2,763 | +62% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.