Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluate new combinations for feasibility and novelty
.claude/skills/yogsoth-ai-combination-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-22 | ✓→✗ | ▼ Worse | 31% | 0% |
| case-23 | ✓→✓ | = Same ✓ | -8% | 0% |
| case-20 | ✓→✓ | = Same ✓ | 124% | 0% |
| case-21 | ✓→✓ | = Same ✓ | 12% | 0% |
| case-03 | ✗→✗ | = Same ✗ | 134% | 0% |
Evaluate proposed combinations for feasibility, novelty, and implementation difficulty.
Subagent — spawned via subagent-spawning/spawn-agent skill.
Combination evaluation requires multi-criteria assessment of each proposed configuration, weighing technical feasibility against novelty and estimating implementation barriers.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 7,779 | 45,186 | +481% | 1 | 1 | 0% | 1,128 | 2,636 | +134% | 0 | 0 | — |
case-04 | fail→fail | 6,483 | 14,708 | +127% | 1 | 1 | 0% | 956 | 2,590 | +171% | 0 | 0 | — |
case-01 | fail→fail | 38,650 | 73,092 | +89% | 1 | 1 | 0% | 1,280 | 1,032 | -19% | 0 | 0 | — |
case-02 | fail→fail | 7,741 | 28,871 | +273% | 1 | 1 | 0% | 1,217 | 4,778 | +293% | 0 | 0 | — |
case-23 | pass→pass | 10,222 | 8,341 | -18% | 1 | 1 | 0% | 2,007 | 1,841 | -8% | 0 | 0 | — |
case-05 | fail→fail | 5,985 | 35,088 | +486% | 1 | 1 | 0% | 954 | 916 | -4% | 0 | 0 | — |
case-06 | fail→fail | 17,781 | 16,298 | -8% | 1 | 1 | 0% | 2,917 | 2,781 | -5% | 0 | 0 | — |
case-07 | fail→fail | 17,820 | 18,657 | +5% | 1 | 1 | 0% | 2,802 | 2,693 | -4% | 0 | 0 | — |
case-08 | fail→fail | 7,088 | 15,874 | +124% | 1 | 1 | 0% | 940 | 1,041 | +11% | 0 | 0 | — |
case-09 | fail→fail | 10,573 | 277,161 | +2521% | 1 | 1 | 0% | 1,605 | 2,621 | +63% | 0 | 0 | — |
case-10 | fail→fail | 12,410 | 22,609 | +82% | 1 | 1 | 0% | 1,829 | 3,075 | +68% | 0 | 0 | — |
case-11 | fail→fail | 6,562 | 8,284 | +26% | 1 | 1 | 0% | 1,038 | 1,462 | +41% | 0 | 0 | — |
case-12 | fail→fail | 7,670 | 7,186 | -6% | 1 | 1 | 0% | 1,222 | 1,223 | +0% | 0 | 0 | — |
case-13 | fail→fail | 10,142 | 26,033 | +157% | 1 | 1 | 0% | 1,464 | 3,679 | +151% | 0 | 0 | — |
case-14 | fail→fail | 9,096 | 12,685 | +39% | 1 | 1 | 0% | 1,348 | 2,254 | +67% | 0 | 0 | — |
case-15 | fail→fail | 10,143 | 19,323 | +91% | 1 | 1 | 0% | 1,567 | 3,094 | +97% | 0 | 0 | — |
case-16 | fail→fail | 12,795 | 11,286 | -12% | 1 | 1 | 0% | 1,987 | 1,985 | -0% | 0 | 0 | — |
case-17 | fail→fail | 10,694 | 11,857 | +11% | 1 | 1 | 0% | 1,605 | 1,949 | +21% | 0 | 0 | — |
case-18 | fail→fail | 3,626 | 5,120 | +41% | 1 | 1 | 0% | 552 | 981 | +78% | 0 | 0 | — |
case-19 | fail→fail | 5,540 | 13,057 | +136% | 1 | 1 | 0% | 905 | 2,222 | +146% | 0 | 0 | — |
case-20 | pass→pass | 17,943 | 42,360 | +136% | 1 | 1 | 0% | 2,753 | 6,163 | +124% | 0 | 0 | — |
case-21 | pass→pass | 11,874 | 12,895 | +9% | 1 | 1 | 0% | 2,104 | 2,364 | +12% | 0 | 0 | — |
case-22 | pass→fail | 19,341 | 24,437 | +26% | 1 | 1 | 0% | 3,007 | 3,942 | +31% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -4 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Other measured skills in the registry, with their headline benchmark lift.