Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Select appropriate baselines for experimental comparison
.claude/skills/yogsoth-ai-baseline-selection/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-18 | ✓→✗ | ▼ Worse | -25% | 0% |
| case-04 | ✓→✓ | = Same ✓ | -7% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 9% | 0% |
| case-01 | ✓→✓ | = Same ✓ | -38% | 0% |
| case-02 | ✓→✓ | = Same ✓ | -31% | 0% |
Select appropriate baselines that provide meaningful comparison points, covering SOTA, simple, and internal baselines.
Subagent — spawned via subagent-spawning/spawn-agent skill.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 13,861 | 12,840 | -7% | 1 | 1 | 0% | 2,234 | 2,082 | -7% | 0 | 0 | — |
case-05 | pass→pass | 15,704 | 17,171 | +9% | 1 | 1 | 0% | 2,575 | 2,814 | +9% | 0 | 0 | — |
case-01 | pass→pass | 18,081 | 10,875 | -40% | 1 | 1 | 0% | 2,837 | 1,770 | -38% | 0 | 0 | — |
case-02 | pass→pass | 16,464 | 10,618 | -36% | 1 | 1 | 0% | 2,541 | 1,748 | -31% | 0 | 0 | — |
case-03 | pass→pass | 16,222 | 12,810 | -21% | 1 | 1 | 0% | 2,815 | 2,256 | -20% | 0 | 0 | — |
case-06 | fail→fail | 18,738 | 12,388 | -34% | 1 | 1 | 0% | 3,020 | 2,082 | -31% | 0 | 0 | — |
case-07 | pass→pass | 17,988 | 12,184 | -32% | 1 | 1 | 0% | 2,669 | 1,956 | -27% | 0 | 0 | — |
case-08 | fail→fail | 16,254 | 10,206 | -37% | 1 | 1 | 0% | 2,506 | 1,572 | -37% | 0 | 0 | — |
case-09 | fail→fail | 16,703 | 11,186 | -33% | 1 | 1 | 0% | 2,546 | 1,724 | -32% | 0 | 0 | — |
case-10 | fail→fail | 23,295 | 16,883 | -28% | 1 | 1 | 0% | 3,608 | 2,786 | -23% | 0 | 0 | — |
case-11 | fail→fail | 17,276 | 13,024 | -25% | 1 | 1 | 0% | 2,645 | 2,172 | -18% | 0 | 0 | — |
case-12 | fail→fail | 18,798 | 13,322 | -29% | 1 | 1 | 0% | 2,890 | 2,049 | -29% | 0 | 0 | — |
case-13 | pass→pass | 18,960 | 15,760 | -17% | 1 | 1 | 0% | 2,953 | 2,411 | -18% | 0 | 0 | — |
case-14 | pass→pass | 20,129 | 14,932 | -26% | 1 | 1 | 0% | 3,032 | 2,260 | -25% | 0 | 0 | — |
case-15 | fail→fail | 17,300 | 15,170 | -12% | 1 | 1 | 0% | 2,538 | 2,402 | -5% | 0 | 0 | — |
case-16 | fail→fail | 17,138 | 14,050 | -18% | 1 | 1 | 0% | 2,713 | 2,373 | -13% | 0 | 0 | — |
case-17 | fail→fail | 22,868 | 18,187 | -20% | 1 | 1 | 0% | 3,692 | 2,861 | -23% | 0 | 0 | — |
case-18 | pass→fail | 16,472 | 12,221 | -26% | 1 | 1 | 0% | 2,624 | 1,966 | -25% | 0 | 0 | — |
case-19 | fail→fail | 17,904 | 12,335 | -31% | 1 | 1 | 0% | 2,806 | 2,081 | -26% | 0 | 0 | — |
case-20 | fail→fail | 21,131 | 16,833 | -20% | 1 | 1 | 0% | 3,296 | 2,747 | -17% | 0 | 0 | — |
case-21 | fail→fail | 18,669 | 13,324 | -29% | 1 | 1 | 0% | 2,600 | 2,109 | -19% | 0 | 0 | — |
case-22 | fail→fail | 17,163 | 11,865 | -31% | 1 | 1 | 0% | 2,703 | 1,790 | -34% | 0 | 0 | — |
case-23 | fail→fail | 17,479 | 13,332 | -24% | 1 | 1 | 0% | 2,601 | 2,154 | -17% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of -4 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Other measured skills in the registry, with their headline benchmark lift.