Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Determine appropriate levels for each experimental factor
.claude/skills/yogsoth-ai-level-specification/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-21 | ✓→✗ | ▼ Worse | 20% | 0% |
| case-20 | ✓→✓ | = Same ✓ | 31% | 0% |
| case-22 | ✓→✓ | = Same ✓ | 39% | 0% |
| case-01 | ✗→✗ | = Same ✗ | 10% | 0% |
| case-02 | ✗→✗ | = Same ✗ | 32% | 0% |
Determine the specific values (levels) at which each factor will be tested, including spacing strategy and justification.
Subagent — spawned via subagent-spawning/spawn-agent skill.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 27,087 | 23,469 | -13% | 1 | 1 | 0% | 3,087 | 3,400 | +10% | 0 | 0 | — |
case-02 | fail→fail | 15,616 | 12,777 | -18% | 1 | 1 | 0% | 2,034 | 2,684 | +32% | 0 | 0 | — |
case-03 | fail→fail | 10,149 | 13,683 | +35% | 1 | 1 | 0% | 1,823 | 2,413 | +32% | 0 | 0 | — |
case-04 | fail→fail | 11,148 | 11,371 | +2% | 1 | 1 | 0% | 2,246 | 2,230 | -1% | 0 | 0 | — |
case-10 | fail→fail | 19,961 | 17,540 | -12% | 1 | 1 | 0% | 2,691 | 2,493 | -7% | 0 | 0 | — |
case-05 | fail→fail | 16,682 | 8,859 | -47% | 1 | 1 | 0% | 2,046 | 1,699 | -17% | 0 | 0 | — |
case-06 | fail→fail | 25,678 | 22,213 | -13% | 1 | 1 | 0% | 3,303 | 2,917 | -12% | 0 | 0 | — |
case-07 | fail→fail | 17,572 | 23,301 | +33% | 1 | 1 | 0% | 3,322 | 3,402 | +2% | 0 | 0 | — |
case-08 | fail→fail | 19,439 | 14,965 | -23% | 1 | 1 | 0% | 2,630 | 2,715 | +3% | 0 | 0 | — |
case-09 | fail→fail | 12,312 | 16,846 | +37% | 1 | 1 | 0% | 2,303 | 2,325 | +1% | 0 | 0 | — |
case-11 | fail→fail | 20,083 | 33,696 | +68% | 1 | 1 | 0% | 3,061 | 1,894 | -38% | 0 | 0 | — |
case-12 | fail→fail | 13,356 | 11,375 | -15% | 1 | 1 | 0% | 1,495 | 1,132 | -24% | 0 | 0 | — |
case-13 | fail→fail | 17,615 | 19,355 | +10% | 1 | 1 | 0% | 2,265 | 2,625 | +16% | 0 | 0 | — |
case-14 | fail→fail | 16,547 | 18,262 | +10% | 1 | 1 | 0% | 2,063 | 2,409 | +17% | 0 | 0 | — |
case-15 | fail→fail | 19,994 | 20,744 | +4% | 1 | 1 | 0% | 2,579 | 3,809 | +48% | 0 | 0 | — |
case-16 | fail→fail | 20,956 | 19,286 | -8% | 1 | 1 | 0% | 2,757 | 2,465 | -11% | 0 | 0 | — |
case-17 | fail→fail | 19,630 | 12,919 | -34% | 1 | 1 | 0% | 2,507 | 2,377 | -5% | 0 | 0 | — |
case-18 | fail→fail | 14,026 | 15,634 | +11% | 1 | 1 | 0% | 2,506 | 1,911 | -24% | 0 | 0 | — |
case-19 | fail→fail | 19,380 | 19,460 | +0% | 1 | 1 | 0% | 2,426 | 2,385 | -2% | 0 | 0 | — |
case-20 | pass→pass | 22,591 | 27,368 | +21% | 1 | 1 | 0% | 2,934 | 3,833 | +31% | 0 | 0 | — |
case-21 | pass→fail | 18,814 | 20,346 | +8% | 1 | 1 | 0% | 2,510 | 3,019 | +20% | 0 | 0 | — |
case-22 | pass→pass | 30,936 | 91,708 | +196% | 1 | 1 | 0% | 6,055 | 8,414 | +39% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of -100 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Other measured skills in the registry, with their headline benchmark lift.