Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Pairwise consistency assessment using Cross-Consistency Assessment (CCA) matrix
.claude/skills/yogsoth-ai-experiment-execution-consistency-pair-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-21 | ✓→✗ | ▼ Worse | 10% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 15% | 0% |
| case-20 | ✓→✓ | = Same ✓ | 24% | 0% |
| case-01 | ✗→✗ | = Same ✗ | 255% | 0% |
| case-02 | ✗→✗ | = Same ✗ | 92% | 0% |
Evaluate all pairwise combinations of parameter values for internal consistency using Cross-Consistency Assessment (CCA). Produce a matrix that enables filtering of the morphological field to retain only plausible configurations.
Subagent — spawned via subagent-spawning/spawn-agent skill.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 9,375 | 27,523 | +194% | 1 | 1 | 0% | 1,518 | 5,395 | +255% | 0 | 0 | — |
case-02 | fail→fail | 8,232 | 15,537 | +89% | 1 | 1 | 0% | 1,488 | 2,857 | +92% | 0 | 0 | — |
case-14 | fail→fail | 21,117 | 23,797 | +13% | 1 | 1 | 0% | 3,832 | 4,651 | +21% | 0 | 0 | — |
case-15 | fail→fail | 19,993 | 18,225 | -9% | 1 | 1 | 0% | 3,950 | 3,576 | -9% | 0 | 0 | — |
case-03 | fail→fail | 9,084 | 51,059 | +462% | 1 | 1 | 0% | 1,392 | 3,384 | +143% | 0 | 0 | — |
case-04 | fail→fail | 24,648 | 21,665 | -12% | 1 | 1 | 0% | 4,615 | 3,851 | -17% | 0 | 0 | — |
case-05 | fail→fail | 21,608 | 24,926 | +15% | 1 | 1 | 0% | 3,832 | 4,973 | +30% | 0 | 0 | — |
case-06 | fail→fail | 42,385 | 21,185 | -50% | 1 | 1 | 0% | 3,838 | 4,321 | +13% | 0 | 0 | — |
case-07 | fail→fail | 21,058 | 14,348 | -32% | 1 | 1 | 0% | 3,821 | 2,713 | -29% | 0 | 0 | — |
case-08 | fail→fail | 28,809 | 30,778 | +7% | 1 | 1 | 0% | 5,080 | 6,318 | +24% | 0 | 0 | — |
case-09 | fail→fail | 42,856 | 24,288 | -43% | 1 | 1 | 0% | 2,191 | 4,667 | +113% | 0 | 0 | — |
case-10 | fail→fail | 40,258 | 22,268 | -45% | 1 | 1 | 0% | 2,029 | 4,406 | +117% | 0 | 0 | — |
case-11 | fail→fail | 28,941 | 25,626 | -11% | 1 | 1 | 0% | 6,170 | 4,994 | -19% | 0 | 0 | — |
case-12 | fail→fail | 36,273 | 15,609 | -57% | 1 | 1 | 0% | 6,083 | 3,013 | -50% | 0 | 0 | — |
case-13 | fail→fail | 20,925 | 20,289 | -3% | 1 | 1 | 0% | 3,735 | 3,902 | +4% | 0 | 0 | — |
case-16 | fail→fail | 32,676 | 16,967 | -48% | 1 | 1 | 0% | 5,498 | 3,171 | -42% | 0 | 0 | — |
case-17 | fail→fail | 25,714 | 20,211 | -21% | 1 | 1 | 0% | 5,354 | 3,479 | -35% | 0 | 0 | — |
case-18 | fail→fail | 10,129 | 12,134 | +20% | 1 | 1 | 0% | 1,900 | 2,294 | +21% | 0 | 0 | — |
case-19 | fail→fail | 30,931 | 21,211 | -31% | 1 | 1 | 0% | 5,347 | 4,038 | -24% | 0 | 0 | — |
case-20 | pass→pass | 14,803 | 19,156 | +29% | 1 | 1 | 0% | 2,593 | 3,222 | +24% | 0 | 0 | — |
case-21 | pass→fail | 22,134 | 22,874 | +3% | 1 | 1 | 0% | 4,141 | 4,549 | +10% | 0 | 0 | — |
case-22 | pass→fail | 20,082 | 23,145 | +15% | 1 | 1 | 0% | 3,706 | 4,251 | +15% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -9 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.