Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluate pairwise value consistency (logical/empirical/normative)
.claude/skills/yogsoth-ai-creative-ideation-consistency-pair-evaluation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-20 | ✓→✓ | = Same ✓ | -47% | 0% |
| case-21 | ✓→✓ | = Same ✓ | 35% | 0% |
| case-22 | ✓→✓ | = Same ✓ | 5% | 0% |
| case-07 | ✗→✗ | = Same ✗ | -18% | 0% |
| case-01 | ✗→✗ | = Same ✗ | 224% | 0% |
Evaluate pairwise consistency of parameter-value combinations, classifying inconsistencies by type.
Subagent — spawned via subagent-spawning/spawn-agent skill.
Consistency evaluation requires careful reasoning about whether two values from different parameters can coexist, considering logical constraints, empirical evidence, and normative conventions.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | pass→pass | 12,549 | 5,744 | -54% | 1 | 1 | 0% | 2,136 | 1,128 | -47% | 0 | 0 | — |
case-07 | fail→fail | 14,373 | 10,471 | -27% | 1 | 1 | 0% | 2,315 | 1,900 | -18% | 0 | 0 | — |
case-01 | fail→fail | 9,418 | 31,605 | +236% | 1 | 1 | 0% | 1,409 | 4,565 | +224% | 0 | 0 | — |
case-02 | fail→fail | 15,381 | 5,488 | -64% | 1 | 1 | 0% | 2,811 | 377 | -87% | 0 | 0 | — |
case-03 | fail→fail | 11,757 | 13,345 | +14% | 1 | 1 | 0% | 1,871 | 2,129 | +14% | 0 | 0 | — |
case-04 | fail→fail | 12,609 | 12,612 | +0% | 1 | 1 | 0% | 1,981 | 2,123 | +7% | 0 | 0 | — |
case-05 | fail→fail | 12,634 | 18,831 | +49% | 1 | 1 | 0% | 2,063 | 1,666 | -19% | 0 | 0 | — |
case-06 | fail→fail | 14,637 | 23,050 | +57% | 1 | 1 | 0% | 2,463 | 3,141 | +28% | 0 | 0 | — |
case-08 | fail→fail | 14,230 | 18,545 | +30% | 1 | 1 | 0% | 2,299 | 3,035 | +32% | 0 | 0 | — |
case-09 | fail→fail | 8,906 | 10,106 | +13% | 1 | 1 | 0% | 1,579 | 1,750 | +11% | 0 | 0 | — |
case-10 | fail→fail | 12,650 | 10,097 | -20% | 1 | 1 | 0% | 2,218 | 1,959 | -12% | 0 | 0 | — |
case-11 | fail→fail | 11,633 | 12,433 | +7% | 1 | 1 | 0% | 1,716 | 1,879 | +9% | 0 | 0 | — |
case-12 | fail→fail | 14,035 | 8,547 | -39% | 1 | 1 | 0% | 2,234 | 1,506 | -33% | 0 | 0 | — |
case-13 | fail→fail | 11,059 | 16,395 | +48% | 1 | 1 | 0% | 1,744 | 1,388 | -20% | 0 | 0 | — |
case-14 | fail→fail | 14,386 | 12,822 | -11% | 1 | 1 | 0% | 2,212 | 2,192 | -1% | 0 | 0 | — |
case-15 | fail→fail | 12,841 | 20,275 | +58% | 1 | 1 | 0% | 2,161 | 3,549 | +64% | 0 | 0 | — |
case-16 | fail→fail | 6,675 | 9,356 | +40% | 1 | 1 | 0% | 1,058 | 1,697 | +60% | 0 | 0 | — |
case-17 | fail→fail | 11,486 | 16,104 | +40% | 1 | 1 | 0% | 1,958 | 2,101 | +7% | 0 | 0 | — |
case-18 | fail→fail | 13,869 | 8,096 | -42% | 1 | 1 | 0% | 2,130 | 1,533 | -28% | 0 | 0 | — |
case-19 | fail→fail | 10,532 | 14,306 | +36% | 1 | 1 | 0% | 1,741 | 2,023 | +16% | 0 | 0 | — |
case-21 | pass→pass | 2,403 | 2,776 | +16% | 1 | 1 | 0% | 507 | 685 | +35% | 0 | 0 | — |
case-22 | pass→pass | 5,554 | 4,991 | -10% | 1 | 1 | 0% | 1,155 | 1,208 | +5% | 0 | 0 | — |
case-23 | fail→fail | 1,655 | 2,021 | +22% | 1 | 1 | 0% | 209 | 439 | +110% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 22 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.