Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Assess evidence quality using GRADE/SOE framework. Rates certainty level and identifies downgrade reasons.
.claude/skills/yogsoth-ai-evidence-grading/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✓→✗ | ▼ Worse | 67% | 0% |
| case-03 | ✓→✓ | = Same ✓ | -5% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 16% | 0% |
| case-02 | ✗→✗ | = Same ✗ | -5% | 0% |
| case-10 | ✗→✗ | = Same ✗ | 15% | 0% |
Assess the quality and certainty of evidence supporting or surrounding a gap.
Subagent — spawned via subagent-spawning/spawn-agent.
One unit = one evidence grading assessment.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. Used by SOPs that declare execution: subagent. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→fail | 14,322 | 12,316 | -14% | 1 | 1 | 0% | 2,150 | 2,042 | -5% | 0 | 0 | — |
case-03 | pass→pass | 18,391 | 17,608 | -4% | 1 | 1 | 0% | 2,883 | 2,731 | -5% | 0 | 0 | — |
case-04 | pass→pass | 22,772 | 23,202 | +2% | 1 | 1 | 0% | 3,788 | 4,397 | +16% | 0 | 0 | — |
case-10 | fail→fail | 8,578 | 10,697 | +25% | 1 | 1 | 0% | 1,376 | 1,588 | +15% | 0 | 0 | — |
case-11 | fail→fail | 9,099 | 14,554 | +60% | 1 | 1 | 0% | 1,327 | 2,415 | +82% | 0 | 0 | — |
case-01 | fail→fail | 10,569 | 25,847 | +145% | 1 | 1 | 0% | 1,411 | 411 | -71% | 0 | 0 | — |
case-05 | pass→fail | 14,633 | 23,197 | +59% | 1 | 1 | 0% | 2,115 | 3,541 | +67% | 0 | 0 | — |
case-06 | fail→fail | 8,393 | 3,264 | -61% | 1 | 1 | 0% | 1,272 | 548 | -57% | 0 | 0 | — |
case-07 | fail→fail | 13,274 | 9,203 | -31% | 1 | 1 | 0% | 1,835 | 1,042 | -43% | 0 | 0 | — |
case-08 | fail→fail | 9,937 | 27,854 | +180% | 1 | 1 | 0% | 1,454 | 3,474 | +139% | 0 | 0 | — |
case-09 | fail→fail | 16,594 | 27,755 | +67% | 1 | 1 | 0% | 2,565 | 2,783 | +8% | 0 | 0 | — |
case-12 | fail→fail | 10,002 | 12,745 | +27% | 1 | 1 | 0% | 1,640 | 2,065 | +26% | 0 | 0 | — |
case-13 | fail→fail | 9,562 | 22,423 | +135% | 1 | 1 | 0% | 1,349 | 3,498 | +159% | 0 | 0 | — |
case-14 | fail→fail | 19,111 | 17,868 | -7% | 1 | 1 | 0% | 2,533 | 1,248 | -51% | 0 | 0 | — |
case-15 | fail→fail | 6,644 | 20,162 | +203% | 1 | 1 | 0% | 990 | 3,018 | +205% | 0 | 0 | — |
case-16 | fail→fail | 17,003 | 6,087 | -64% | 1 | 1 | 0% | 2,431 | 878 | -64% | 0 | 0 | — |
case-17 | fail→fail | 9,921 | 13,063 | +32% | 1 | 1 | 0% | 1,576 | 2,115 | +34% | 0 | 0 | — |
case-18 | fail→fail | 22,459 | 13,551 | -40% | 1 | 1 | 0% | 3,254 | 444 | -86% | 0 | 0 | — |
case-19 | fail→fail | 20,642 | 21,642 | +5% | 1 | 1 | 0% | 1,967 | 2,014 | +2% | 0 | 0 | — |
case-20 | fail→fail | 18,341 | 18,247 | -1% | 1 | 1 | 0% | 2,701 | 2,788 | +3% | 0 | 0 | — |
case-21 | fail→fail | 16,880 | 11,132 | -34% | 1 | 1 | 0% | 2,610 | 1,776 | -32% | 0 | 0 | — |
case-22 | fail→fail | 16,774 | 7,637 | -54% | 1 | 1 | 0% | 2,494 | 458 | -82% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 19 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.