Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run CASP (8 study-type variants), JBI (~6 variants), or AMSTAR-2 quality-appraisal checklists — each ending in the tool's own required integrated judgment, not just item tallies. Also runs a proposal "rhetorical-completeness-check" mode (entry_mode="completeness_check") that instead diffs unit-classification's rhetorical labels against a target checklist's expected label set. Use this after study-design-tool-gate has dispatched to CASP/JBI/AMSTAR-2 (mode a), or directly after unit-classification
.claude/skills/yogsoth-ai-quality-appraisal-checklist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -37% | 0% |
CASP/JBI/AMSTAR-2 item-level appraisal + each tool's own required overall synthesis. Also hosts the proposal rhetorical-completeness-check as a second entry mode (see below) rather than as its own SOP file.
Subagent — spawned via spawn-agent skill.
references/item-sets.md — full item lists per tool/variant, read before drafting; kept out of this SKILL.md body per Progressive Disclosure (CASP alone has 8 variants).
This plan's header documents a correction found while planning: the pipeline graph (context/2026-08-07-13-42-sop-pipeline-graph.html) never defines rhetorical-completeness-check as a node in its nodes array — it only appears as a method label on the edge unit-classification → quality-appraisal-checklist, with the edge's own inline comment ("M13修订") describing it as an entry mode into THIS SOP, not a standalone one. The design spec (§4) listed it as a separate file, but per the spec's own stated rule that the graph is the source of truth over the transcription, this SOP's entry_mode parameter is the correct home for it — building a 5th file here would silently inflate the buildable-SOP count past the spec's own stated total of 30.
<!-- BEGIN available-tables (generated) -->
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | fail→fail | 25,793 | 13,346 | -48% | 1 | 1 | 0% | 3,442 | 1,694 | -51% | 0 | 0 | — |
case-01 | fail→fail | 18,365 | 32,665 | +78% | 1 | 1 | 0% | 2,296 | 5,370 | +134% | 0 | 0 | — |
case-02 | fail→fail | 13,550 | 27,277 | +101% | 1 | 1 | 0% | 1,556 | 4,448 | +186% | 0 | 0 | — |
case-08 | fail→pass | 20,266 | 9,786 | -52% | 1 | 1 | 0% | 2,300 | 1,236 | -46% | 0 | 0 | — |
case-03 | fail→fail | 12,806 | 13,678 | +7% | 1 | 1 | 0% | 1,260 | 1,754 | +39% | 0 | 0 | — |
case-04 | fail→fail | 11,661 | 18,569 | +59% | 1 | 1 | 0% | 1,133 | 2,976 | +163% | 0 | 0 | — |
case-05 | fail→fail | 16,394 | 7,675 | -53% | 1 | 1 | 0% | 1,762 | 846 | -52% | 0 | 0 | — |
case-06 | pass→pass | 17,232 | 14,574 | -15% | 1 | 1 | 0% | 1,944 | 1,890 | -3% | 0 | 0 | — |
case-07 | pass→pass | 15,960 | 14,618 | -8% | 1 | 1 | 0% | 1,741 | 2,011 | +16% | 0 | 0 | — |
case-14 | pass→fail | 12,661 | 13,673 | +8% | 1 | 1 | 0% | 1,339 | 1,859 | +39% | 0 | 0 | — |
case-09 | fail→fail | 18,633 | 12,077 | -35% | 1 | 1 | 0% | 2,239 | 1,601 | -28% | 0 | 0 | — |
case-10 | fail→fail | 16,929 | 11,953 | -29% | 1 | 1 | 0% | 1,780 | 1,480 | -17% | 0 | 0 | — |
case-11 | fail→fail | 13,818 | 7,045 | -49% | 1 | 1 | 0% | 1,475 | 696 | -53% | 0 | 0 | — |
case-12 | fail→pass | 17,633 | 10,596 | -40% | 1 | 1 | 0% | 1,919 | 1,472 | -23% | 0 | 0 | — |
case-13 | fail→pass | 17,429 | 10,778 | -38% | 1 | 1 | 0% | 1,858 | 1,248 | -33% | 0 | 0 | — |
case-16 | fail→pass | 18,911 | 11,545 | -39% | 1 | 1 | 0% | 2,051 | 1,560 | -24% | 0 | 0 | — |
case-17 | fail→fail | 18,028 | 6,946 | -61% | 1 | 1 | 0% | 1,964 | 674 | -66% | 0 | 0 | — |
case-18 | pass→pass | 13,330 | 8,920 | -33% | 1 | 1 | 0% | 1,312 | 973 | -26% | 0 | 0 | — |
case-19 | fail→pass | 12,578 | 7,747 | -38% | 1 | 1 | 0% | 1,287 | 815 | -37% | 0 | 0 | — |
case-20 | pass→pass | 16,484 | 18,823 | +14% | 1 | 1 | 0% | 2,269 | 3,148 | +39% | 0 | 0 | — |
case-21 | pass→pass | 22,708 | 30,270 | +33% | 1 | 1 | 0% | 4,104 | 6,146 | +50% | 0 | 0 | — |
case-22 | pass→pass | 18,204 | 24,882 | +37% | 1 | 1 | 0% | 2,229 | 3,855 | +73% | 0 | 0 | — |
case-23 | fail→fail | 19,184 | 23,641 | +23% | 1 | 1 | 0% | 3,169 | 4,230 | +33% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +17 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.