Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Classify a paper's study design (RCT, cohort, case-control, diagnostic-accuracy, systematic-review, animal-study, prediction-model, etc., or not_applicable) and dispatch to the correct downstream bias-risk/quality/reporting tool and specific variant (CASP has 8 variants, JBI ~6, RoB2 has parallel/cluster/crossover versions). Use this as the mandatory first step before running ANY of CASP, JBI, AMSTAR-2, NOS, RoB2, ROBINS-I, QUADAS-2, CONSORT, STROBE, ARRIVE, SPIRIT, TRIPOD, or engineering-config
.claude/skills/yogsoth-ai-study-design-tool-gate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 637% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-20 | ✗→✓ | ▲ Improved | -70% | 0% |
| case-09 | ✓→✗ | ▼ Worse | 14% | 0% |
Classifies study design and dispatches to the right bias-risk/quality/reporting tool + variant — or determines none applies. Added per coverage-audit M11: the original graph had no node representing this dispatch decision at all; every A1/A2 tool was drawn as if it started with no gate.
Subagent — spawned via spawn-agent skill.
references/tool-dispatch-table.md — the full dispatch table (every study_design → tool + variant mapping). Read before drafting the prompt's decision, not summarized inline here (kept out of this SKILL.md body per Progressive Disclosure).
This is worth restating: most of these tools carry medical/clinical assumptions baked into their domains, and forcing a dispatch onto a paper that has no matching study design produces a meaningless result, not a conservative one. Do not treat a high not_applicable rate across a batch of CS/ML papers as a sign this SOP is failing to trigger correctly.
<!-- BEGIN available-tables (generated) -->
| SOP | When to use | | --- | --- | | spawn-agent | Spawn a customized CC subagent with full MCP tool access. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 13,864 | 23,749 | +71% | 1 | 1 | 0% | 1,541 | 3,575 | +132% | 0 | 0 | — |
case-02 | fail→pass | 9,672 | 35,353 | +266% | 1 | 1 | 0% | 794 | 5,855 | +637% | 0 | 0 | — |
case-03 | pass→pass | 13,893 | 18,281 | +32% | 1 | 1 | 0% | 1,718 | 2,997 | +74% | 0 | 0 | — |
case-04 | pass→pass | 14,893 | 13,160 | -12% | 1 | 1 | 0% | 1,704 | 1,873 | +10% | 0 | 0 | — |
case-05 | pass→pass | 10,519 | 13,634 | +30% | 1 | 1 | 0% | 872 | 1,749 | +101% | 0 | 0 | — |
case-06 | pass→pass | 13,885 | 16,518 | +19% | 1 | 1 | 0% | 1,508 | 2,179 | +44% | 0 | 0 | — |
case-07 | pass→pass | 14,751 | 11,037 | -25% | 1 | 1 | 0% | 1,940 | 1,407 | -27% | 0 | 0 | — |
case-08 | pass→pass | 11,140 | 13,941 | +25% | 1 | 1 | 0% | 933 | 1,769 | +90% | 0 | 0 | — |
case-09 | pass→fail | 9,472 | 21,156 | +123% | 1 | 1 | 0% | 828 | 943 | +14% | 0 | 0 | — |
case-10 | pass→pass | 11,701 | 21,248 | +82% | 1 | 1 | 0% | 1,175 | 3,080 | +162% | 0 | 0 | — |
case-11 | pass→fail | 13,023 | 22,783 | +75% | 1 | 1 | 0% | 1,324 | 1,047 | -21% | 0 | 0 | — |
case-12 | pass→pass | 11,727 | 12,897 | +10% | 1 | 1 | 0% | 1,092 | 1,670 | +53% | 0 | 0 | — |
case-13 | pass→pass | 11,070 | 11,700 | +6% | 1 | 1 | 0% | 1,029 | 1,456 | +41% | 0 | 0 | — |
case-14 | pass→pass | 12,964 | 15,519 | +20% | 1 | 1 | 0% | 1,234 | 2,129 | +73% | 0 | 0 | — |
case-15 | fail→pass | 13,674 | 10,001 | -27% | 1 | 1 | 0% | 1,496 | 1,140 | -24% | 0 | 0 | — |
case-16 | pass→pass | 8,486 | 12,513 | +47% | 1 | 1 | 0% | 567 | 1,547 | +173% | 0 | 0 | — |
case-17 | pass→pass | 12,370 | 12,837 | +4% | 1 | 1 | 0% | 1,398 | 1,683 | +20% | 0 | 0 | — |
case-18 | fail→fail | 24,351 | 33,255 | +37% | 1 | 1 | 0% | 3,211 | 2,700 | -16% | 0 | 0 | — |
case-19 | fail→pass | 20,204 | 12,902 | -36% | 1 | 1 | 0% | 2,232 | 1,512 | -32% | 0 | 0 | — |
case-20 | fail→pass | 17,256 | 6,865 | -60% | 1 | 1 | 0% | 1,915 | 578 | -70% | 0 | 0 | — |
case-21 | pass→pass | 18,588 | 13,294 | -28% | 1 | 1 | 0% | 2,186 | 1,743 | -20% | 0 | 0 | — |
case-22 | pass→pass | 12,099 | 14,472 | +20% | 1 | 1 | 0% | 1,298 | 1,960 | +51% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 20 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.