Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Explain why different studies reach different conclusions — heterogeneity investigation protocol. Budget: 30 studies, 30 effect sizes, 50 web searches.
.claude/skills/yogsoth-ai-heterogeneity-investigation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 157% | 0% |
Design a protocol to investigate and explain between-study heterogeneity — why studies of the same question reach different conclusions.
When a meta-analysis reveals substantial heterogeneity (I2 > 50%, significant Q-test, large tau2), this strategy designs the investigation protocol: subgroup analyses, meta-regression, moderator identification, and outlier diagnostics. Produces the investigation plan, not the computation.
| Resource | Floor | Target | |----------|-------|--------| | Studies identified | 20 | 30 | | Effect sizes extracted | 20 | 30 | | Web searches | 35 | 50 | | Moderator candidates | 5 | 10+ | | Quality assessments | 15 | 30 |
Budget gate: cannot exit until 80% of floor met.
<HARD-GATE>
| Metric | Current | Floor | Target | Status |
|--------|---------|-------|--------|--------|
| Studies found | 0 | 20 | 30 | BLOCKED |
| Effect sizes planned | 0 | 20 | 30 | BLOCKED |
| Web searches done | 0 | 35 | 50 | BLOCKED |
| Moderators identified | 0 | 5 | 10+ | BLOCKED |
| Quality assessed | 0 | 15 | 30 | BLOCKED |
</HARD-GATE>| Tactic | When to Use | |--------|-------------| | effect-size-extraction | Extract effect sizes with full study characteristics | | quality-assessment-protocol | Assess whether quality explains heterogeneity | | evidence-synthesis-planning | Plan subgroup and meta-regression models |
| SOP | When to Use | |-----|-------------| | pico-formulation | Frame the heterogeneity question | | inclusion-criteria-design | Broad inclusion to capture variation | | effect-size-planning | Standardize for comparability | | data-extraction-form | Rich moderator variable extraction | | risk-of-bias-assessment | RoB as potential moderator | | heterogeneity-source-analysis | Core SOP — classify heterogeneity sources | | sensitivity-analysis-design | Outlier removal, influence diagnostics | | publication-bias-assessment | Bias as heterogeneity source | | meta-analysis-synthesis | Final investigation protocol |
pico-formulation emphasizing variation in P/I/C/Oinclusion-criteria-design with broad criteria (capture variation)effect-size-extraction with rich study-level covariatesheterogeneity-source-analysis to generate moderator hypothesesquality-assessment-protocol (RoB as moderator)evidence-synthesis-planning for subgroup + meta-regressionmeta-analysis-synthesis for investigation protocolWeb searches focus on domain knowledge about why results might differ (methodological, clinical, statistical heterogeneity).
yamlprotocol: question: [Why do studies of X reach different conclusions?] heterogeneity_metrics: [I2, tau2, Q-test, prediction interval] moderator_candidates: clinical: [population, intervention details, outcome timing] methodological: [study design, RoB, measurement tools] statistical: [effect size type, analysis method, sample size] investigation_plan: subgroup_analyses: [categorical moderators] meta_regression: [continuous moderators] outlier_diagnostics: [influence analysis, Baujat plot] sensitivity: [leave-one-out, cumulative by quality] a_priori_hypotheses: [pre-specified moderator hypotheses] multiple_testing: [correction strategy] reporting: PRISMA-2020 + heterogeneity reporting guidelines
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | effect-size-extraction | Systematically extract effect sizes and conditions from papers for meta-analytic synthesis | | evidence-synthesis-planning | Plan the statistical synthesis approach — model selection, heterogeneity strategy, and reporting | | quality-assessment-protocol | Methodological quality and bias risk assessment of included studies using validated tools |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | data-extraction-form | Design structured data extraction form for systematic meta-analysis data collection | | effect-size-planning | Determine effect size types and calculation methods for meta-analytic synthesis | | heterogeneity-source-analysis | Identify and classify sources of between-study heterogeneity (clinical, methodological, statistical) | | inclusion-criteria-design | Define inclusion/exclusion criteria for systematic study selection in meta-analysis | | meta-analysis-synthesis | Produce final meta-analysis protocol document assembling all planning outputs into PRISMA-compliant protocol | | pico-formulation | Construct PICO/PECO framework for the meta-analysis research question | | publication-bias-assessment | Plan funnel plots, Egger's test, trim-and-fill, p-curve, and selection model analyses for publication bias | | risk-of-bias-assessment | Assess methodological bias using RoB2, PROBAST, or QUADAS-2 validated tools | | sensitivity-analysis-design | Design leave-one-out, influence diagnostics, subgroup analyses, and robustness checks |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 28,400 | 29,207 | +3% | 1 | 1 | 0% | 4,823 | 5,503 | +14% | 0 | 0 | — |
case-02 | fail→pass | 18,455 | 24,804 | +34% | 1 | 1 | 0% | 3,205 | 5,477 | +71% | 0 | 0 | — |
case-03 | fail→pass | 23,878 | 25,937 | +9% | 1 | 1 | 0% | 3,849 | 5,310 | +38% | 0 | 0 | — |
case-04 | fail→fail | 14,089 | 26,893 | +91% | 1 | 1 | 0% | 2,451 | 5,125 | +109% | 0 | 0 | — |
case-05 | pass→fail | 13,874 | 19,997 | +44% | 1 | 1 | 0% | 2,316 | 4,520 | +95% | 0 | 0 | — |
case-06 | pass→pass | 4,129 | 10,980 | +166% | 1 | 1 | 0% | 806 | 3,322 | +312% | 0 | 0 | — |
case-07 | fail→pass | 16,368 | 2,869 | -82% | 1 | 1 | 0% | 2,325 | 1,751 | -25% | 0 | 0 | — |
case-08 | fail→pass | 13,158 | 1,809 | -86% | 1 | 1 | 0% | 596 | 1,534 | +157% | 0 | 0 | — |
case-09 | fail→pass | 28,996 | 2,875 | -90% | 1 | 1 | 0% | 1,529 | 1,774 | +16% | 0 | 0 | — |
case-10 | fail→fail | 14,107 | 10,309 | -27% | 1 | 1 | 0% | 2,102 | 2,781 | +32% | 0 | 0 | — |
case-11 | fail→pass | 20,868 | 2,934 | -86% | 1 | 1 | 0% | 1,800 | 1,780 | -1% | 0 | 0 | — |
case-12 | pass→pass | 14,342 | 10,853 | -24% | 1 | 1 | 0% | 1,879 | 3,118 | +66% | 0 | 0 | — |
case-13 | pass→fail | 14,959 | 46,923 | +214% | 1 | 1 | 0% | 2,374 | 7,909 | +233% | 0 | 0 | — |
case-14 | pass→pass | 10,512 | 15,502 | +47% | 1 | 1 | 0% | 1,582 | 3,722 | +135% | 0 | 0 | — |
case-15 | fail→pass | 11,344 | 2,816 | -75% | 1 | 1 | 0% | 1,615 | 1,743 | +8% | 0 | 0 | — |
case-16 | pass→pass | 13,661 | 4,288 | -69% | 1 | 1 | 0% | 1,952 | 2,029 | +4% | 0 | 0 | — |
case-17 | fail→fail | 14,234 | 19,685 | +38% | 1 | 1 | 0% | 2,033 | 4,327 | +113% | 0 | 0 | — |
case-18 | pass→fail | 18,722 | 20,575 | +10% | 1 | 1 | 0% | 2,754 | 2,573 | -7% | 0 | 0 | — |
case-19 | fail→fail | 16,560 | 37,071 | +124% | 1 | 1 | 0% | 2,429 | 7,488 | +208% | 0 | 0 | — |
case-20 | pass→pass | 22,282 | 9,533 | -57% | 1 | 1 | 0% | 1,978 | 3,023 | +53% | 0 | 0 | — |
case-21 | pass→pass | 13,456 | 8,916 | -34% | 1 | 1 | 0% | 1,987 | 2,768 | +39% | 0 | 0 | — |
case-22 | pass→pass | 10,201 | 2,830 | -72% | 1 | 1 | 0% | 1,436 | 1,728 | +20% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 19 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.