Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Cross-Study Statistical Synthesis Campaign — 5 strategies for systematic collection and methodological planning of multi-study evidence synthesis. Covers pairwise, network, cumulative meta-analysis, heterogeneity investigation, and bias detection. Stops at protocol design (no computation).
.claude/skills/yogsoth-ai-meta-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 63% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 104% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 36% | 0% |
Cross-study statistical synthesis — systematic collection and methodological planning for multi-study evidence synthesis. This campaign orchestrates the full meta-analysis planning pipeline from PICO formulation through protocol design, stopping before computation.
| Trigger | Strategy | When | |---------|----------|------| | Two-arm comparison | pairwise-synthesis | Comparing method A vs B across studies | | Multi-method comparison | network-comparison | Comparing N>=3 methods with indirect evidence | | Temporal evidence tracking | cumulative-tracking | Tracking how evidence evolves over time | | Inconsistent findings | heterogeneity-investigation | Studies disagree — explain why | | Evidence reliability | bias-detection | Assess systematic biases in evidence body |
| Strategy | Purpose | |----------|---------| | pairwise-synthesis | Paired meta-analysis protocol for two-method comparison | | network-comparison | Network meta-analysis for N-method simultaneous comparison | | cumulative-tracking | Cumulative meta-analysis tracking evidence over time | | heterogeneity-investigation | Explain inter-study variation in findings | | bias-detection | Systematic bias assessment of evidence body |
| Tactic | Purpose | |--------|---------| | effect-size-extraction | Extract effect sizes and conditions from papers | | quality-assessment-protocol | Methodological quality and bias risk assessment | | evidence-synthesis-planning | Plan the statistical synthesis approach |
| SOP | Purpose | |-----|---------| | pico-formulation | Construct PICO/PECO framework | | inclusion-criteria-design | Define inclusion/exclusion criteria | | effect-size-planning | Determine effect size types and calculation methods | | data-extraction-form | Design structured data extraction form | | risk-of-bias-assessment | Assess methodological bias (RoB2/PROBAST/QUADAS-2) | | heterogeneity-source-analysis | Identify and classify heterogeneity sources | | publication-bias-assessment | Plan funnel plots, Egger's test, trim-and-fill, p-curve | | sensitivity-analysis-design | Design leave-one-out, influence diagnostics, subgroups | | evidence-network-construction | Build evidence network graph for NMA | | meta-analysis-synthesis | Produce final meta-analysis protocol document |
| Strategy | Studies | Effect Sizes | Web Searches | |----------|---------|--------------|--------------| | pairwise-synthesis | 30 | 30 | 40 | | network-comparison | 50 | 80 | 60 | | cumulative-tracking | 40 | 40 | 30 | | heterogeneity-investigation | 30 | 30 | 50 | | bias-detection | 40 | 40 | 40 |
| MCP Server | Tools | |------------|-------| | brave-search | brave_web_search, brave_llm_context | | apify | rag-web-browser, google-scholar-scraper | | alphaxiv | get_paper_content, answer_pdf_queries | | semantic-scholar | ss_relevance_search, ss_paper, ss_references, ss_citations |
context/meta-analysis/ subdirectories per strategy<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | bias-detection | Assess systematic biases in the evidence body — publication bias, reporting bias, and selective outcome reporting. Budget: 40 studies, 40 effect sizes, 40 web searches. | | cumulative-tracking | Track evidence accumulation over time — cumulative meta-analysis protocol design. Budget: 40 studies, 40 effect sizes, 30 web searches. | | heterogeneity-investigation | Explain why different studies reach different conclusions — heterogeneity investigation protocol. Budget: 30 studies, 30 effect sizes, 50 web searches. | | network-comparison | Compare N methods simultaneously including indirect evidence — network meta-analysis protocol design. Budget: 50 studies, 80 effect sizes, 60 web searches. | | pairwise-synthesis | Compare two methods across multiple studies — paired meta-analysis protocol design. Budget: 30 studies, 30 effect sizes, 40 web searches. |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | fail→pass | 35,518 | 13,050 | -63% | 1 | 1 | 0% | 4,924 | 3,495 | -29% | 0 | 0 | — |
case-06 | pass→pass | 28,885 | 14,510 | -50% | 1 | 1 | 0% | 4,348 | 2,775 | -36% | 0 | 0 | — |
case-01 | fail→fail | 36,820 | 20,231 | -45% | 1 | 1 | 0% | 6,370 | 1,916 | -70% | 0 | 0 | — |
case-02 | fail→fail | 54,141 | 12,623 | -77% | 1 | 1 | 0% | 8,260 | 1,702 | -79% | 0 | 0 | — |
case-03 | fail→fail | 41,475 | 37,963 | -8% | 1 | 1 | 0% | 6,598 | 5,611 | -15% | 0 | 0 | — |
case-04 | fail→pass | 45,566 | 30,862 | -32% | 1 | 1 | 0% | 8,267 | 5,807 | -30% | 0 | 0 | — |
case-05 | fail→pass | 24,107 | 27,548 | +14% | 1 | 1 | 0% | 2,966 | 4,827 | +63% | 0 | 0 | — |
case-07 | pass→pass | 23,767 | 35,689 | +50% | 1 | 1 | 0% | 3,730 | 7,542 | +102% | 0 | 0 | — |
case-08 | pass→fail | 64,740 | 20,693 | -68% | 1 | 1 | 0% | 8,252 | 2,024 | -75% | 0 | 0 | — |
case-09 | pass→pass | 19,029 | 21,031 | +11% | 1 | 1 | 0% | 2,177 | 3,488 | +60% | 0 | 0 | — |
case-10 | pass→pass | 16,918 | 4,534 | -73% | 1 | 1 | 0% | 2,003 | 2,042 | +2% | 0 | 0 | — |
case-11 | pass→pass | 61,166 | 51,694 | -15% | 1 | 1 | 0% | 5,929 | 3,567 | -40% | 0 | 0 | — |
case-13 | pass→pass | 30,163 | 27,385 | -9% | 1 | 1 | 0% | 4,085 | 3,861 | -5% | 0 | 0 | — |
case-14 | pass→pass | 25,591 | 5,585 | -78% | 1 | 1 | 0% | 2,123 | 2,215 | +4% | 0 | 0 | — |
case-15 | fail→fail | 11,712 | 8,035 | -31% | 1 | 1 | 0% | 993 | 1,735 | +75% | 0 | 0 | — |
case-16 | fail→pass | 41,118 | 8,775 | -79% | 1 | 1 | 0% | 1,390 | 2,831 | +104% | 0 | 0 | — |
case-21 | fail→fail | 14,698 | 3,377 | -77% | 1 | 1 | 0% | 1,548 | 1,790 | +16% | 0 | 0 | — |
case-17 | fail→fail | 17,558 | 7,660 | -56% | 1 | 1 | 0% | 1,802 | 1,667 | -7% | 0 | 0 | — |
case-18 | fail→pass | 32,473 | 2,917 | -91% | 1 | 1 | 0% | 1,307 | 1,776 | +36% | 0 | 0 | — |
case-19 | fail→pass | 49,686 | 7,888 | -84% | 1 | 1 | 0% | 3,457 | 1,684 | -51% | 0 | 0 | — |
case-20 | pass→pass | 18,563 | 2,974 | -84% | 1 | 1 | 0% | 2,023 | 1,814 | -10% | 0 | 0 | — |
case-22 | fail→pass | 13,093 | 3,102 | -76% | 1 | 1 | 0% | 1,326 | 1,745 | +32% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 17 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.