Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Compare two methods across multiple studies — paired meta-analysis protocol design. Budget: 30 studies, 30 effect sizes, 40 web searches.
.claude/skills/yogsoth-ai-pairwise-synthesis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✓→✗ | ▼ Worse | 16% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 233% | 0% |
| case-10 | ✓→✗ | ▼ Worse | 176% | 0% |
| case-11 | ✓→✗ | ▼ Worse | -33% | 0% |
| case-13 | ✓→✗ | ▼ Worse | -56% | 0% |
Design a pairwise meta-analysis protocol comparing method A vs method B across multiple independent studies.
When two methods (interventions, algorithms, approaches) have been compared across multiple studies, synthesize the evidence into a single quantitative estimate of relative performance. This strategy produces the complete protocol — stopping before computation.
| Resource | Floor | Target | |----------|-------|--------| | Studies identified | 20 | 30 | | Effect sizes extracted | 20 | 30 | | Web searches | 25 | 40 | | Quality assessments | 15 | 30 |
Budget gate: cannot exit until 80% of floor met.
<HARD-GATE>
| Metric | Current | Floor | Target | Status |
|--------|---------|-------|--------|--------|
| Studies found | 0 | 20 | 30 | BLOCKED |
| Effect sizes planned | 0 | 20 | 30 | BLOCKED |
| Web searches done | 0 | 25 | 40 | BLOCKED |
| Quality assessed | 0 | 15 | 30 | BLOCKED |
</HARD-GATE>| Tactic | When to Use | |--------|-------------| | effect-size-extraction | After study identification, extract paired effect sizes | | quality-assessment-protocol | Assess RoB for each included study | | evidence-synthesis-planning | After extraction, plan the statistical model |
| SOP | When to Use | |-----|-------------| | pico-formulation | First — frame the comparison question | | inclusion-criteria-design | After PICO — define what studies qualify | | effect-size-planning | Determine which effect size metric to use | | data-extraction-form | Design the extraction template | | risk-of-bias-assessment | Per-study quality assessment | | sensitivity-analysis-design | Plan robustness checks | | publication-bias-assessment | Assess reporting bias | | meta-analysis-synthesis | Final protocol assembly |
pico-formulation to structure the comparisoninclusion-criteria-design to define eligibilityeffect-size-extraction tactic for each studyquality-assessment-protocol tactic for bias riskevidence-synthesis-planning tactic for model selectionmeta-analysis-synthesis to produce protocolIterate steps 3-5 until budget floor is met. Check state ledger before each iteration.
yamlprotocol: question: [PICO-structured question] inclusion_criteria: [eligibility rules] studies_included: [list with metadata] effect_size_type: [SMD/OR/RR/MD] model: [fixed-effect/random-effects] heterogeneity_plan: [I2, tau2, subgroup, meta-regression] sensitivity_plan: [leave-one-out, influence diagnostics] bias_assessment_plan: [funnel plot, Egger's, trim-and-fill] reporting: PRISMA-2020
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | effect-size-extraction | Systematically extract effect sizes and conditions from papers for meta-analytic synthesis | | evidence-synthesis-planning | Plan the statistical synthesis approach — model selection, heterogeneity strategy, and reporting | | quality-assessment-protocol | Methodological quality and bias risk assessment of included studies using validated tools |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | data-extraction-form | Design structured data extraction form for systematic meta-analysis data collection | | effect-size-planning | Determine effect size types and calculation methods for meta-analytic synthesis | | inclusion-criteria-design | Define inclusion/exclusion criteria for systematic study selection in meta-analysis | | meta-analysis-synthesis | Produce final meta-analysis protocol document assembling all planning outputs into PRISMA-compliant protocol | | pico-formulation | Construct PICO/PECO framework for the meta-analysis research question | | publication-bias-assessment | Plan funnel plots, Egger's test, trim-and-fill, p-curve, and selection model analyses for publication bias | | risk-of-bias-assessment | Assess methodological bias using RoB2, PROBAST, or QUADAS-2 validated tools | | sensitivity-analysis-design | Design leave-one-out, influence diagnostics, subgroup analyses, and robustness checks |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 38,465 | 46,950 | +22% | 1 | 1 | 0% | 8,266 | 9,938 | +20% | 0 | 0 | — |
case-02 | fail→fail | 44,253 | 134,726 | +204% | 1 | 1 | 0% | 6,919 | 9,421 | +36% | 0 | 0 | — |
case-03 | fail→fail | 44,967 | 10,573 | -76% | 1 | 1 | 0% | 7,506 | 2,199 | -71% | 0 | 0 | — |
case-04 | pass→fail | 39,552 | 41,231 | +4% | 1 | 1 | 0% | 8,115 | 9,393 | +16% | 0 | 0 | — |
case-05 | pass→pass | 48,598 | 41,181 | -15% | 1 | 1 | 0% | 8,250 | 9,399 | +14% | 0 | 0 | — |
case-06 | pass→fail | 19,700 | 76,508 | +288% | 1 | 1 | 0% | 2,814 | 9,369 | +233% | 0 | 0 | — |
case-07 | fail→fail | 30,442 | 9,117 | -70% | 1 | 1 | 0% | 5,502 | 2,925 | -47% | 0 | 0 | — |
case-08 | fail→fail | 60,115 | 20,287 | -66% | 1 | 1 | 0% | 8,219 | 2,253 | -73% | 0 | 0 | — |
case-09 | fail→fail | 18,240 | 11,856 | -35% | 1 | 1 | 0% | 2,972 | 2,529 | -15% | 0 | 0 | — |
case-10 | pass→fail | 18,216 | 39,686 | +118% | 1 | 1 | 0% | 3,387 | 9,365 | +176% | 0 | 0 | — |
case-11 | pass→fail | 21,732 | 18,704 | -14% | 1 | 1 | 0% | 3,970 | 2,649 | -33% | 0 | 0 | — |
case-12 | fail→fail | 26,091 | 10,207 | -61% | 1 | 1 | 0% | 4,290 | 2,974 | -31% | 0 | 0 | — |
case-13 | pass→fail | 45,043 | 7,508 | -83% | 1 | 1 | 0% | 5,570 | 2,463 | -56% | 0 | 0 | — |
case-14 | pass→pass | 21,453 | 44,262 | +106% | 1 | 1 | 0% | 3,762 | 9,369 | +149% | 0 | 0 | — |
case-15 | fail→fail | 23,257 | 8,539 | -63% | 1 | 1 | 0% | 3,483 | 2,601 | -25% | 0 | 0 | — |
case-16 | fail→fail | 27,924 | 43,860 | +57% | 1 | 1 | 0% | 3,555 | 9,363 | +163% | 0 | 0 | — |
case-17 | fail→fail | 21,550 | 46,797 | +117% | 1 | 1 | 0% | 3,455 | 9,360 | +171% | 0 | 0 | — |
case-18 | fail→fail | 43,081 | 41,847 | -3% | 1 | 1 | 0% | 4,432 | 9,368 | +111% | 0 | 0 | — |
case-19 | fail→fail | 68,200 | 40,183 | -41% | 1 | 1 | 0% | 8,214 | 9,362 | +14% | 0 | 0 | — |
case-20 | fail→fail | 30,223 | 44,896 | +49% | 1 | 1 | 0% | 4,826 | 8,985 | +86% | 0 | 0 | — |
case-21 | fail→fail | 42,907 | 11,893 | -72% | 1 | 1 | 0% | 5,495 | 2,952 | -46% | 0 | 0 | — |
case-22 | fail→fail | 30,024 | 41,689 | +39% | 1 | 1 | 0% | 5,094 | 9,368 | +84% | 0 | 0 | — |
case-23 | fail→fail | 32,389 | 12,281 | -62% | 1 | 1 | 0% | 5,478 | 2,330 | -57% | 0 | 0 | — |
case-24 | fail→fail | 30,320 | 42,154 | +39% | 1 | 1 | 0% | 4,917 | 9,360 | +90% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 23 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -21 percentage points is the difference between those two pass rates over the 23 comparable cases. 8 cases got worse with the skill loaded, and they are included in that figure.
Other measured skills in the registry, with their headline benchmark lift.