Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Track evidence accumulation over time — cumulative meta-analysis protocol design. Budget: 40 studies, 40 effect sizes, 30 web searches.
.claude/skills/yogsoth-ai-cumulative-tracking/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 112% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 162% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-23 | ✗→✓ | ▲ Improved | 59% | 0% |
Design a cumulative meta-analysis protocol tracking how evidence evolves over time as new studies are published.
Cumulative meta-analysis adds studies one-by-one in chronological order, showing when the evidence became conclusive, whether early studies were misleading, and how the pooled estimate stabilized. This strategy produces the protocol for temporal evidence tracking.
| Resource | Floor | Target | |----------|-------|--------| | Studies identified | 28 | 40 | | Effect sizes extracted | 28 | 40 | | Web searches | 20 | 30 | | Temporal coverage (years) | 5 | 10+ | | Quality assessments | 20 | 40 |
Budget gate: cannot exit until 80% of floor met.
<HARD-GATE>
| Metric | Current | Floor | Target | Status |
|--------|---------|-------|--------|--------|
| Studies found | 0 | 28 | 40 | BLOCKED |
| Effect sizes planned | 0 | 28 | 40 | BLOCKED |
| Web searches done | 0 | 20 | 30 | BLOCKED |
| Year range covered | 0 | 5 | 10+ | BLOCKED |
| Quality assessed | 0 | 20 | 40 | BLOCKED |
</HARD-GATE>| Tactic | When to Use | |--------|-------------| | effect-size-extraction | Extract effect sizes with publication dates | | quality-assessment-protocol | Assess quality evolution over time | | evidence-synthesis-planning | Plan cumulative pooling approach |
| SOP | When to Use | |-----|-------------| | pico-formulation | Frame the temporal evidence question | | inclusion-criteria-design | Define eligibility with temporal scope | | effect-size-planning | Standardize effect sizes for temporal pooling | | data-extraction-form | Template with mandatory date fields | | risk-of-bias-assessment | Per-study assessment (track quality trends) | | heterogeneity-source-analysis | Time-varying heterogeneity | | sensitivity-analysis-design | First-study effect, vintage analysis | | publication-bias-assessment | Time-lag bias assessment | | meta-analysis-synthesis | Final cumulative protocol assembly |
pico-formulation with temporal dimension explicitinclusion-criteria-design with date range requirementseffect-size-extraction tactic with date metadataquality-assessment-protocol noting temporal trendsevidence-synthesis-planning for cumulative modelmeta-analysis-synthesis for final protocolEnsure no temporal gaps. Flag periods with no publications.
yamlprotocol: question: [PICO with temporal dimension] temporal_scope: [start_year - end_year] inclusion_criteria: [eligibility with date requirements] studies_included: - [study, year, effect_size, cumulative_n] chronological_order: [sorted study list] effect_size_type: [consistent metric across time] model: [random-effects with cumulative pooling] temporal_analyses: - cumulative_forest_plot - first_study_effect_test - evidence_stabilization_point - vintage_regression time_lag_bias: [assessment plan] quality_trend: [RoB evolution over time] reporting: PRISMA-2020 + temporal extension
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | effect-size-extraction | Systematically extract effect sizes and conditions from papers for meta-analytic synthesis | | evidence-synthesis-planning | Plan the statistical synthesis approach — model selection, heterogeneity strategy, and reporting | | quality-assessment-protocol | Methodological quality and bias risk assessment of included studies using validated tools |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | data-extraction-form | Design structured data extraction form for systematic meta-analysis data collection | | effect-size-planning | Determine effect size types and calculation methods for meta-analytic synthesis | | heterogeneity-source-analysis | Identify and classify sources of between-study heterogeneity (clinical, methodological, statistical) | | inclusion-criteria-design | Define inclusion/exclusion criteria for systematic study selection in meta-analysis | | meta-analysis-synthesis | Produce final meta-analysis protocol document assembling all planning outputs into PRISMA-compliant protocol | | pico-formulation | Construct PICO/PECO framework for the meta-analysis research question | | publication-bias-assessment | Plan funnel plots, Egger's test, trim-and-fill, p-curve, and selection model analyses for publication bias | | risk-of-bias-assessment | Assess methodological bias using RoB2, PROBAST, or QUADAS-2 validated tools | | sensitivity-analysis-design | Design leave-one-out, influence diagnostics, subgroup analyses, and robustness checks |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 29,796 | 40,456 | +36% | 1 | 1 | 0% | 5,941 | 7,524 | +27% | 0 | 0 | — |
case-02 | fail→fail | 31,982 | 34,593 | +8% | 1 | 1 | 0% | 6,243 | 7,511 | +20% | 0 | 0 | — |
case-08 | fail→pass | 14,355 | 2,427 | -83% | 1 | 1 | 0% | 784 | 1,662 | +112% | 0 | 0 | — |
case-03 | pass→fail | 21,710 | 12,697 | -42% | 1 | 1 | 0% | 3,437 | 2,334 | -32% | 0 | 0 | — |
case-04 | pass→fail | 21,491 | 10,456 | -51% | 1 | 1 | 0% | 3,717 | 2,063 | -44% | 0 | 0 | — |
case-05 | pass→fail | 26,904 | 6,958 | -74% | 1 | 1 | 0% | 4,626 | 2,377 | -49% | 0 | 0 | — |
case-06 | pass→fail | 22,252 | 37,527 | +69% | 1 | 1 | 0% | 3,703 | 7,472 | +102% | 0 | 0 | — |
case-07 | fail→pass | 29,942 | 20,224 | -32% | 1 | 1 | 0% | 1,973 | 5,173 | +162% | 0 | 0 | — |
case-09 | fail→fail | 11,885 | 2,506 | -79% | 1 | 1 | 0% | 601 | 1,694 | +182% | 0 | 0 | — |
case-10 | fail→fail | 16,480 | 2,653 | -84% | 1 | 1 | 0% | 1,063 | 1,670 | +57% | 0 | 0 | — |
case-11 | fail→fail | 25,622 | 2,986 | -88% | 1 | 1 | 0% | 1,098 | 1,835 | +67% | 0 | 0 | — |
case-12 | pass→fail | 10,206 | 17,874 | +75% | 1 | 1 | 0% | 1,704 | 3,426 | +101% | 0 | 0 | — |
case-13 | pass→fail | 16,701 | 28,727 | +72% | 1 | 1 | 0% | 3,522 | 7,452 | +112% | 0 | 0 | — |
case-14 | fail→pass | 11,316 | 4,360 | -61% | 1 | 1 | 0% | 1,830 | 2,011 | +10% | 0 | 0 | — |
case-15 | pass→pass | 10,312 | 5,265 | -49% | 1 | 1 | 0% | 1,649 | 2,201 | +33% | 0 | 0 | — |
case-16 | pass→pass | 8,926 | 33,497 | +275% | 1 | 1 | 0% | 1,461 | 7,444 | +410% | 0 | 0 | — |
case-17 | pass→fail | 12,811 | 17,220 | +34% | 1 | 1 | 0% | 2,438 | 2,115 | -13% | 0 | 0 | — |
case-18 | pass→pass | 13,389 | 11,538 | -14% | 1 | 1 | 0% | 1,909 | 2,923 | +53% | 0 | 0 | — |
case-19 | pass→pass | 15,103 | 3,780 | -75% | 1 | 1 | 0% | 1,429 | 1,944 | +36% | 0 | 0 | — |
case-20 | pass→fail | 12,457 | 2,423 | -81% | 1 | 1 | 0% | 1,905 | 1,700 | -11% | 0 | 0 | — |
case-21 | pass→fail | 13,002 | 2,809 | -78% | 1 | 1 | 0% | 1,875 | 1,706 | -9% | 0 | 0 | — |
case-22 | fail→pass | 13,411 | 3,016 | -78% | 1 | 1 | 0% | 2,146 | 1,758 | -18% | 0 | 0 | — |
case-23 | fail→pass | 11,436 | 9,065 | -21% | 1 | 1 | 0% | 1,647 | 2,617 | +59% | 0 | 0 | — |
case-24 | fail→fail | 9,888 | 2,518 | -75% | 1 | 1 | 0% | 1,500 | 1,673 | +12% | 0 | 0 | — |
case-25 | fail→fail | 8,639 | 1,869 | -78% | 1 | 1 | 0% | 1,335 | 1,487 | +11% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 17 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -20 percentage points is the difference between those two pass rates over the 17 comparable cases. 10 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.