Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Assess systematic biases in the evidence body — publication bias, reporting bias, and selective outcome reporting. Budget: 40 studies, 40 effect sizes, 40 web searches.
.claude/skills/yogsoth-ai-bias-detection/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-04 | ✓→✗ | ▼ Worse | 215% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -23% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 189% | 0% |
| case-07 | ✓→✗ | ▼ Worse | -23% | 0% |
Design a protocol to systematically assess biases that threaten the validity of meta-analytic conclusions.
Bias in the evidence body (publication bias, outcome reporting bias, citation bias, time-lag bias, language bias) can invalidate pooled estimates. This strategy designs the complete bias detection and adjustment protocol — funnel plots, statistical tests, sensitivity analyses, and GRADE certainty downgrading.
| Resource | Floor | Target | |----------|-------|--------| | Studies identified | 28 | 40 | | Effect sizes extracted | 28 | 40 | | Web searches | 28 | 40 | | Bias domains assessed | 5 | 8 | | Quality assessments | 20 | 40 |
Budget gate: cannot exit until 80% of floor met.
<HARD-GATE>
| Metric | Current | Floor | Target | Status |
|--------|---------|-------|--------|--------|
| Studies found | 0 | 28 | 40 | BLOCKED |
| Effect sizes planned | 0 | 28 | 40 | BLOCKED |
| Web searches done | 0 | 28 | 40 | BLOCKED |
| Bias domains assessed | 0 | 5 | 8 | BLOCKED |
| Quality assessed | 0 | 20 | 40 | BLOCKED |
</HARD-GATE>| Tactic | When to Use | |--------|-------------| | effect-size-extraction | Extract effect sizes with precision (SE, CI) | | quality-assessment-protocol | Full RoB2 assessment per study | | evidence-synthesis-planning | Plan bias-adjusted models |
| SOP | When to Use | |-----|-------------| | pico-formulation | Frame the evidence assessment question | | inclusion-criteria-design | Include grey literature, preprints | | effect-size-planning | Ensure precision metrics extracted | | data-extraction-form | Template capturing reporting completeness | | risk-of-bias-assessment | Per-study RoB (core of this strategy) | | publication-bias-assessment | Core SOP — funnel plots, statistical tests | | sensitivity-analysis-design | Trim-and-fill, selection models | | heterogeneity-source-analysis | Bias as heterogeneity driver | | meta-analysis-synthesis | Final bias assessment protocol |
pico-formulation for the evidence reliability questioninclusion-criteria-design maximizing source diversity (grey lit, preprints, registries)effect-size-extraction with precision metrics (SE, CI, N)quality-assessment-protocol for comprehensive RoB2publication-bias-assessment for statistical detection planheterogeneity-source-analysis for bias-driven heterogeneitysensitivity-analysis-design for bias-adjustment methodsmeta-analysis-synthesis for final protocolWeb searches target: trial registries, grey literature databases, dissertation repositories, conference abstracts.
yamlprotocol: question: [Is the evidence body for X biased?] bias_domains: publication_bias: visual: [funnel plot, contour-enhanced funnel] statistical: [Egger's test, Begg's test, Peters' test] adjustment: [trim-and-fill, Copas selection model, PET-PEESE] outcome_reporting_bias: detection: [registry-publication comparison] tool: [ROB-ME, ORBIT] time_lag_bias: detection: [time-to-publication analysis] citation_bias: detection: [citation network analysis] language_bias: mitigation: [multi-language search strategy] small_study_effects: detection: [funnel asymmetry, regression tests] adjustment: [limit meta-analysis] grey_literature_search: [databases, registries, contacts] grade_assessment: domain: publication_bias downgrading_criteria: [when to downgrade certainty] sensitivity_plan: [selection model, 3PSM, p-curve, z-curve] reporting: PRISMA-2020 + ROB-ME guidelines
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | effect-size-extraction | Systematically extract effect sizes and conditions from papers for meta-analytic synthesis | | evidence-synthesis-planning | Plan the statistical synthesis approach — model selection, heterogeneity strategy, and reporting | | quality-assessment-protocol | Methodological quality and bias risk assessment of included studies using validated tools |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | data-extraction-form | Design structured data extraction form for systematic meta-analysis data collection | | effect-size-planning | Determine effect size types and calculation methods for meta-analytic synthesis | | heterogeneity-source-analysis | Identify and classify sources of between-study heterogeneity (clinical, methodological, statistical) | | inclusion-criteria-design | Define inclusion/exclusion criteria for systematic study selection in meta-analysis | | meta-analysis-synthesis | Produce final meta-analysis protocol document assembling all planning outputs into PRISMA-compliant protocol | | pico-formulation | Construct PICO/PECO framework for the meta-analysis research question | | publication-bias-assessment | Plan funnel plots, Egger's test, trim-and-fill, p-curve, and selection model analyses for publication bias | | risk-of-bias-assessment | Assess methodological bias using RoB2, PROBAST, or QUADAS-2 validated tools | | sensitivity-analysis-design | Design leave-one-out, influence diagnostics, subgroup analyses, and robustness checks |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,812 | 10,084 | -59% | 1 | 1 | 0% | 4,438 | 1,942 | -56% | 0 | 0 | — |
case-02 | fail→pass | 22,695 | 38,181 | +68% | 1 | 1 | 0% | 3,806 | 7,159 | +88% | 0 | 0 | — |
case-03 | fail→fail | 30,836 | 9,004 | -71% | 1 | 1 | 0% | 4,879 | 1,927 | -61% | 0 | 0 | — |
case-04 | pass→fail | 11,974 | 36,338 | +203% | 1 | 1 | 0% | 2,420 | 7,634 | +215% | 0 | 0 | — |
case-05 | pass→fail | 14,973 | 11,194 | -25% | 1 | 1 | 0% | 2,840 | 2,187 | -23% | 0 | 0 | — |
case-06 | pass→fail | 11,877 | 26,521 | +123% | 1 | 1 | 0% | 2,174 | 6,274 | +189% | 0 | 0 | — |
case-07 | pass→fail | 18,931 | 15,725 | -17% | 1 | 1 | 0% | 3,041 | 2,353 | -23% | 0 | 0 | — |
case-08 | pass→fail | 16,952 | 12,036 | -29% | 1 | 1 | 0% | 2,497 | 2,361 | -5% | 0 | 0 | — |
case-09 | pass→pass | 7,337 | 36,175 | +393% | 1 | 1 | 0% | 1,237 | 7,627 | +517% | 0 | 0 | — |
case-10 | pass→fail | 24,682 | 9,560 | -61% | 1 | 1 | 0% | 4,059 | 1,953 | -52% | 0 | 0 | — |
case-11 | pass→fail | 22,530 | 12,189 | -46% | 1 | 1 | 0% | 3,644 | 2,051 | -44% | 0 | 0 | — |
case-12 | pass→pass | 12,282 | 18,520 | +51% | 1 | 1 | 0% | 2,043 | 4,609 | +126% | 0 | 0 | — |
case-13 | pass→fail | 17,515 | 10,898 | -38% | 1 | 1 | 0% | 2,697 | 2,059 | -24% | 0 | 0 | — |
case-14 | pass→fail | 9,725 | 14,968 | +54% | 1 | 1 | 0% | 1,583 | 2,310 | +46% | 0 | 0 | — |
case-15 | pass→fail | 18,355 | 12,143 | -34% | 1 | 1 | 0% | 2,650 | 1,983 | -25% | 0 | 0 | — |
case-16 | pass→fail | 16,666 | 9,762 | -41% | 1 | 1 | 0% | 2,481 | 2,129 | -14% | 0 | 0 | — |
case-17 | fail→fail | 15,131 | 11,975 | -21% | 1 | 1 | 0% | 2,536 | 2,164 | -15% | 0 | 0 | — |
case-18 | pass→fail | 11,348 | 9,931 | -12% | 1 | 1 | 0% | 1,897 | 2,142 | +13% | 0 | 0 | — |
case-19 | pass→fail | 7,186 | 10,922 | +52% | 1 | 1 | 0% | 1,161 | 2,125 | +83% | 0 | 0 | — |
case-20 | pass→fail | 15,284 | 42,395 | +177% | 1 | 1 | 0% | 2,302 | 8,202 | +256% | 0 | 0 | — |
case-21 | pass→fail | 5,586 | 14,558 | +161% | 1 | 1 | 0% | 972 | 1,985 | +104% | 0 | 0 | — |
case-22 | pass→fail | 11,623 | 10,750 | -8% | 1 | 1 | 0% | 1,888 | 2,336 | +24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 5 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -68 percentage points is the difference between those two pass rates over the 5 comparable cases. 17 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.