Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Systematic review and meta-analysis pipeline for medical research. Covers protocol registration (PROSPERO), search strategy, screening, data extraction, risk of bias assessment (QUADAS-2/ROBINS-I), statistical synthesis (bivariate/HSROC for DTA, random-effects for intervention), and PRISMA-compliant reporting. Supports both DTA and intervention meta-analyses.
.claude/skills/aperivue-meta-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 136% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 591% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 518% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 537% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 576% | 0% |
You are helping a medical researcher conduct a systematic review and meta-analysis. You support the full pipeline from protocol development to submission-ready manuscript, with specialized support for diagnostic test accuracy (DTA) meta-analyses.
${CLAUDE_SKILL_DIR}/references/)${CLAUDE_SKILL_DIR}/references/PROSPERO_template.md -- field-by-field guide with word limits, pitfalls checklist${CLAUDE_SKILL_DIR}/references/icmje_coi_guide.md -- batch generation, python-docx pitfalls, form structure${CLAUDE_SKILL_DIR}/references/r_templates.md${CLAUDE_SKILL_DIR}/references/checklists/PRISMA_DTA.md -- 27-item checklistQUADAS3.md -- current recommended DTA tool: 6 phases, 4 domains, 20 signalling questions, assessed per accuracy estimateQUADAS2.md -- the 2011 tool: 4 domains + 10 signalling questions (use when appraising or reproducing a review that used it)ROBINS_I.md -- 7 domains + pre-assessment + synthesis recommendationRoB2.md -- 5 domains + signalling questions + overall judgmentPROBAST.md -- 4 domains + AI extension + validation studiesNOS.md -- Cohort (8 items) + Case-control (8 items) + star interpretationJBI_Case_Series.md -- 10-item critical appraisal checklist for case series${CLAUDE_SKILL_DIR}/references/phase9_circulation.md -- thread continuity, attachment scope, recipient structure, 7-day window${CLAUDE_SKILL_DIR}/references/phase10_recovery.md -- trigger conditions, 12-step rebuild sprint, PROSPERO amendment, re-circulation framing${CLAUDE_SKILL_DIR}/references/data_integrity_checklist.md -- DI-1~DI-9 extraction/synthesis guardrails (prior anonymized MA projects)${CLAUDE_SKILL_DIR}/references/review_orchestration.md -- RO-1~RO-5 circulation discipline (extends phase9_circulation.md)${CLAUDE_SKILL_DIR}/references/submission_package_drift.md -- multi-journal folder hygiene, DO_NOT_EDIT_HERE gate, _build.sh pattern${CLAUDE_SKILL_DIR}/references/post_submission_release_ops.md -- Zenodo DOI gating, tag-cleanup gates, reject-retarget versioning${CLAUDE_SKILL_DIR}/references/empirical_lessons.md -- 16 accumulated SR-MA peer-review / submission lessons (2026-05/06) that drive the Phase 4 extraction-form schema, Phase 4c QC, and Phase 8 submission gates. Load before designing the extraction form and before submission.${CLAUDE_SKILL_DIR}/templates/)templates/extraction_form_v2.md) -- dual-extractor schema with source_page_ref, source_verbatim_quote, cohort_source, overlap_flag_reviewer1/2, sample_n_dta_pool vs sample_n_prognostic_pool columns. Required for SR-MA targeting high-impact radiology / medical AI journals.templates/supplementary_8file_checklist.md) -- S1-S8 mandatory package (PRISMA, PROSPERO, search strategy, exclusion list, extraction table, per-study x per-domain RoB, subgroup forests, sensitivity / publication bias) with a submission-gate bash check.${CLAUDE_SKILL_DIR}/scripts/)screening_reconcile.py -- Phase 3f ID-set screening reconciliation.check_pool_consistency.py -- pool-composition / PRISMA count consistency.cohort_overlap_check.py -- shared-database cohort-overlap detection.extract_assist.py -- Phase 4 AI-assisted extraction suggestions (page ref + verbatim quote, AI_SUGGESTED/needs_review); human-confirm then dta_extraction_qc.py. Challenge card: scripts/extract_assist_challenge/.dta_extraction_qc.py -- 2x2 cell ↔ source sens/spec QC on the confirmed extraction CSV.| Type | RoB Tool | Statistical Model | Reporting Guideline | |------|----------|-------------------|-------------------| | DTA (diagnostic test accuracy) | QUADAS-3 (QUADAS-2 for legacy reviews) | Bivariate / HSROC | PRISMA-DTA | | Intervention (treatment effect) | RoB 2 (RCT) / ROBINS-I (NRSI) | Random-effects (DL/REML) | PRISMA 2020 | | Prognostic (prediction model) | QUIPS / PROBAST | Random-effects | PRISMA 2020 | | Observational (prevalence/association) | NOS / JBI | Random-effects | MOOSE |
Auto-detect type from the research question or accept user specification.
Goal: Produce a PROSPERO-ready protocol document.
QUADAS-3's first two phases are review-level and belong in the protocol: phase 1 states the synthesis question(s) (population, index test(s), target condition — a review may have more than one), and phase 2 defines the ideal test accuracy trial for each: objective, participants, index test(s), definition of the target condition, analysis. Every later risk-of-bias and applicability judgement is made against that trial. Write the review-specific guidance for answering each signalling question here too, with clinical and methodological input, and publish it as a web appendix. Defining the ideal trial after seeing the studies is not an assessment — it is a judgement fitted to the results. See references/checklists/QUADAS3.md.
${CLAUDE_SKILL_DIR}/references/PROSPERO_template.md for field-by-field guidanceCRD42 + 9 digits (14 characters total), e.g. CRD42024500001. Validate any ID that appears in the manuscript or registration doc with grep -oE 'CRD42[0-9]+' and assert a 14-character length / ^CRD42\d{9}$ — a 15-character ID (a stray digit) is a transcription error a reviewer will check against the live record.7_Submission/ or equivalent directoryGoal: Develop and validate reproducible search strategies.
/search-lit:Save search strategies as a structured document, one section per database, with date of search, number of results, and any limits applied.
Deduplicate by DOI first, then PMID. Save raw counts for PRISMA flow.
Goal: Systematic title/abstract and full-text screening with two independent reviewers.
3a. Round 1 — initial title/abstract screening (single reviewer). Define the exclusion codes from the protocol (E1=Not target population, E2=Not intervention, E3=Ineligible type, E4=Non-human, E5=Duplicate). Mark every record INCLUDE / EXCLUDE / MAYBE with a reason code → round1_{date}.tsv.
3b. Round 2 — dual independent title/abstract screening. A second independent reviewer (or AI as a documented second-pass tool with human verification) re-screens all R1 records. Compute Cohen's κ and report it in Methods. round2_tag = INCLUDE / EXCLUDE / MAYBE, where MAYBE means disagreement or either reviewer flagged uncertainty → round2_tag, round2_reason columns.
3c. Round 3 — adjudication of disagreements (first reviewer). Build the R3 sheet with all MAYBE records first, then INCLUDE records for a brief confirmation pass. The first reviewer independently adjudicates each row (round3_decision, plus round3_reason only when overturning R2). Optional AI-assisted pre-screening can compress the effort — but AI suggestions are not decisions: the reviewer independently confirms or overturns every one. Template, sort priority, and the required Methods boilerplate are in the reference file.
3d. Round 4 — full-text screening. Retrieve full texts for round3_decision = INCLUDE (use /fulltext-retrieval), apply the full-text exclusion codes (F1=No extractable outcome, F2=No comparative data, F3=Cannot separate target population, F4=Inadequate sample/follow-up, F5=Full-text unavailable), with two independent reviewers, Cohen's κ, and consensus or a third reviewer for disagreements. Flag comparative studies for priority extraction.
3e. PRISMA flow. Track counts at every stage (R1 → R2 → R3 → R4 → final included); generate the diagram with /make-figures once the numbers are final.
3f. Post-consensus count reconciliation gate (MANDATORY before Phase 5 write-up). Reconcile the counts from the raw ID sets, never from prose summaries, and record the canonical totals in one source-of-truth file:
bashpython "${CLAUDE_SKILL_DIR}/scripts/screening_reconcile.py" \ --screening 2_Screening/fulltext_screening.tsv \ --consensus 2_Screening/consensus_decisions.tsv \ --table1 6_Tables/table1_studies.csv \ --output 2_Screening/screening_consensus.json
Downstream stages consume screening_consensus.json for counts and ID sets; the Markdown consensus document remains the human explanation. Three hard rules:
narrative-only studies") that does not match the enumerable set (A ∪ C) \ B \ T.
cite the added/removed IDs. A transition claim with no enumerable ID set is a P0 and blocks the Phase 5 hand-off.
STAGE_TRANSFER_LOSS is a P0. Exit 1 when a record is included at screening but absentfrom the consensus artifact altogether — no adjudication was ever recorded. An exclusion is a decision; silence is a gap. Never let it settle into narrative-only (why: reference file).
The set algebra, the reconciliation-table template, and the failure pattern it exists for (a manuscript ships counts the ID sets do not support, with every downstream artifact echoing the same unreconciled prose total) are in the reference file.
3f.5 Pool composition lock (MANDATORY at adjudication freeze). Once 3f passes, freeze the pool into a single source-of-truth YAML that every downstream artifact can be checked against:
bashcp "${CLAUDE_SKILL_DIR}/templates/FINAL_POOL_LOCK.yaml.template" 2_Data/FINAL_POOL_LOCK.yaml # fill counts + UID lists from 3f, compute the SHA-256 over the sorted UID list, # and COMMIT THE LOCK before any Phase 4 extraction
k included from the extraction TSV at manuscript build time — alwaysreference final_pool_n from the lock.
arm-separable from both-arm rows: a study contributing one arm must not have its full-cohort count folded into a pooled total. A hand-carried headline total that does not re-derive from the locked per-study values is a P0.
FINAL_POOL_LOCK_v2.yaml, and propagate to every artifact.
Read on demand:
| File | Read it when | Cost if read blindly | |---|---|---| | references/phase3_screening_detail.md | you are executing a screening round, using AI pre-screening, or a reconciliation/lock gate fired | ~3,600 tokens; the round procedures are needed one round at a time, not all at invocation |
Goal: Create standardized extraction forms and extract 2x2 or effect-size data.
4.0 Entry gate (MANDATORY) — pool composition lock ↔ adjudication TSV. Before any extraction work begins, confirm the round-3 adjudication TSV and FINAL_POOL_LOCK.yaml (Phase 3f.5) agree on which UIDs are included:
bashpython "${CLAUDE_SKILL_DIR}/scripts/check_pool_consistency.py" \ --lock 2_Data/FINAL_POOL_LOCK.yaml \ --adjudication-tsv 2_Screening/round3_adjudication.tsv \ --decision-col round3_decision --uid-col uid \ --include-labels "INCLUDE,INCLUDE_MIXED" \ --out qc/pool_consistency.json
The gate fails closed: any UID disagreement blocks extraction. Resolve by re-freezing the lock with the corrected UID set (and propagating downstream) or by correcting a mis-labelled TSV row. Do NOT proceed with a mismatch — the extraction matrix will not align with the locked pool, and the drift surfaces as a fabrication-grade red flag at peer review.
> Failure-mode cross-ref → references/data_integrity_checklist.md DI-1~DI-5 are mandatory > during extraction (2x2 arm-swap, KM audit trail, methodology mismatch, PRISMA 5-way drift, > single-source k).
Extraction form. For an SR-MA targeting high-impact radiology / medical AI journals use ${CLAUDE_SKILL_DIR}/templates/extraction_form_v2.md — its dual-extractor, source-page-reference, and verbatim-quote columns are what close the 2x2 cell-swap and cohort-overlap blind spots. The DTA and intervention field lists are in the reference file.
AI-drafted starting document — treat as hallucination-suspect. If a mentor or collaborator shared an AI-drafted study list, 2x2 set, or effect estimates (even flagged "for reference only"): save it with a _DO_NOT_USE_VERBATIM suffix and re-verify every N, denominator, event count, OR/CI, and author/year against the source PDF. Trust hierarchy: source PDF + own analysis stdout > the mentor's direct text > the attached AI draft — never promote a draft up that ladder. Procedure and precedent: reference file.
4b. Special cases (KM reconstruction, composite exposure). When studies report outcomes only as Kaplan-Meier curves, or the intervention is a composite of techniques, load ${CLAUDE_SKILL_DIR}/references/phase4_km_composite.md for the WebPlotDigitizer → IPDfromKM procedure (cite Guyot et al. 2012, doi:10.1186/1471-2288-12-9) and the 4-path composite-exposure decision tree. Pre-specify a sensitivity analysis excluding composite-exposure studies.
Cross-verification (≥2 independent reviewers). Report inter-reviewer agreement (% or Cohen's κ) at title/abstract and full-text stages. Verify denominator consistency — the denominator may differ across outcomes within one study, so for each outcome back-calculate event ÷ denominator and confirm it reproduces the paper's reported percentage. Distinguish KM-curve estimates from raw event counts and record the data source (Table / KM / text). Log every consensus decision in {project}/consensus_log.md, then lock the dataset; later changes need a dated justification.
4c. Extraction QC & cohort overlap. After dual-extractor consensus, run both before locking:
bash# 2x2 cell integrity: validates TP/FN/TN/FP against source-reported sens/spec (catches arm-swap) python3 "${CLAUDE_SKILL_DIR}/scripts/dta_extraction_qc.py" \ --input 2_Extraction/extraction.csv --tolerance 0.02 \ --out 2_Extraction/qc/dta_extraction_qc.tsv # cohort overlap: shared public DB / same institution+period / same first author ±2y python3 "${CLAUDE_SKILL_DIR}/scripts/cohort_overlap_check.py" \ --input 2_Extraction/studies.csv --enrich \ --out 2_Extraction/qc/cohort_overlap.md
Any FLAG_SWAP / FLAG_MISMATCH requires third-reviewer adjudication before Phase 6. A confirmed flag is not resolved until the extraction form itself is edited — a flag corrected only in a review note silently re-enters synthesis, so re-run the QC and confirm zero open flags before locking. HIGH-confidence overlap pairs require a Limitations acknowledgment plus a sensitivity analysis excluding one of the pair. Cross-links: /peer-review Phase 2A P1 + P2.
Read on demand:
| File | Read it when | Cost if read blindly | |---|---|---| | references/phase4_extraction_detail.md | building the extraction form, an AI draft was shared, you want the optional extract_assist.py scaffolding, or a QC flag fired | ~4,700 tokens; a clean dual-extraction with no AI draft needs none of it | | references/phase4_km_composite.md | studies report only KM curves, or the exposure is composite | ~2,200 tokens |
Goal: Guide structured RoB assessment with the appropriate tool.
DTA: this phase runs QUADAS-3 phases 3–6 (flow diagram, identify the estimates to assess, assess, overall judgement). Phases 1–2 — the synthesis question and the ideal test accuracy trial — were written in Phase 1 above. If they were not, stop and write them before judging anything; they are the comparator every judgement is made against.
Select tool based on meta-analysis type (see table above), then read the corresponding checklist:
| Tool | Checklist File | |------|---------------| | QUADAS-3 (DTA, current) | ${CLAUDE_SKILL_DIR}/references/checklists/QUADAS3.md | | QUADAS-2 (DTA, legacy) | ${CLAUDE_SKILL_DIR}/references/checklists/QUADAS2.md | | RoB 2 (RCT) | ${CLAUDE_SKILL_DIR}/references/checklists/RoB2.md | | ROBINS-I (NRSI) | ${CLAUDE_SKILL_DIR}/references/checklists/ROBINS_I.md | | PROBAST (Prediction) | ${CLAUDE_SKILL_DIR}/references/checklists/PROBAST.md | | NOS (Observational) | ${CLAUDE_SKILL_DIR}/references/checklists/NOS.md | | JBI (Case Series) | ${CLAUDE_SKILL_DIR}/references/checklists/JBI_Case_Series.md |
For AI/ML prediction models, also apply PROBAST+AI extensions.
Output: Summary table + traffic light plot (use /make-figures).
Goal: Execute meta-analysis and generate publication-ready outputs.
> Failure-mode cross-ref → references/data_integrity_checklist.md DI-6/DI-7/DI-9 are the consistency gate (CSV ↔ script ↔ prose; single-source k; 3-way numeric reconciliation before Stage 4).
IMPORTANT: Always use R for meta-analysis (packages: meta, metafor, mada). See ${CLAUDE_SKILL_DIR}/references/r_templates.md for full code templates.
| Analysis family | Primary tool | Key output | |-----------------|-------------|-----------| | DTA | mada::reitsma() (bivariate) | Pooled Se/Sp + SROC with confidence/prediction regions | | Intervention | meta::metagen() / meta::metabin() | Pooled OR/RR, I², Egger's test, leave-one-out | | Dual (comparative + single-arm) | metabin + metaprop | PRIMARY vs SECONDARY per pre-specified protocol |
Load-on-demand: Read ${CLAUDE_SKILL_DIR}/references/phase6_statistical_synthesis.md for the full R code templates, the dual-approach decision table (comparative vs single-arm), practical cautions (method.tau, HK CI, zero-cell correction), publication-bias test power, sensitivity-analysis menu, and error-handling rules.
Three checks before the pool is written up — each is a Methods sentence, not only a setting. R and detail in the same reference:
analysis off the inverse-variance default onto Peto / Mantel-Haenszel without a zero-cell correction / GLMM. Inverse-variance methods including DerSimonian-Laird are to be avoided for rare events, and so are 0.5 continuity corrections with them.
exists — never derived from Cochran's Q or I². "A random-effects model was used because I² was 65%" is a reviewer catch, not a rationale.
readers, thresholds, or time points from the same participants need one pre-specified estimate per study, a multivariate model, or robust variance estimation — not independent pooling.
Goal: Catch numerical hallucinations that survived the forward pipeline (CSV → .R → manuscript).
The failure pattern — treat this as a lived near-miss, not hypothetical: > A safety outcome is reported with its arm-level events, and therefore its p-value, > direction-reversed relative to what the primary-source Table actually recorded. > The extraction CSV is correct; the R script's Fisher exact > matrix() was hand-typed after a column in the source Table was misread. Internal > consistency checks passed because every downstream artifact (Abstract, Discussion, > Table, forest caption) echoed the same wrong number. The reversal was caught only on > a second-pass audit with random extraction sampling against the primary paper.
Non-negotiable rules:
read.csv(...) + subset / filter. Never copy a 2x2 table from a paper's Table intomatrix(c(...), ...) by eye.
matrix, c(), ordata.frame line MUST carry a comment citing the exact CSV row + column OR the exact primary-source Table/Page coordinate. Example: r # source: data_extraction_final.csv row <N> (<first-author> <year>), cols <event_arm1>=0, <event_arm2>=1 # verified against primary source Table <X>, page <P> fisher.test(matrix(c(0, 45, 1, 55), nrow = 2, byrow = FALSE))
comparative analysis while the full cohort of that study appears elsewhere, extraction_consensus_log.md must carry an explicit row for the arm-specific values. Pooled totals and arm-specific values MUST NOT share a row.
from the Results section of the draft manuscript and trace each back to (a) the R output log and (b) the original paper's Table/Figure.
peer_review_<vN>_internal.md:| Claim (manuscript line) | R output file:line | Primary source (paper, Table/Fig, page) | Match? | |---|---|---|---|
sensitivity script — MUST be wrapped inline as [VERIFY-CSV] in the manuscript until the Phase 2.5a audit in /self-review clears it.
reported effect size (Cohen's dz/f, AUC, OR, HR, β, sens/spec, ICC) MUST be re-derived from the modified dataset. If a sensitivity-table effect size is identical to the primary analysis to two decimals across ≥4 values, the recomputation almost certainly did not run and the primary values were transcribed — re-run the script on the modified data.
effect sizes are byte-identical while the inputs differ, that is the tell. Probability of ≥4 independent values coinciding to 2 decimals by chance is ≈ (0.01)^4 — essentially zero.
byte-identical to the primary tables while the underlying means/SDs differ — the sensitivity analysis was never actually recomputed. Internal consistency cannot see it.
fixed, resolved, or corrected, that status isonly valid if it carries the re-run evidence: a timestamp and the relevant stdout / output-file line showing the corrected value, or the commit that changed it. A bare "fixed in v10" with no re-run artifact does NOT clear the finding — re-run the script and attach the output.
was fixed (e.g., a major-comparison N still reading the old total after a "fixed" note). The outcome-denominator cross-check (/self-review Phase 2.5b, the cohort-arithmetic / pool-lock assertions) must pass against the current outputs before any "fixed" status is accepted.
When this phase triggers: every time Phase 6 outputs change (first draft, revision, reviewer- requested re-analysis). Not optional on "minor" re-runs — the precedent reversal above occurred inside a "minor" revision-era re-analysis.
Goal: Assess certainty of the body of evidence.
For DTA meta-analysis, apply GRADE-DTA framework:
For intervention meta-analysis, apply standard GRADE.
Certainty is assessed per outcome, not once for the review. The five domains resolve differently for each outcome — an outcome pooled from 12 studies with narrow CIs and one pooled from 3 with a wide CI do not share a rating, and a single review-level "moderate certainty" sentence tells a reader nothing about the outcome they came for. Rate every outcome carried into the Summary of Findings table, and state the reason for each downgrade (which domain, why) rather than the resulting label alone.
Output: Summary of Findings table — one row per outcome, carrying the pooled estimate with its precision alongside the certainty rating (high / moderate / low / very low).
Goal: Generate PRISMA-compliant manuscript sections.
> Failure-mode cross-ref → references/submission_package_drift.md — apply the _build.sh pattern + DO_NOT_EDIT_HERE gate when staging multi-journal submission folders.
/check-reporting with PRISMA-DTA or PRISMA 2020, thenrun it a second time over the abstract with PRISMA 2020 for Abstracts — 12 items, its own denominator. One run does not cover both.
/write-paper with meta-analysis type selected/make-figures for:PMID:35213097) scored 24 SR/MAs against PRISMA 2020 and found 24 of 42 items reported by fewer than 80%. The checklist itself lives in /check-reporting; what follows is where drafts actually fail, so check these by hand before the compliance run rather than after it:
| PRISMA item | What is missing | Observed | |---|---|---| | 20a | For each synthesis, a brief summary of the contributing studies' characteristics and risk of bias — not one global paragraph covering all pools | 0/24 | | 27 | Data availability: which of the extraction forms, extracted data, analysis dataset, and analytic code are public, and where | 0/24 | | 24a–c | Registration number, where the protocol can be read, and any amendment — an explicit "not registered" satisfies 24a | 0/24 | | 22 / 15 | Certainty of evidence per outcome, and the method used to assess it | 9% | | 13f / 20d | Sensitivity analysis: method and result | 28% | | 18 | Risk of bias per study, shown study-by-study rather than as a pooled proportion | 32% | | 13d | Rationale for the synthesis model (see Phase 6 check 2) | 35% | | 16b | Studies that look eligible but were excluded, cited individually with the reason | 25% | | Abstract #3, #12 | Eligibility criteria and registration inside the structured abstract | 0/24 each |
The abstract items are the cheapest of these and the most reliably forgotten. PRISMA 2020 devotes a separate 12-item instrument to the abstract — item 2 of the main checklist does nothing but defer to it — so a manuscript can satisfy all 42 main-text items and still fail most of the twelve. /check-reporting carries it as PRISMA_2020_Abstracts.md; run it as its own pass and report its score separately, because folding twelve items into a 42-item total is how they stay invisible.
locked dataset, analysis code, RoB judgments) and where — repository, DOI, or supplementary file. "Available from the corresponding author on reasonable request" satisfies few journals now and no longer satisfies item 27. If a Zenodo DOI is minted post-acceptance, references/post_submission_release_ops.md covers propagating it back into this statement.
/check-reporting output ("Assessed by: <tool>", JSON blocks, "READY FOR SUBMISSION" verdicts, action-item lists), search-development planning docs (decision logs, expected-yield estimates, [Check on execution] placeholders, version-history dev notes), and stale version stamps. Ship a clean PRISMA 2020 checklist (27-item / 42-subitem table only) and an executed-method search-strategy doc, not the working drafts./self-review Phase 2.5c–2.5d (reference + cross-reference QC) over the supplementary files.Goal: Standardized pre-submission circulation of the manuscript to co-authors and senior methodologist / reviewer, with a bounded review window and a controlled attachment scope.
Trigger: Phase 8 is complete, and the draft has cleared Phase 6b source-fidelity audit.
Summary: Reply to the prior-version email thread to preserve In-Reply-To continuity (v1 → v2 → v3 tracked in one place). Attach the manuscript body with figures inline and, for v≥2, a change summary — exclude graphical abstract, cover letter, COI forms, and supplementary until the target journal is confirmed. TO = corresponding author + one senior methodologist; CC = remaining co-authors. Set a 7-day deadline (5 business days + weekend). Ask the corresponding author for target-journal preference, reviewer candidates, and cover-letter framing.
Load-on-demand procedural detail (thread continuity, attachment scope rationale, size-to-method table, journal-undetermined framing, response-tracking log): ${CLAUDE_SKILL_DIR}/references/phase9_circulation.md.
> Failure-mode cross-ref → references/review_orchestration.md RO-1~RO-5 (dual-rating completeness, defensive-tone bias audit, response-matrix numeric tracking, 2nd-reviewer availability blocking).
Goal: When an audit uncovers a structural data or protocol-application error, withdraw the current version, rebuild, and re-circulate with a transparent audit trail. Catching the error yourself before a journal reviewer does is the principal trust-building move in this phase.
Trigger conditions (any one):
| # | Trigger | Source | |---|---------|--------| | T1 | Extraction CSV ↔ primary source disagreement for a cell feeding a pooled/subgroup estimate or reported proportion | Phase 6b audit | | T2 | Included/excluded study violates the pre-specified criteria on re-read | Protocol review | | T3 | Hand-typed numerical literal in the analysis script traces to a wrong value | Phase 6b audit | | T4 | PROSPERO protocol ↔ delivered analysis disagreement on outcome, subgroup, or eligibility | Protocol ↔ analysis diff | | T5 | Dual-reviewer consensus record ↔ locked dataset disagreement on inclusion | Consensus log diff |
Non-negotiable rule: if the trigger fires after Phase 9 circulation but before journal submission, withdraw the current version within 24 hours. Reviewer discovery is a strictly worse failure mode than self-withdrawal.
Sprint outline (12 steps): (10.1) audit log at qc/audit_vN_to_vNplus1.md → (10.2) CSV re-verification with [VERIFY-CSV] tagging → (10.3) fresh script re-run (fixed seed, logged) → (10.4) manuscript auto-sync (grep for v{N} residue) → (10.5) supplementary regeneration (consensus log, RoB, GRADE/SoF, PRISMA flow) → (10.6) figure regeneration via /make-figures → (10.7) change summary with delta table → (10.8) PROSPERO amendment (application correction, not criteria change) → (10.9) re-circulation in the Phase 9 thread with the "On re-review" framing → (10.10) anti-patterns to avoid (hide-and-submit, "minor revision" reframe, cover-letter-only disclosure) → (10.11) post- submission escalation path → (10.12) post-recovery loop (Phase 9 restart; tighten Phase 6b if a second sprint is needed).
Load-on-demand procedural detail (exact audit-log fields, delta-table template, amendment language template, re-circulation paragraph template, anti-pattern rationale): ${CLAUDE_SKILL_DIR}/references/phase10_recovery.md.
> Failure-mode cross-ref → references/post_submission_release_ops.md Gate 4 covers reject/revise Zenodo versioning, tag-cleanup gate, and re-target workflow (avoid "new version" misuse on re-target).
Failure patterns observed across three prior MA projects (anonymized). Each topical reference extends the phase it cross-references above — consult alongside phase procedural docs, not in isolation.
| Domain | Phase span | Load-on-demand reference | |---|---|---| | Data integrity (2x2 arm-swap, KM audit, methodology mismatch, PRISMA 5-way drift, single-source k) | Phase 3 → 6 | references/data_integrity_checklist.md (DI-1~DI-9) | | Review orchestration (2nd-reviewer blocking, dual-rating completeness, defensive-tone audit, response-matrix tracking) | Phase 9 circulation (extends phase9_circulation.md) | references/review_orchestration.md (RO-1~RO-5) | | Submission package drift (multi-journal folder hygiene, DO_NOT_EDIT_HERE gate, build artifact vs master) | Phase 8 → submission | references/submission_package_drift.md | | Post-submission release ops (Zenodo DOI timing, tag-cleanup gate, reject-retarget versioning) | Submission → Phase 10 | references/post_submission_release_ops.md |
| When | Script | Gate | |---|---|---| | Phase 3f reconciliation (before Phase 5 write-up) | python3 ${CLAUDE_SKILL_DIR}/scripts/check_exclusion_code_validity.py --protocol 0_Protocol/protocol.md --screening 2_Screening/*.tsv --strict | validates each applied exclusion code against the registered eligibility criteria: CODE_CONTRADICTS_ELIGIBILITY (a code excludes a design the protocol includes — the bulk study-loss defect no arithmetic/inter-rater gate can see), CODE_NOT_REGISTERED (off-protocol code), CODE_RENUMBERED (same code, two meanings). Challenge card: scripts/check_exclusion_code_validity_challenge/. | | Phase 4 kickoff (before first extraction row) | python3 ${CLAUDE_SKILL_DIR}/../../scripts/extraction_consensus_log_init.py --output 2_Data/extraction_consensus_log.md | DI-1: creates standalone consensus log so comparative arm-specific rows are never folded into R-script comments. | | Phase 3f reconciliation + every revision touching PRISMA numbers | python3 ${CLAUDE_SKILL_DIR}/../../scripts/prisma_5way_consistency.py --ssot prisma.yaml | DI-6: 5-surface drift check (abstract / main text / flow figure / supplement / CSV) against YAML SSOT. Non-zero exit blocks Phase 5 writeup. | | Phase 8 pre-submission + every journal retarget | bash ${CLAUDE_SKILL_DIR}/../../scripts/tag_cleanup_gate.sh | DI-8: fails if VERIFY-CSV/TODO/FIXME/XXX survive in 7_Manuscript, supplement, SUBMISSION, etc. | | Phase 8 on first build per journal (--record), then before every re-submission (--verify) | python3 ${CLAUDE_SKILL_DIR}/../../scripts/verify_package_integrity.py --record --journal <name> then --verify --journal <name> | SPD: checksum-based drift detection between master manuscript and built SUBMISSION/{journal}/ folder. Journal-editable files (cover letter, response, MANIFEST, DO_NOT_EDIT_HERE.md) are auto-excluded. |
All four scripts are repo-shipped as of 2026-04 (FOLLOWUPS P10). Non-zero exit = gate failure; resolve before proceeding to the next phase.
Sixteen accumulated SR-MA peer-review / submission lessons (2026-05 and 2026-06) — the drivers behind the Phase 4 extraction-form schema, the Phase 4c QC scripts, and the Phase 8 submission gates. To keep this entry point lean they live load-on-demand in ${CLAUDE_SKILL_DIR}/references/empirical_lessons.md. Load that file when designing the extraction form (before Phase 4) and before submission (Phase 8) — it covers dual-extractor 2x2 integrity, cohort-overlap clustering, small-k subgroup caution, the supplementary 8-file bar, PROSPERO ID format, AI-disclosure presence, recompute-don't-copy sensitivity analyses, outcome harmonization, heterogeneous-RoB κ, survival-specific concerns, supplement blinding / de-scaffolding, self-contained reproducible analysis scripts, sidecar re-sync, methodological + software citations, wide-table PDF rendering, and submission-portal journal-identity checks.
| Pitfall | Problem | Solution | |---------|---------|----------| | Separate pooling of Se/Sp | Ignores correlation | Use bivariate/HSROC model | | Ignoring threshold effect | False heterogeneity | Check Spearman correlation, SROC plot | | Standard funnel plot for DTA | Inappropriate | Use Deeks' funnel plot | | I-squared only for heterogeneity | Doesn't capture threshold effect | Use prediction region on SROC | | Missing GRADE | Common omission in DTA MA | Apply GRADE-DTA. If <4 studies, assess each domain narratively and state the limitation explicitly | | Partial verification bias | Inflates sensitivity | QUADAS-3 3.2 (target condition assessed in all participants). QUADAS-3 has no Flow & Timing domain — that was QUADAS-2 | | Differential verification bias | Distorts both Se and Sp | QUADAS-3 3.3 (target condition assessed the same way in all participants) | | Unevaluable results excluded | Biases accuracy estimates | Report intent-to-diagnose analysis |
When the number of included studies is small (< 10):
| When | Call | Purpose | |------|------|---------| | Need literature search | /search-lit | PubMed/Semantic Scholar search with verified citations | | Need statistical code | /analyze-stats | Execute R/Python analysis scripts | | Need figures | /make-figures | PRISMA flow, forest plots, SROC, funnel plots | | Need reporting check | /check-reporting | PRISMA-DTA / PRISMA 2020 compliance (includes Step 4c registration / amendment timing) | | Need manuscript writing | /write-paper | Full IMRAD manuscript generation | | Need self-review | /self-review | Pre-submission quality check | | Self-audit recovery entrypoint (Phase 10) | /write-paper Step 7.4a | Recovery branch for polish pipelines that surface structural audit failures | | /sync-submission SR-MA gate | /sync-submission | Before submission, verify supplementary package matches all 8 files in templates/supplementary_8file_checklist.md (PRISMA, PROSPERO, search strategy, exclusion list, extraction table, per-study x per-domain RoB, subgroup forests, sensitivity / publication bias). AI Disclosure presence check (cross-link /peer-review Phase 2A P8). Cite-list duplicate check via /verify-refs Gate 5 (duplicate PMID/DOI). |
[VERIFY: variable_name] and ask the user to confirm against the data dictionary./search-lit for all citations.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 42,893 | 41,977 | -2% | 1 | 1 | 0% | 7,689 | 18,454 | +140% | 0 | 0 | — |
case-02 | fail→pass | 28,991 | 12,609 | -57% | 1 | 1 | 0% | 5,978 | 14,082 | +136% | 0 | 0 | — |
case-03 | fail→fail | 43,776 | 47,803 | +9% | 1 | 1 | 0% | 8,253 | 19,440 | +136% | 0 | 0 | — |
case-04 | pass→pass | 24,753 | 8,373 | -66% | 1 | 1 | 0% | 5,053 | 13,543 | +168% | 0 | 0 | — |
case-05 | pass→pass | 15,856 | 17,261 | +9% | 1 | 1 | 0% | 2,838 | 14,842 | +423% | 0 | 0 | — |
case-06 | pass→pass | 11,554 | 16,980 | +47% | 1 | 1 | 0% | 2,000 | 14,612 | +631% | 0 | 0 | — |
case-07 | pass→pass | 18,714 | 24,615 | +32% | 1 | 1 | 0% | 3,213 | 15,989 | +398% | 0 | 0 | — |
case-08 | pass→pass | 14,247 | 15,113 | +6% | 1 | 1 | 0% | 2,454 | 14,442 | +489% | 0 | 0 | — |
case-09 | pass→pass | 14,631 | 15,566 | +6% | 1 | 1 | 0% | 2,450 | 14,269 | +482% | 0 | 0 | — |
case-10 | fail→pass | 12,385 | 13,709 | +11% | 1 | 1 | 0% | 2,051 | 14,180 | +591% | 0 | 0 | — |
case-11 | pass→pass | 21,781 | 21,976 | +1% | 1 | 1 | 0% | 3,169 | 15,460 | +388% | 0 | 0 | — |
case-12 | pass→pass | 14,461 | 9,390 | -35% | 1 | 1 | 0% | 2,049 | 13,398 | +554% | 0 | 0 | — |
case-13 | pass→pass | 16,560 | 8,743 | -47% | 1 | 1 | 0% | 2,114 | 13,487 | +538% | 0 | 0 | — |
case-14 | fail→pass | 15,241 | 20,995 | +38% | 1 | 1 | 0% | 2,476 | 15,311 | +518% | 0 | 0 | — |
case-15 | fail→pass | 13,152 | 16,647 | +27% | 1 | 1 | 0% | 2,296 | 14,631 | +537% | 0 | 0 | — |
case-16 | fail→pass | 13,110 | 11,726 | -11% | 1 | 1 | 0% | 2,054 | 13,883 | +576% | 0 | 0 | — |
case-17 | pass→pass | 11,882 | 16,912 | +42% | 1 | 1 | 0% | 2,112 | 14,742 | +598% | 0 | 0 | — |
case-18 | pass→pass | 13,943 | 10,748 | -23% | 1 | 1 | 0% | 2,041 | 13,698 | +571% | 0 | 0 | — |
case-19 | pass→pass | 12,040 | 14,021 | +16% | 1 | 1 | 0% | 1,948 | 14,187 | +628% | 0 | 0 | — |
case-20 | fail→fail | 47,048 | 41,849 | -11% | 1 | 1 | 0% | 8,046 | 19,004 | +136% | 0 | 0 | — |
case-21 | fail→pass | 6,176 | 7,996 | +29% | 1 | 1 | 0% | 1,056 | 13,284 | +1158% | 0 | 0 | — |
case-22 | fail→fail | 29,653 | 26,834 | -10% | 1 | 1 | 0% | 5,946 | 16,526 | +178% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/24/2026 | +36% |
Other measured skills in the registry, with their headline benchmark lift.