▸case-01 We are conducting a meta-analysis on 12 randomized controlled trials measuring depression scores using different scales (HAM-D, BDI). We need to calculate the primary pooled summary effect size across all studies. A team member suggested using a simple unweighted arithmetic mean of study mean differences. How should the primary pooled effect size be calculated for continuous outcomes across varying measurement scales? | pass→pass | 18,083 | 17,172 | -5% | 1 | 1 | 0% | 2,271 | 2,417 | +6% | 0 | 0 | — |
▸case-02 We are launching a systematic review on SGLT2 inhibitors and cardiovascular outcomes in type 2 diabetes. We need a search strategy for MEDLINE/PubMed. Someone suggested typing 'SGLT2 diabetes heart' into the search bar. Provide the formal search strategy structure required for systematic literature retrieval. | pass→pass | 21,523 | 22,354 | +4% | 1 | 1 | 0% | 3,115 | 3,440 | +10% | 0 | 0 | — |
▸case-03 We are evaluating 8 non-randomized observational studies comparing bariatric surgery versus medical management for NASH. A reviewer suggested using the Cochrane RoB 2.0 tool for all studies. How should methodological quality be assessed for non-randomized studies of interventions? | fail→fail | 14,672 | 21,266 | +45% | 1 | 1 | 0% | 2,433 | 2,891 | +19% | 0 | 0 | — |
▸case-04 In our meta-analysis of 18 clinical trials for hypertension, two studies show extreme effect sizes that skew the pooled estimate. An analyst recommends permanently deleting these two studies from the main review findings. How should the plan address identified potential outliers without corrupting the primary synthesis? | pass→pass | 20,429 | 17,152 | -16% | 1 | 1 | 0% | 2,383 | 2,652 | +11% | 0 | 0 | — |
▸case-05 Our primary meta-analysis includes 25 trials on stroke rehabilitation, 7 of which are judged to be at high risk of bias due to lack of blinding. The team wants to restrict the primary synthesis exclusively to low-risk studies from the start. How should high-risk studies be handled when structuring the robustness testing plan? | fail→fail | 20,006 | 14,102 | -30% | 1 | 1 | 0% | 2,365 | 2,583 | +9% | 0 | 0 | — |
▸case-06 In a systematic review of antiplatelet therapy, substantial clinical heterogeneity is expected. A colleague argues that because a random-effects model was selected as primary, no model comparison is needed. How should model assumption dependence be evaluated in the sensitivity plan? | pass→pass | 23,258 | 23,427 | +1% | 1 | 1 | 0% | 2,689 | 3,233 | +20% | 0 | 0 | — |
▸case-07 For 5 out of 20 trials in our pain management meta-analysis, standard deviations were not reported and had to be estimated using baseline SDs and correlation coefficients. An author suggests treating estimated SDs as exact known constants without further check. How should missing parameter imputation be evaluated in the analysis plan? | pass→pass | 22,209 | 14,379 | -35% | 1 | 1 | 0% | 2,664 | 2,441 | -8% | 0 | 0 | — |
▸case-08 To check if small-study effects or publication bias distort our meta-analytic outcome, a team member suggests relying solely on Rosenthal's fail-safe N threshold. How should publication bias sensitivity be designed to overcome known flaws of fail-safe N? | fail→pass | 17,591 | 23,504 | +34% | 1 | 1 | 0% | 2,973 | 3,377 | +14% | 0 | 0 | — |
▸case-09 We plan to test whether treatment response differs between pediatric and adult cohorts across 14 trials. A reviewer suggests running separate meta-analyses in each subgroup and declaring a difference if one subgroup p-value is significant and the other is not. How should subgroup sensitivity analyses be specified? | fail→fail | 21,749 | 18,528 | -15% | 1 | 1 | 0% | 2,879 | 3,205 | +11% | 0 | 0 | — |
▸case-10 Our meta-analysis on sepsis mortality includes studies with minimum follow-up ranging from 28 days to 1 year. The lead author wants to arbitrarily exclude any study under 90 days follow-up. How should uncertainty in trial inclusion criteria thresholding be structured in the sensitivity framework? | pass→pass | 21,747 | 21,110 | -3% | 1 | 1 | 0% | 2,689 | 2,801 | +4% | 0 | 0 | — |
▸case-11 Three cluster-randomized trials in our community health meta-analysis did not report intra-class correlation coefficients (ICCs), so we borrowed an ICC of 0.05 from literature. A statistician suggested using this single ICC value as definitive. How should unmeasured design effect parameters be tested for sensitivity? | pass→pass | 16,779 | 25,208 | +50% | 1 | 1 | 0% | 2,837 | 3,576 | +26% | 0 | 0 | — |
▸case-12 When identifying influential studies in a meta-analysis of 30 vaccine trials, a researcher proposes selecting studies for sensitivity analysis based purely on visual inspection of the forest plot. What formal diagnostic metrics should be specified in the plan to identify candidate studies for sensitivity testing? | pass→pass | 17,374 | 20,474 | +18% | 1 | 1 | 0% | 2,797 | 2,800 | +0% | 0 | 0 | — |
▸case-13 In a meta-analysis of multi-arm trials, covariance between multiple treatment arms sharing a single control group requires an assumed correlation coefficient rho = 0.5. A co-author states that rho = 0.5 is standard so no further verification is required. How should multi-arm correlation assumptions be stress-tested? | fail→fail | 22,444 | 20,668 | -8% | 1 | 1 | 0% | 2,848 | 2,750 | -3% | 0 | 0 | — |
▸case-14 Our meta-analysis pools observational cohort studies on air pollution and asthma hospitalizations. A referee asks how robust the observed relative risk of 1.25 is against potential unmeasured confounding. What quantitative sensitivity approach should be included in the protocol? | pass→pass | 21,576 | 21,442 | -1% | 1 | 1 | 0% | 2,800 | 3,067 | +10% | 0 | 0 | — |
▸case-15 In a psychiatric intervention meta-analysis, several trials had attrition rates exceeding 20%. The team performed a complete-case analysis as primary. Someone proposes assuming all dropped participants had successful outcomes. How should missing outcome data assumptions be tested in sensitivity analysis? | pass→pass | 21,388 | 20,082 | -6% | 1 | 1 | 0% | 2,719 | 2,496 | -8% | 0 | 0 | — |
▸case-16 In a synthesis of 22 oncology trials, heterogeneity is high (I-squared = 78%). An analyst suggests randomly removing 3 trials until I-squared drops below 50%. How should contribution to heterogeneity be systematically mapped for sensitivity evaluation? | fail→fail | 20,258 | 15,377 | -24% | 1 | 1 | 0% | 2,476 | 2,762 | +12% | 0 | 0 | — |
▸case-17 A meta-analysis on dietary interventions includes peer-reviewed journal articles and unpublished conference abstracts. A contributor suggests excluding all conference abstracts from the entire review project. How should publication status eligibility be evaluated within the analysis protocol? | pass→pass | 17,262 | 16,938 | -2% | 1 | 1 | 0% | 2,532 | 2,006 | -21% | 0 | 0 | — |
▸case-18 When meta-analyzing osteoarthritis pain relief, trials measured outcomes at different timepoints between 4 weeks and 12 weeks. The primary outcome was defined as 8-week pain. How should sensitivity analysis handle studies that reported outcomes slightly outside this exact target window? | pass→pass | 21,719 | 21,126 | -3% | 1 | 1 | 0% | 2,604 | 2,746 | +5% | 0 | 0 | — |
▸case-19 For a meta-analysis with few studies (N=5) and rare events, the primary model uses DerSimonian-Laird random effects. A reviewer points out that DerSimonian-Laird underestimates variance when study count is low. How should small-sample estimator risk be evaluated in the sensitivity plan? | pass→pass | 24,489 | 23,646 | -3% | 1 | 1 | 0% | 3,252 | 3,256 | +0% | 0 | 0 | — |
▸case-20 In a safety meta-analysis of surgical complications, 6 studies have zero events in one arm. The primary approach adds a constant 0.5 continuity correction to all sparse cells. An investigator suggests no sensitivity testing is needed because 0.5 is standard software default. How should sparse event adjustments be stress-tested? | pass→fail | 23,277 | 22,359 | -4% | 1 | 1 | 0% | 3,070 | 2,717 | -11% | 0 | 0 | — |
▸case-21 In a meta-analysis of pharmaceutical efficacy, 10 studies were industry-funded and 5 were publicly funded. A team member proposes omitting industry-funded trials completely from the review. How should potential funding source bias be systematically addressed? | pass→pass | 20,559 | 18,107 | -12% | 1 | 1 | 0% | 2,375 | 2,312 | -3% | 0 | 0 | — |
▸case-22 After completing five sensitivity analyses testing outlier removal, RoB restriction, and model choice, how should the conclusions and sensitivity results be structured in the final synthesis report? | pass→pass | 26,708 | 19,193 | -28% | 1 | 1 | 0% | 2,801 | 2,439 | -13% | 0 | 0 | — |