▸case-01 I'm drafting a study proposal on whether mindfulness app usage reduces burnout in junior doctors. Can you evaluate the strength of this specific research question using a standard five-point appraisal checklist? Give me a structured quality assessment of the question itself, focusing on its formulation rather than trying to map it into PICO slots or assessing study outcomes. | fail→fail | 20,306 | 15,766 | -22% | 1 | 1 | 0% | 2,382 | 1,879 | -21% | 0 | 0 | — |
▸case-02 We are writing a systematic review protocol and need to assess a candidate research question: 'Does fine-tuning LLMs on synthetic domain data improve diagnostic accuracy in rare disease triage without compromising safety?' Please perform a quality appraisal on this proposed question across its core evaluative criteria. Output a detailed judgment report assessing the validity and soundness of the question itself. | pass→fail | 29,961 | 31,022 | +4% | 1 | 1 | 0% | 4,013 | 1,485 | -63% | 0 | 0 | — |
▸case-03 Please appraise the following research question for our grant application: 'To what extent does gamified peer-review improve critical thinking skills in undergraduate online courses?' I need an independent checklist evaluation of the question's quality and viability before we execute the trial, keeping the appraisal strictly focused on the question's formulation rather than study results or execution details. | fail→pass | 20,698 | 18,935 | -9% | 1 | 1 | 0% | 2,514 | 2,396 | -5% | 0 | 0 | — |
▸case-04 We need to structure our clinical trial protocol for comparing high-flow nasal cannula versus non-invasive ventilation in acute respiratory failure. Please break down our core inquiry into Population, Intervention, Comparison, and Outcome slots for registration in ClinicalTrials.gov. | pass→pass | 21,895 | 18,462 | -16% | 1 | 1 | 0% | 3,060 | 2,655 | -13% | 0 | 0 | — |
▸case-05 We are conducting a systematic review on cognitive behavioral therapy for insomnia. Here is a published randomized trial (Smith et al., 2023) detailing baseline randomization, blinding procedures, and loss to follow-up. Please rate the risk of bias in the study's methodological execution using the Cochrane Risk of Bias 2 tool. | fail→fail | 11,923 | 40,941 | +243% | 1 | 1 | 0% | 1,348 | 4,345 | +222% | 0 | 0 | — |
▸case-06 We have compiled pooled effect sizes and confidence intervals from ten clinical trials evaluating SGLT2 inhibitors in chronic kidney disease. Please evaluate the overall certainty of this evidence body across risk of bias, inconsistency, indirectness, imprecision, and publication bias using the GRADE approach. | pass→pass | 19,334 | 32,038 | +66% | 1 | 1 | 0% | 2,997 | 5,120 | +71% | 0 | 0 | — |
▸case-07 Our oncology group wants to evaluate: 'Does adjuvant immunotherapy prolong 5-year overall survival in resected stage III melanoma patients compared to observation?' Many team members want us to just fill out the PICO slots and call that our question assessment. Please perform a formal quality appraisal of this research question. | pass→pass | 24,915 | 31,366 | +26% | 1 | 1 | 0% | 3,316 | 4,162 | +26% | 0 | 0 | — |
▸case-08 A proposed study asks: 'Does a Mediterranean diet intervention reduce HbA1c in type 2 diabetes patients over 12 months?' Some reviewers want us to evaluate whether the diet will actually work and produce statistically significant HbA1c reductions. Please provide the appropriate question-level quality appraisal without evaluating expected study results. | pass→fail | 17,873 | 27,507 | +54% | 1 | 1 | 0% | 1,985 | 1,274 | -36% | 0 | 0 | — |
▸case-09 We are submitting a proposal: 'Can smartphone-based ECG monitoring detect silent atrial fibrillation in rural hypertension clinics?' Reviewers asked whether our nursing staff training schedule and battery supply logistics are feasible. We want to separate the checklist quality appraisal of the research question itself from the assessment of study execution details. | pass→pass | 22,192 | 15,789 | -29% | 1 | 1 | 0% | 2,701 | 1,855 | -31% | 0 | 0 | — |
▸case-10 We are preparing a grant for a health informatics project asking: 'Does automated EHR alerting reduce inappropriate vancomycin ordering in pediatric ICUs?' Execute a full quality assessment of this research question formulation. | fail→fail | 27,100 | 46,348 | +71% | 1 | 1 | 0% | 3,218 | 6,698 | +108% | 0 | 0 | — |
▸case-11 We have a candidate question for a mental health trial: 'Does exercise therapy reduce depressive symptoms in adolescents?' Our advisor suggested merging the PICO slot mapping (Patient, Intervention, Comparator, Outcome) and the FINER appraisal into a single combined table. Please perform the question quality appraisal maintaining proper framework separation. | pass→pass | 22,465 | 23,271 | +4% | 1 | 1 | 0% | 2,811 | 3,094 | +10% | 0 | 0 | — |
▸case-12 An environmental health study asks: 'Does long-term exposure to ambient PM2.5 increase cardiovascular mortality in elderly urban populations?' Analysts want to fill out the PECO (Population, Exposure, Comparator, Outcome) fields. Perform the research question quality appraisal, ensuring you do not confuse slot filling with question appraisal. | pass→pass | 21,468 | 20,927 | -3% | 1 | 1 | 0% | 2,607 | 2,996 | +15% | 0 | 0 | — |
▸case-13 A qualitative health research question asks: 'How do nurse practitioners experience moral distress in rural emergency departments?' Some researchers suggested using SPIDER to judge question quality. Perform a quality appraisal of this research question using the appropriate checklist framework for question appraisal rather than qualitative slot framing. | pass→pass | 22,135 | 17,613 | -20% | 1 | 1 | 0% | 2,688 | 2,187 | -19% | 0 | 0 | — |
▸case-14 A clinical trial lead wants to rate the question 'Does early mobility therapy shorten ICU length of stay?' using a simplified 3-point scale: Clarity, Relevance, and Cost. Provide the standard full quality appraisal of this research question. | pass→pass | 20,166 | 25,689 | +27% | 1 | 1 | 0% | 2,469 | 3,442 | +39% | 0 | 0 | — |
▸case-15 An investigator asks: 'Does daily hyperbaric oxygen therapy reverse cognitive decline in early-stage Alzheimer's disease?' The team is worried about feasibility and wants to only judge feasibility. Provide a complete five-part question quality appraisal so all key evaluative dimensions are covered. | pass→pass | 22,347 | 29,571 | +32% | 1 | 1 | 0% | 2,784 | 3,987 | +43% | 0 | 0 | — |
▸case-16 A proposed study asks: 'Does continuous glucose monitoring improve glycemic control in pregnant women with gestational diabetes compared to standard fingerstick monitoring?' The review board is confusing the novelty of the research question with the technical difficulty of sensor calibration. Appraise the research question itself. | pass→pass | 19,558 | 27,707 | +42% | 1 | 1 | 0% | 2,469 | 2,383 | -3% | 0 | 0 | — |
▸case-17 Consider the research question: 'Does withholding standard antiretroviral therapy in asymptomatic HIV patients slow immune exhaustion?' Provide a formal quality appraisal of this proposed research question. | pass→pass | 21,304 | 24,817 | +16% | 1 | 1 | 0% | 2,581 | 3,187 | +23% | 0 | 0 | — |
▸case-18 A cardiology question states: 'Does adding spironolactone to standard heart failure therapy reduce 30-day readmissions in elderly patients?' A reviewer claims we should wait for trial outcome data before judging question relevance. Appraise the research question formulation now. | pass→pass | 20,092 | 20,145 | +0% | 1 | 1 | 0% | 2,412 | 2,574 | +7% | 0 | 0 | — |
▸case-19 Appraise the research question: 'Does an audit-and-feedback implementation strategy increase guideline-concordant antibiotic prescribing in urgent care clinics?' Please carry out this quality assessment. | fail→fail | 19,630 | 22,151 | +13% | 1 | 1 | 0% | 2,353 | 3,004 | +28% | 0 | 0 | — |
▸case-20 We need an appraisal of the research question: 'Is telemedicine follow-up cost-effective compared to in-person clinic visits for post-operative orthopedic patients in rural counties?' Evaluate the question formulation. | fail→fail | 19,130 | 43,918 | +130% | 1 | 1 | 0% | 2,337 | 4,450 | +90% | 0 | 0 | — |
▸case-21 Assess the research question: 'Does virtual reality exposure therapy reduce PTSD symptoms in combat veterans?' Ensure each criteria judgment in the appraisal is independent and not collapsed into a single pass/fail score or study outcome prediction. | pass→pass | 19,208 | 23,354 | +22% | 1 | 1 | 0% | 2,454 | 3,133 | +28% | 0 | 0 | — |
▸case-22 We are evaluating the research question: 'Does school-based mindfulness intervention reduce anxiety symptoms in middle school students?' The protocol draft currently mixes PICO slot definitions, FINER judgments, and trial execution steps in one bulleted list. Provide a clean quality appraisal focused solely on question judgment. | pass→pass | 18,884 | 13,981 | -26% | 1 | 1 | 0% | 2,258 | 1,456 | -36% | 0 | 0 | — |