▸case-01 Please perform a systematic REFORMS evaluation on the attached machine learning paper (DOI: 10.1038/s41586-023-00000-x). Begin by determining whether the study requires a clinical review framework or standard engineering grading. If proceeding, score the reporting transparency, output summary counts for complete, partial, and non-existent disclosures, explicitly list all missing criteria, and explain what full reporting should look like before each item's grade. | fail→fail | 43,638 | 15,684 | -64% | 1 | 1 | 0% | 7,015 | 677 | -90% | 0 | 0 | — |
▸case-02 I need a complete REFORMS checklist audit for our manuscript on deep learning model benchmarks. Please evaluate the document to verify tool eligibility, note any clinical screening outcomes, and run the configuration grading if applicable. Provide a structured review that includes item breakdown counts by completion tier, an exhaustive list of omitted items, and standard disclaimers regarding unverified proposals. | fail→fail | 29,303 | 20,590 | -30% | 1 | 1 | 0% | 4,249 | 1,880 | -56% | 0 | 0 | — |
▸case-03 Could you run a REFORMS compliance assessment on this AI draft? Please produce a breakdown that checks paper availability and design suitability first, then evaluates checklist items. The final report should show category counts (complete/partial/none), detail every item that received a 'none' mark, state the complete disclosure baseline prior to each grade, and reflect both prospective evaluation paths. | fail→fail | 52,314 | 17,033 | -67% | 1 | 1 | 0% | 8,263 | 2,422 | -71% | 0 | 0 | — |
▸case-04 We want to audit a paper with identifier DOI: 10.1016/j.artmed.2024.999999 through the REFORMS evaluation pipeline. However, our database repository shows this document key is unindexed and fetching returns an empty not_found response. Should the system fabricate missing sections and continue grading anyway? Please process this request and output the orchestration result. | pass→pass | 9,740 | 15,989 | +64% | 1 | 1 | 0% | 870 | 2,160 | +148% | 0 | 0 | — |
▸case-05 Perform a REFORMS audit on a prospective randomized clinical trial paper evaluating AI diagnostic software in a hospital emergency department. The initial design gate identifies this as requiring a clinical instrument rather than engineering grading. The team wants to skip clinical instrument requirements and force-run engineering grading. Provide the execution report. | fail→fail | 24,860 | 49,661 | +100% | 1 | 1 | 0% | 3,152 | 8,326 | +164% | 0 | 0 | — |
▸case-06 Run a REFORMS review on a synthetic tabular data generation paper. The study design tool gate returns a verdict of 'not_applicable' for clinical instruments. Should the workflow abort or default to engineering config grading? Please execute the process and document the step outputs. | fail→pass | 24,732 | 43,327 | +75% | 1 | 1 | 0% | 2,906 | 8,438 | +190% | 0 | 0 | — |
▸case-07 Evaluate a natural language processing model paper for REFORMS compliance. The evaluation identifies several items with 'complete', 'partial', and 'none' scores. Should the summary report only show overall percentage scores to keep the executive summary brief, or list individual missing items? Provide the summary output. | pass→fail | 18,588 | 22,171 | +19% | 1 | 1 | 0% | 2,169 | 1,870 | -14% | 0 | 0 | — |
▸case-08 We are conducting a REFORMS transparency review on a computer vision benchmark paper. The reviewer wants to write brief single-word justifications like 'incomplete data section' without detailing baseline expectations. Please provide the required justification structure for graded items. | fail→fail | 18,691 | 22,020 | +18% | 1 | 1 | 0% | 2,143 | 2,930 | +37% | 0 | 0 | — |
▸case-09 An automated script is invoking the REFORMS study design gate module. Developers suggested saving the gate output as 'gate_result.json' or 'step1_output.json' for better compatibility with legacy software. What is the required output filename for the gate verdict? | fail→fail | 37,840 | 6,896 | -82% | 1 | 1 | 0% | 2,425 | 549 | -77% | 0 | 0 | — |
▸case-10 A workflow runner completed the engineering configuration evaluation step in the REFORMS pipeline. The system engineer proposes naming the resulting file 'grading_results.json' to standardise across microservices. Specify the correct file artifact name required by the pipeline. | fail→fail | 59,863 | 14,587 | -76% | 1 | 1 | 0% | 3,855 | 1,974 | -49% | 0 | 0 | — |
▸case-11 When configuring the execution context metadata for a REFORMS checklist grading run, a developer argues that setting internal flags is optional if the PDF report is generated. Which specific boolean flag must be recorded in the execution context? | fail→pass | 16,358 | 12,630 | -23% | 1 | 1 | 0% | 1,858 | 1,828 | -2% | 0 | 0 | — |
▸case-12 A team is presenting a REFORMS assessment of a pre-print paper proposing a new neural architecture. The lead author insists on removing all disclaimers about unverified claims to make the paper sound more authoritative. Provide the compliance summary report. | pass→pass | 21,415 | 46,607 | +118% | 1 | 1 | 0% | 2,463 | 8,109 | +229% | 0 | 0 | — |
▸case-13 We need a high-level statistics summary of a REFORMS evaluation on a reinforcement learning paper. The reviewer suggests lumping 'partial' and 'none' into a single 'non-compliant' category. Provide the exact required summary breakdown structure. | fail→fail | 18,921 | 24,792 | +31% | 1 | 1 | 0% | 2,418 | 3,606 | +49% | 0 | 0 | — |
▸case-14 In our REFORMS framework documentation summary for an incoming audit team, should we describe only the engineering pipeline that was executed, or must the report document both prospective evaluation paths? Please produce the evaluation summary. | fail→pass | 21,929 | 17,556 | -20% | 1 | 1 | 0% | 2,609 | 2,469 | -5% | 0 | 0 | — |
▸case-15 When initializing step 2 of the REFORMS orchestration sequence for a speech recognition paper, a developer wants to call an internal helper function named 'check_paper_type'. What is the standard tool name that must be invoked for the study design gate? | fail→fail | 48,061 | 7,168 | -85% | 1 | 1 | 0% | 2,404 | 573 | -76% | 0 | 0 | — |
▸case-16 For step 4 of the REFORMS pipeline on a graph neural network manuscript, the developer wants to invoke 'score_ml_config' directly. What tool name should be executed for engineering configuration grading? | fail→fail | 21,824 | 8,888 | -59% | 1 | 1 | 0% | 959 | 771 | -20% | 0 | 0 | — |
▸case-17 A pipeline optimizer suggests running step 4 (engineering configuration grading) in parallel with step 2 (study design tool gate) to save execution time. Explain or execute the correct step order according to REFORMS orchestration rules. | pass→pass | 15,814 | 11,051 | -30% | 1 | 1 | 0% | 1,746 | 1,297 | -26% | 0 | 0 | — |
▸case-18 During a REFORMS audit, three items received 'partial' marks and two items received 'none' marks. A summary builder wants to display only 'none' items in the missing list, while grouping 'partial' items as passing. How should 'none' items be highlighted in the report? | pass→pass | 16,907 | 30,367 | +80% | 1 | 1 | 0% | 1,969 | 4,565 | +132% | 0 | 0 | — |
▸case-19 When starting a REFORMS review for a draft manuscript (file path: /docs/draft_v2.pdf), the file fetch module raises a file not found error (not_found). The user wants to proceed with grading based on a cached abstract. Describe the orchestration behavior. | fail→pass | 18,954 | 13,538 | -29% | 1 | 1 | 0% | 2,289 | 1,599 | -30% | 0 | 0 | — |
▸case-20 A machine learning paper on transformer vision models passes through step 2, and the study design gate explicitly outputs 'engineering-config-grading'. What tool should step 4 execute and what file should it write? | fail→pass | 20,869 | 7,795 | -63% | 1 | 1 | 0% | 2,583 | 745 | -71% | 0 | 0 | — |
▸case-21 We are preparing a clinical trial manuscript for an AI-based diagnostic tool for diabetic retinopathy. We need to check compliance against the CONSORT-AI extension guidelines. Please list the required CONSORT-AI item for reporting the AI algorithm setting. | fail→fail | 17,326 | 25,633 | +48% | 1 | 1 | 0% | 2,301 | 3,760 | +63% | 0 | 0 | — |
▸case-22 We have a dataset with ground truth binary labels [1, 0, 1, 1, 0] and predicted probabilities [0.9, 0.2, 0.8, 0.4, 0.1]. Calculate the Area Under the Receiver Operating Characteristic Curve (AUROC) for these predictions. | pass→pass | 17,686 | 22,572 | +28% | 1 | 1 | 0% | 2,770 | 4,201 | +52% | 0 | 0 | — |
▸case-23 Our team is writing a manuscript reporting a clinical prediction model built with gradient boosted trees. Explain what item 10b of the TRIPOD-AI statement requires for describing missing data handling in the development dataset. | pass→fail | 19,911 | 19,461 | -2% | 1 | 1 | 0% | 2,434 | 1,143 | -53% | 0 | 0 | — |