▸case-21 Analyze the machine learning hypothesis: 'Fine-tuning Llama-3-8B on 50,000 domain-specific medical QA pairs will yield a higher USMLE Step 1 score than prompting base GPT-4o zero-shot.' Output JSON with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 11,345 | 4,961 | -56% | 1 | 1 | 0% | 2,095 | 1,338 | -36% | 0 | 0 | — |
▸case-01 I want to evaluate the hypothesis: 'A weekly 10-minute cold plunge reduces systemic inflammation markers in healthy adults.' Please run a scientific refutation analysis on this. Generate a JSON object containing the ID assigned to the hypothesis, the statement itself, two or more positive predictions, a specific concrete observation that would disprove the claim, an assessment of how testable it is, a final verdict outcome, recommendations for revising it if necessary, and relevant notes. | fail→pass | 12,369 | 11,610 | -6% | 1 | 1 | 0% | 2,054 | 1,515 | -26% | 0 | 0 | — |
▸case-02 Can you review my product hypothesis: 'Adding social proof badges to the landing page increases conversion rates among first-time visitors.' Provide your analysis as a JSON string detailing the hypothesis identifier, the statement, a list of observable predictions, a refutation condition describing what observation proves it wrong, testability level, overall verdict status, advice on how to rewrite it if it lacks falsifiability, and additional notes. | fail→pass | 7,744 | 5,093 | -34% | 1 | 1 | 0% | 1,473 | 1,352 | -8% | 0 | 0 | — |
▸case-03 Assess whether this statement is scientifically testable: 'Universal basic income improves overall human happiness.' Produce a structured JSON response with fields for the hypothesis code, full statement, positive expected outcomes, a refuting scenario that would disprove the premise, technical and ethical testability rating, the logical verdict classification, suggested edits if it isn't falsifiable, and any supplementary notes. | fail→fail | 12,554 | 9,002 | -28% | 1 | 1 | 0% | 2,240 | 1,828 | -18% | 0 | 0 | — |
▸case-04 I am working on a clinical trial idea regarding cardiovascular health, but I do not have a formulated hypothesis statement yet. Please execute a refutation analysis and output JSON detailing the hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 10,662 | 2,229 | -79% | 1 | 1 | 0% | 1,896 | 772 | -59% | 0 | 0 | — |
▸case-05 Evaluate the falsifiability of this claim: 'An undetectable immaterial entity controls quantum spin state flips without leaving any physical trace or measurable signal.' Output a JSON object with fields hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→fail | 8,176 | 5,799 | -29% | 1 | 1 | 0% | 1,402 | 1,280 | -9% | 0 | 0 | — |
▸case-06 Analyze the statement: 'Either it will rain tomorrow in Seattle or it will not rain tomorrow in Seattle.' Evaluate whether this claim can be refuted by empirical observation, providing a JSON report with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 9,230 | 6,540 | -29% | 1 | 1 | 0% | 1,620 | 1,488 | -8% | 0 | 0 | — |
▸case-07 Evaluate the statement: 'Abstract expressionist painting is inherently superior to Renaissance realism.' Provide a refutation analysis formatted as JSON containing hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 7,285 | 5,886 | -19% | 1 | 1 | 0% | 1,136 | 1,318 | +16% | 0 | 0 | — |
▸case-08 Assess the scientific hypothesis: 'Objects of different masses fall with identical accelerations in a vacuum near Earth's surface.' Provide a refutation assessment in JSON with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 9,750 | 6,011 | -38% | 1 | 1 | 0% | 1,703 | 1,373 | -19% | 0 | 0 | — |
▸case-09 Run a refutation analysis on the clinical statement: 'Daily administration of 500mg metformin reduces fasting blood glucose by at least 15 mg/dL in type 2 diabetic adults over 12 weeks compared to placebo.' Format as a JSON object with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 11,147 | 7,585 | -32% | 1 | 1 | 0% | 2,094 | 1,921 | -8% | 0 | 0 | — |
▸case-10 Analyze the claim: 'Plant-based diets make people feel much better.' Format your analysis into JSON with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 9,647 | 7,303 | -24% | 1 | 1 | 0% | 1,754 | 1,767 | +1% | 0 | 0 | — |
▸case-11 Calculate the required sample size per variation for an e-commerce checkout A/B test with a baseline conversion rate of 5%, a minimum detectable effect of 10% relative uplift (to 5.5%), 80% statistical power, and a significance level alpha of 0.05. Show your formula and calculation. | pass→fail | 11,385 | 6,635 | -42% | 1 | 1 | 0% | 2,609 | 1,624 | -38% | 0 | 0 | — |
▸case-12 Draft an IRB-compliant clinical research protocol outline for a Phase II trial investigating a novel monoclonal antibody for mild-to-moderate Alzheimer's disease, including study design, inclusion/exclusion criteria, primary/secondary endpoints, and safety monitoring. | pass→fail | 30,902 | 11,667 | -62% | 1 | 1 | 0% | 5,612 | 2,331 | -58% | 0 | 0 | — |
▸case-13 Formulate a systematic PubMed boolean search string and search strategy to identify randomized controlled trials published between 2015 and 2024 regarding sleep hygiene interventions for adolescent insomnia. | pass→fail | 13,701 | 4,619 | -66% | 1 | 1 | 0% | 2,667 | 1,092 | -59% | 0 | 0 | — |
▸case-14 Check the falsifiability of: 'Changing the checkout button color from blue to green will increase completed transactions by 3% on our online retail site.' Output JSON with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | pass→pass | 9,155 | 6,878 | -25% | 1 | 1 | 0% | 1,707 | 1,590 | -7% | 0 | 0 | — |
▸case-15 Evaluate this engineering claim: 'Replacing Redis with an in-memory LRU cache in Go will reduce P99 API response latency by at least 20ms under 10,000 QPS load.' Return JSON containing hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 15,969 | 5,817 | -64% | 1 | 1 | 0% | 2,620 | 1,464 | -44% | 0 | 0 | — |
▸case-16 Assess the hypothesis: 'Mercury in retrograde causes cloud server downtime unless system engineers project positive mental thoughts.' Return a JSON object with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→fail | 8,591 | 8,524 | -1% | 1 | 1 | 0% | 1,492 | 1,741 | +17% | 0 | 0 | — |
▸case-17 Assess the statement: 'If Napoleon had avoided battle at Waterloo, the French Empire would have survived continuously into the 20th century.' Output JSON with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 12,902 | 8,781 | -32% | 1 | 1 | 0% | 2,197 | 1,864 | -15% | 0 | 0 | — |
▸case-18 Evaluate the theoretical physics hypothesis: 'Spacetime contains 11 compactified dimensions at the Planck scale that leave no observable signatures below 10^19 GeV collision energy.' Format as JSON with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | pass→fail | 11,826 | 12,416 | +5% | 1 | 1 | 0% | 2,025 | 2,502 | +24% | 0 | 0 | — |
▸case-19 I have two hypotheses: 1) 'Supplementing 5000 IU Vitamin D3 daily reduces self-reported winter upper respiratory infection frequency by 20% in adults living above 45 degrees north latitude.' 2) 'Vitamin D enhances cosmic aura.' Please evaluate hypothesis H1 according to falsifiability standards and output JSON with hypothesis_id set to 'H1', statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | pass→pass | 9,516 | 6,441 | -32% | 1 | 1 | 0% | 1,941 | 1,657 | -15% | 0 | 0 | — |
▸case-20 Evaluate the behavioral economics claim: 'Offering a 10% discount with a 24-hour countdown timer yields a higher conversion rate than offering an unconditional 15% discount without a timer.' Return a JSON object with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 11,534 | 5,085 | -56% | 1 | 1 | 0% | 1,934 | 1,331 | -31% | 0 | 0 | — |
▸case-22 Check the falsifiability of: 'Inoculating corn seeds with mycorrhizal fungi increases kernel yield per acre by at least 8% under drought conditions (<15 inches seasonal rainfall).' Format output as JSON with hypothesis_id, statement, positive_predictions, falsification_scenario, testability, verdict, revision_suggestion, and notes. | fail→pass | 8,397 | 7,719 | -8% | 1 | 1 | 0% | 1,560 | 1,868 | +20% | 0 | 0 | — |