▸case-10 In our survey experiment about corporate ESG claims, we are worried respondents will guess our hypothesis and report pro-environmental attitudes to look good. We plan to rely on a standard direct post-treatment question. What demand-mitigation design features should be built into our hypothesis and design architecture? | fail→pass | 19,650 | 27,366 | +39% | 1 | 1 | 0% | 3,131 | 8,369 | +167% | 0 | 0 | — |
▸case-17 We are setting up a field experiment sending encouraging emails to registered voters. Before drafting our hypotheses, what core methodological checks must we conduct regarding causal inference? | fail→pass | 16,116 | 16,135 | +0% | 1 | 1 | 0% | 2,651 | 6,221 | +135% | 0 | 0 | — |
▸case-18 We are designing a survey experiment on minimum wage policy attitudes using a 2x2 factorial design. What graphical tool and path check should we use prior to formalizing our hypotheses? | fail→pass | 15,592 | 12,166 | -22% | 1 | 1 | 0% | 2,779 | 5,656 | +104% | 0 | 0 | — |
▸case-19 When writing our pre-analysis plan for a conjoint experiment on housing policy, how should we formally define the target quantity versus the empirical estimator? | pass→pass | 17,973 | 17,856 | -1% | 1 | 1 | 0% | 3,326 | 6,629 | +99% | 0 | 0 | — |
▸case-20 We are formulating hypotheses for a political campaign framing experiment. Is it better to test our primary frame against a pure control group or against an active competing campaign message? | pass→pass | 12,196 | 16,622 | +36% | 1 | 1 | 0% | 2,094 | 6,506 | +211% | 0 | 0 | — |
▸case-21 We are measuring the prevalence of vote-buying in a local election using a list experiment alongside a direct survey question. How should this hypothesis be categorized and specified? | pass→pass | 12,445 | 17,150 | +38% | 1 | 1 | 0% | 2,252 | 6,734 | +199% | 0 | 0 | — |
▸case-22 We are conducting a 10-country study on trust in democratic institutions. We expect the treatment effect to vary slightly across countries depending on corruption levels, but we do not have specific directional predictions or per-country power for all 10 nations. How should we treat these cross-country variations in our hypothesis structure? | fail→pass | 12,639 | 18,071 | +43% | 1 | 1 | 0% | 2,232 | 6,607 | +196% | 0 | 0 | — |
▸case-01 I am preparing a pre-analysis plan for a survey experiment evaluating how presenting local vs. global economic costs alters public support for carbon taxation. Could you draft a formal causal hypothesis framework for this study? I need a breakdown that articulates the underlying data generation logic, defines the precise counterfactual and estimand, lays out the core hypothesis at conceptual, operational, and statistical levels, specifies the smallest meaningful effect size, and classifies our predictions into primary, secondary, and exploratory tiers. | fail→pass | 34,872 | 31,020 | -11% | 1 | 1 | 0% | 6,237 | 9,735 | +56% | 0 | 0 | — |
▸case-08 We are running a two-experiment study on public support for defense spending. Experiment 1 tests H1 (threat framing vs control) and H2 (cost framing vs control). In Experiment 2, we test whether combining threat and cost framing produces a synergistic effect. Should we label the Experiment 2 hypothesis as H1 or restart numbering? | pass→pass | 9,724 | 19,272 | +98% | 1 | 1 | 0% | 1,800 | 7,269 | +304% | 0 | 0 | — |
▸case-09 We are conducting a survey experiment providing factual information about tariff rates to measure changes in protectionist trade policy preferences. We plan to measure trade policy preferences immediately after treatment. How should we structure the causal estimands for this information experiment? | pass→pass | 17,345 | 26,129 | +51% | 1 | 1 | 0% | 3,074 | 8,237 | +168% | 0 | 0 | — |
▸case-02 We are designing a study where Experiment 1 tests whether a financial literacy message changes retirement savings intentions, and Experiment 2 predicts that adding a low-intensity social proof message produces no additional effect beyond the basic message. Please help us write out the hypothesis architecture for Experiment 2. We need a formal specification that handles the null prediction cleanly, contrasts active versus passive reference groups, defines what evidence would refute our theory versus competing explanations, and details the statistical decision procedure for bounds testing. | fail→fail | 22,013 | 33,023 | +50% | 1 | 1 | 0% | 4,304 | 9,735 | +126% | 0 | 0 | — |
▸case-03 I am running a multi-stage information provision experiment studying how learning about wage inequality impacts tax redistribution preferences. I need a comprehensive hypothesis architecture document to map out the study. Please provide a structured design outline that separates belief updates from policy preference outcomes, addresses how sample-level findings map to broader population claims, details strategies to minimize survey demand artifacts, and maps each theoretical claim directly to its corresponding regression specification. | fail→fail | 31,718 | 32,232 | +2% | 1 | 1 | 0% | 6,222 | 9,720 | +56% | 0 | 0 | — |
▸case-04 I have already formulated our hypotheses, estimands, and theoretical framework for a survey experiment on media bias. Can you draft the complete formatted R markdown file for registering our PAP on EGAP, including all R code chunks for data cleaning and power simulations? | fail→fail | 21,306 | 30,463 | +43% | 1 | 1 | 0% | 4,782 | 9,689 | +103% | 0 | 0 | — |
▸case-05 We have finalized our causal hypothesis and estimands for a study on civic engagement. Can you generate the publication-ready Methods section prose in APSA style conforming to DA-RT transparency standards for our manuscript submission? | fail→fail | 16,280 | 29,979 | +84% | 1 | 1 | 0% | 2,994 | 9,681 | +223% | 0 | 0 | — |
▸case-06 Here is our hypothesis framework for an item count study. Please construct the exact survey questionnaire item text and random allocation Javascript code for Qualtrics to deliver the 4-item vs 5-item lists. | fail→fail | 15,120 | 29,079 | +92% | 1 | 1 | 0% | 2,975 | 9,681 | +225% | 0 | 0 | — |
▸case-07 We are testing whether candidate charisma increases voter persuasion in a vignette experiment. We want to claim that charisma has 'a positive effect' on a 1-7 scale. Can you write a formal hypothesis specification? Note that we plan to use a basic p < 0.05 test without declaring effect thresholds or model specifications. | fail→pass | 8,127 | 26,114 | +221% | 1 | 1 | 0% | 1,597 | 8,054 | +404% | 0 | 0 | — |
▸case-11 We want to test if economic anxiety increases support for protectionist trade policies. Can you write down this hypothesis for our pre-analysis plan? A standard one-sentence statement like 'Economic anxiety increases support for protectionism' should be sufficient. | pass→pass | 2,344 | 25,712 | +997% | 1 | 1 | 0% | 410 | 8,086 | +1872% | 0 | 0 | — |
▸case-12 Our survey experiment examines 15 different subgroup interaction effects alongside our main treatment effect of candidate race on vote choice. Should we classify all 15 subgroup interactions as confirmatory hypotheses in our pre-analysis plan? | fail→pass | 12,622 | 18,697 | +48% | 1 | 1 | 0% | 2,080 | 6,617 | +218% | 0 | 0 | — |
▸case-13 We hypothesize that adding neutral disclaimers to a public health message causes no change in compliance intentions compared to the message without disclaimers. How should we formalize this hypothesis test in our pre-analysis plan? | pass→pass | 13,464 | 26,688 | +98% | 1 | 1 | 0% | 2,587 | 8,431 | +226% | 0 | 0 | — |
▸case-14 We are running a vignette experiment on Mechanical Turk using a convenience sample of 1,000 respondents. We plan to claim that our estimated effect represents the Population Average Treatment Effect (PATE) for all US voters without additional assumptions. How should our hypothesis framework define this estimand? | pass→pass | 13,034 | 18,503 | +42% | 1 | 1 | 0% | 2,359 | 6,855 | +191% | 0 | 0 | — |
▸case-15 We are testing a group-threat theory versus a procedural fairness theory in police-community interactions. Our main hypothesis is that procedural fairness increases cooperation. How should our hypothesis architecture handle the competing group-threat explanation? | pass→pass | 14,387 | 30,927 | +115% | 1 | 1 | 0% | 2,384 | 9,007 | +278% | 0 | 0 | — |
▸case-16 In a survey experiment testing the impact of highlighting foreign aid spending, our treatment text describes $40 billion in aid and mentions that aid helps prevent international conflict. Control gets no text. What key assumption about treatment contrasts must be explicitly stated? | fail→pass | 8,322 | 11,864 | +43% | 1 | 1 | 0% | 1,342 | 5,649 | +321% | 0 | 0 | — |
▸case-23 We are running an experiment on a rare population of 150 local mayors, so our sample size is strictly fixed and underpowered for standard alpha = 0.05. How should we structure our statistical testing logic for this fixed-N study? | pass→pass | 16,718 | 20,666 | +24% | 1 | 1 | 0% | 2,862 | 7,056 | +147% | 0 | 0 | — |