▸case-17 An experiment testing $5 off vs 10% off coupons produced redemption counts in `coupons_exp.csv`. Can you run a standard Fisher's exact test to check significance? | fail→fail | 17,965 | 15,004 | -16% | 1 | 1 | 0% | 496 | 329 | -34% | 0 | 0 | — |
▸case-01 We just wrapped up our checkout flow experiment, and the raw metrics for the control and treatment groups are saved in `checkout_data.csv`. Could you perform the standard pre-registered statistical testing on these experimental results and let me know if the difference is meaningful? | fail→fail | 21,068 | 15,577 | -26% | 1 | 1 | 0% | 1,297 | 535 | -59% | 0 | 0 | — |
▸case-02 I need to evaluate our new search ranking algorithm against the baseline. I've attached the click-through rate measurements across our test splits. Please run the required statistical testing procedures on these experiment metrics to see if the improvement is statistically sound. | fail→fail | 8,927 | 17,055 | +91% | 1 | 1 | 0% | 636 | 439 | -31% | 0 | 0 | — |
▸case-03 Here are the metric outputs from our mobile app redesign trial comparing average session duration between user cohorts. Please execute our statistical testing process on these results and give me a full summary of the statistical evaluation. | fail→fail | 12,113 | 16,895 | +39% | 1 | 1 | 0% | 1,150 | 630 | -45% | 0 | 0 | — |
▸case-04 We tested a new collaborative filtering recommender against matrix factorization in an A/B test. We want to check if the click-through rates are significantly different. Would a two-sample Z-test or chi-square test be appropriate here, or how should we evaluate these pre-registered experiment results? | fail→fail | 21,350 | 37,211 | +74% | 1 | 1 | 0% | 2,723 | 6,837 | +151% | 0 | 0 | — |
▸case-10 We measured step-1 to step-3 completion rates in our new user onboarding flow across splits. I'm going to run `scipy.stats.ttest_ind` on the completion binary flags. Please execute our statistical evaluation process. | fail→fail | 19,688 | 30,676 | +56% | 1 | 1 | 0% | 4,099 | 5,056 | +23% | 0 | 0 | — |
▸case-05 We completed a landing page trial (`landing_page_exp.parquet`) comparing Variant A and Variant B. Normally we calculate standard Student's t-test p-values. Please run our pre-registered testing workflow to evaluate conversion rate lift. | fail→fail | 29,051 | 10,092 | -65% | 1 | 1 | 0% | 4,677 | 273 | -94% | 0 | 0 | — |
▸case-06 Our pricing page test measured Average Revenue Per User (ARPU) across Control and Variant C. Revenue data is heavily right-skewed with many zeroes. Should we run Welch's t-test or Mann-Whitney U test on `arpu_results.csv` to execute our pre-registered testing? | fail→fail | 14,889 | 15,579 | +5% | 1 | 1 | 0% | 2,539 | 291 | -89% | 0 | 0 | — |
▸case-07 We deployed an API caching optimization and measured 95th percentile latency (p95) across test groups (`latency_metrics.json`). I was planning to run a standard t-test on the log-transformed latencies. Please execute our statistical test suite. | fail→fail | 14,448 | 16,009 | +11% | 1 | 1 | 0% | 185 | 297 | +61% | 0 | 0 | — |
▸case-08 We ran a subject line test on 50,000 subscribers (`email_exp_logs.csv`). Can you compute the standard chi-square test of independence and p-value to determine if Option B beat Option A? | fail→fail | 12,813 | 10,321 | -19% | 1 | 1 | 0% | 2,242 | 259 | -88% | 0 | 0 | — |
▸case-09 We evaluated a personalized push notification campaign versus control (`push_results.parquet`). I'm preparing to calculate the Z-score and standard two-proportion confidence intervals. Please run our pre-registered experiment test procedure. | fail→fail | 16,866 | 9,950 | -41% | 1 | 1 | 0% | 839 | 247 | -71% | 0 | 0 | — |
▸case-11 We want to evaluate conversion rate changes in `ad_campaign_eval.csv` and determine if the effect size falls within a Region of Practical Equivalence. Execute our pre-registered testing workflow on these metrics. | fail→fail | 10,591 | 5,356 | -49% | 1 | 1 | 0% | 146 | 194 | +33% | 0 | 0 | — |
▸case-12 7-day retention data for our dark mode feature flag experiment is available in `retention_data.csv`. Should we use a standard two-sample t-test to evaluate retention, or what is our established process? | fail→fail | 19,758 | 9,148 | -54% | 1 | 1 | 0% | 2,551 | 366 | -86% | 0 | 0 | — |
▸case-13 We tested a new payment gateway router and collected transaction failure events (`gateway_errors.parquet`). The baseline error rate is 0.5%. Can you run a standard Z-test for proportions to test if the new gateway significantly increased failures? | fail→fail | 18,140 | 8,601 | -53% | 1 | 1 | 0% | 3,440 | 726 | -79% | 0 | 0 | — |
▸case-14 We evaluated normalized discounted cumulative gain (NDCG@10) for two search ranking models across 5,000 queries (`ndcg_eval.csv`). I was thinking of running an ANOVA test. Please run our pre-registered statistical test on these results. | fail→fail | 16,086 | 16,020 | -0% | 1 | 1 | 0% | 3,266 | 309 | -91% | 0 | 0 | — |
▸case-15 We collected 5-star product ratings after updating the review submission UI (`ratings_ab.csv`). Please conduct our pre-registered statistical analysis to see if treatment ratings differ from control. | fail→fail | 16,525 | 10,718 | -35% | 1 | 1 | 0% | 2,326 | 432 | -81% | 0 | 0 | — |
▸case-16 Page load metrics for control and treatment variants in our CDN experiment are saved in `load_times.csv`. I'm about to run Welch's t-test on mean load times. Please execute our statistical testing protocol. | fail→fail | 24,492 | 11,413 | -53% | 1 | 1 | 0% | 4,743 | 421 | -91% | 0 | 0 | — |
▸case-18 We measured 30-day cancellation rates in `churn_trial_metrics.csv` for two billing cadence cohorts. Please run our official pre-registered statistical testing process on these results. | fail→fail | 8,248 | 10,333 | +25% | 1 | 1 | 0% | 1,797 | 262 | -85% | 0 | 0 | — |
▸case-19 We are planning a two-sample A/B test for checkout conversion rate ($p_1 = 0.10$, $p_2 = 0.12$) with power $0.80$ and $\alpha = 0.05$. What is the formula for Cohen's $h$ arcsine transformation used to calculate the effect size for statistical power analysis? | pass→pass | 12,946 | 10,482 | -19% | 1 | 1 | 0% | 1,644 | 2,252 | +37% | 0 | 0 | — |
▸case-20 During an A/B experiment, we need to verify that user assignment between Control and Variant matches the expected 50/50 split. Which statistical hypothesis test function in `scipy.stats` should be run on the observed sample counts to detect Sample Ratio Mismatch (SRM)? | pass→pass | 12,612 | 11,000 | -13% | 1 | 1 | 0% | 1,464 | 1,046 | -29% | 0 | 0 | — |
▸case-21 We have a raw event table `user_events_raw` with columns `event_id`, `user_id`, and `event_timestamp`. Write a SQL query using a window function to deduplicate records by keeping only the row with the minimum `event_timestamp` for each `event_id`. | pass→pass | 11,086 | 5,736 | -48% | 1 | 1 | 0% | 1,140 | 942 | -17% | 0 | 0 | — |
▸case-22 We are applying K-Means clustering to `user_demographics.csv` to segment users based on purchase behavior. Which function in `sklearn.metrics` calculates the Silhouette Coefficient to evaluate cluster separation across different values of $k$? | pass→pass | 8,711 | 8,616 | -1% | 1 | 1 | 0% | 692 | 630 | -9% | 0 | 0 | — |