▸case-01 I need a comprehensive methodology blueprint for building a time-series demand forecasting model across multiple regional warehouses. Outline the recommended workflow covering data preprocessing, model selection criteria, backtesting strategies, and model drift alerts. | pass→pass | 31,581 | 27,565 | -13% | 1 | 1 | 0% | 4,561 | 6,558 | +44% | 0 | 0 | — |
▸case-02 We ran an online experiment testing 12 different checkout button colors simultaneously against a control group to find which color maximizes conversion rate. Walk through the statistical procedure for analyzing the conversion metrics across all groups. | fail→fail | 18,523 | 18,759 | +1% | 1 | 1 | 0% | 3,143 | 5,042 | +60% | 0 | 0 | — |
▸case-03 We introduced a premium loyalty program and want to measure whether enrolling in it causes customers to spend more money. We have historical transaction data for customers who opted in versus those who did not, but enrollment was completely voluntary. How should we measure the true causal impact? | pass→pass | 16,244 | 18,152 | +12% | 1 | 1 | 0% | 2,507 | 4,432 | +77% | 0 | 0 | — |
▸case-04 We are building a churn prediction model for a subscription service where only 0.5% of active users churn each month. Outline the data sampling and evaluation strategy to ensure the model reliably catches churners without being overwhelmed by majority class accuracy. | pass→pass | 16,671 | 13,810 | -17% | 1 | 1 | 0% | 2,528 | 4,051 | +60% | 0 | 0 | — |
▸case-05 We trained a complex gradient boosted decision tree model to grant or deny small business loans, but regulators require us to explain the exact drivers behind every individual decision. What techniques should we implement to provide localized feature attribution for each applicant? | pass→pass | 21,121 | 17,476 | -17% | 1 | 1 | 0% | 3,064 | 4,838 | +58% | 0 | 0 | — |
▸case-06 Our marketing team wants to evaluate the ROI of linear TV ads, print billboards, and podcasts alongside digital search ads. User tracking cookies cannot track offline exposures. Design an analytical framework to estimate channel effectiveness and optimize overall budget allocation. | pass→pass | 22,174 | 23,235 | +5% | 1 | 1 | 0% | 3,484 | 5,536 | +59% | 0 | 0 | — |
▸case-07 We plan to test a new homepage design. Baseline conversion rate is 5% and we want to detect a minimum relative lift of 10% with 80% power at a 0.05 significance level. Explain how to determine sample size and the risks of checking p-values daily. | pass→pass | 16,802 | 17,339 | +3% | 1 | 1 | 0% | 3,104 | 4,889 | +58% | 0 | 0 | — |
▸case-08 We are evaluating a daily retail sales forecasting model using historical data from 2020 to 2023. Detail the cross-validation strategy we should use during hyperparameter tuning. | pass→pass | 18,896 | 22,120 | +17% | 1 | 1 | 0% | 3,059 | 5,958 | +95% | 0 | 0 | — |
▸case-09 Our dataset contains a high-cardinality categorical variable (ZIP code with over 15,000 unique values) that we want to feed into a gradient boosting model. What encoding techniques should we use instead of standard one-hot encoding? | pass→pass | 13,607 | 16,161 | +19% | 1 | 1 | 0% | 2,203 | 4,145 | +88% | 0 | 0 | — |
▸case-10 Our credit scoring model has been deployed in production for 6 months. What monitoring infrastructure and statistical metrics should we implement to detect when input data patterns shift away from training distributions before prediction quality drops? | pass→pass | 24,611 | 30,122 | +22% | 1 | 1 | 0% | 3,728 | 5,243 | +41% | 0 | 0 | — |
▸case-11 We want to analyze customer tenure for a SaaS product where many current subscribers have not yet churned (right-censored data). We want to estimate the probability of retention over time across different cohorts. What statistical modeling approach is best suited? | pass→pass | 10,719 | 14,350 | +34% | 1 | 1 | 0% | 1,915 | 4,216 | +120% | 0 | 0 | — |
▸case-12 In a checkout flow experiment, our key metric is Average Revenue Per User (ARPU), which is calculated as total revenue divided by total visitors. The metric exhibits strong right-skewness and zero-inflation. How should we perform hypothesis testing on this ratio metric? | pass→pass | 19,240 | 17,677 | -8% | 1 | 1 | 0% | 3,150 | 4,634 | +47% | 0 | 0 | — |
▸case-13 We are running a discount email campaign with limited budget. We want to target only 'persuadable' customers who buy ONLY IF they receive a discount, avoiding 'sure things' (who buy anyway) and 'sleeping dogs' (who get annoyed). What modeling strategy targets this specific incremental effect? | pass→pass | 14,216 | 13,081 | -8% | 1 | 1 | 0% | 2,375 | 3,989 | +68% | 0 | 0 | — |
▸case-14 We need to dynamically allocate web traffic between 4 headline variants for a breaking news article that will only be relevant for 48 hours. We want to maximize total clicks while minimizing traffic wasted on underperforming variants during the test. What approach should we use? | pass→pass | 13,936 | 17,394 | +25% | 1 | 1 | 0% | 2,029 | 4,585 | +126% | 0 | 0 | — |
▸case-15 We have transactional e-commerce data containing purchase timestamps, order values, and product categories for 500,000 customers. We want to group customers into actionable business segments. Outline the feature engineering and clustering methodology. | pass→pass | 20,056 | 16,958 | -15% | 1 | 1 | 0% | 3,104 | 4,430 | +43% | 0 | 0 | — |
▸case-16 We are building a collaborative filtering recommendation engine for a video streaming site. How should the system generate recommendations for brand-new users who have no viewing history? | pass→pass | 16,433 | 14,963 | -9% | 1 | 1 | 0% | 2,384 | 3,915 | +64% | 0 | 0 | — |
▸case-17 We have high-dimensional gene expression data (20,000 features per sample) and want to visualize the cluster structures in 2D to spot non-linear relationships across cell types. What visualization/reduction technique is best suited? | pass→pass | 12,859 | 12,552 | -2% | 1 | 1 | 0% | 2,184 | 3,947 | +81% | 0 | 0 | — |
▸case-18 We launched a regional marketing campaign exclusively in the state of Ohio. No random control group exists, and other states have different baseline economic conditions. How can we construct a rigorous statistical baseline to estimate the net impact of the campaign? | pass→pass | 21,756 | 15,012 | -31% | 1 | 1 | 0% | 2,234 | 4,058 | +82% | 0 | 0 | — |
▸case-19 We are tuning a gradient boosted tree model with 12 continuous and discrete hyperparameters. Grid search is computationally prohibitive. What efficient, sample-effective optimization methodology should we implement? | pass→pass | 14,531 | 16,720 | +15% | 1 | 1 | 0% | 2,373 | 4,689 | +98% | 0 | 0 | — |
▸case-20 We need to flag fraudulent credit card transactions in real-time. Unlabeled transaction logs contain high volume, high dimensional streaming data where anomalies are rare (< 0.1%). What unsupervised algorithm suite should we deploy? | pass→pass | 18,878 | 20,004 | +6% | 1 | 1 | 0% | 2,863 | 4,913 | +72% | 0 | 0 | — |
▸case-21 We are using Difference-in-Differences (DiD) to evaluate the impact of a state-level policy change on employment. What core statistical assumption must be formally tested before trusting the DiD estimate, and how is it evaluated? | pass→pass | 12,654 | 9,742 | -23% | 1 | 1 | 0% | 2,140 | 3,350 | +57% | 0 | 0 | — |
▸case-22 Write an ANSI SQL query to calculate a 7-day moving average of daily sales from a table named `daily_sales` containing `sale_date` and `amount` columns. | pass→pass | 11,828 | 7,495 | -37% | 1 | 1 | 0% | 1,026 | 2,965 | +189% | 0 | 0 | — |
▸case-23 Explain the mathematical relationship between variance and standard deviation for a sample dataset, including the exact formula connecting them. | pass→pass | 8,896 | 8,950 | +1% | 1 | 1 | 0% | 1,619 | 3,606 | +123% | 0 | 0 | — |
▸case-24 Provide a Python code snippet using scikit-learn that sets up a Pipeline containing a StandardScaler and LogisticRegression, then fits it on `X_train` and `y_train`. | pass→pass | 3,452 | 3,347 | -3% | 1 | 1 | 0% | 686 | 2,332 | +240% | 0 | 0 | — |