▸case-01 Our SaaS platform is considering introducing an interactive in-app setup checklist to boost trial-to-paid conversion. Before building out the full backend workflow, we need to test whether users will actually engage with it and if it reduces drop-off. Please analyze this feature concept and break down a series of quick validation tests for our underlying assumptions. Present the plan detailing each assumption, proposed experiment, success metric, and target threshold. | fail→pass | 15,967 | 14,902 | -7% | 1 | 1 | 0% | 2,800 | 3,107 | +11% | 0 | 0 | — |
▸case-02 We want to add an AI summary generator to our e-commerce merchant dashboard. The engineering lead wants to spend 6 weeks building LLM integration, prompt pipelines, and background queues immediately. We want to test if merchants actually want this button before building the infrastructure. Create an experiment plan. We are tempted to just run an email survey asking merchants if they would like AI summaries. | pass→pass | 17,721 | 15,133 | -15% | 1 | 1 | 0% | 2,669 | 2,911 | +9% | 0 | 0 | — |
▸case-03 Our checkout team wants to test a new single-page checkout flow on our live e-commerce store to see if it increases conversion rate. The product manager proposes releasing the new design to 50% of all holiday traffic immediately without safety guardrails. Outline an experiment plan for this checkout change. | pass→pass | 18,959 | 14,068 | -26% | 1 | 1 | 0% | 2,775 | 2,925 | +5% | 0 | 0 | — |
▸case-04 We are thinking about sending daily AI-curated news summaries via push notifications in our mobile app. The team proposed running an in-app poll asking users 'Would you read a daily AI news digest? Yes/No'. Design a better validation plan for this proposal. | pass→pass | 13,580 | 14,273 | +5% | 1 | 1 | 0% | 2,266 | 2,980 | +32% | 0 | 0 | — |
▸case-05 A customer support software company wants to test an automated refund processing agent. The CTO wants to build a full automated banking API integration first. How can we test whether customers trust and use automated refunds with minimal engineering effort? | pass→pass | 17,667 | 14,942 | -15% | 1 | 1 | 0% | 2,295 | 2,957 | +29% | 0 | 0 | — |
▸case-06 Our design tool team wants to add real-time multiplayer cursor collaboration. We aren't sure if WebSockets can handle 100 concurrent users per document without latency lag, nor if users need multiplayer on small documents. Create an experiment design covering both concerns. | fail→pass | 20,215 | 17,218 | -15% | 1 | 1 | 0% | 3,467 | 3,087 | -11% | 0 | 0 | — |
▸case-07 We are redesigning the main navigation menu of a complex B2B analytics platform. The team wants to launch the new navigation directly to production and measure customer support ticket volume over 3 months. Provide a lower-effort, lower-risk experiment design. | fail→pass | 13,378 | 15,326 | +15% | 1 | 1 | 0% | 2,233 | 2,631 | +18% | 0 | 0 | — |
▸case-08 We want to add receipt optical character recognition (OCR) scanning to our expense management software. The PM wants to survey users with the question: 'How much do you like OCR features on a scale of 1-5?'. Redesign this validation effort to ensure actionable learning. | pass→pass | 13,025 | 17,430 | +34% | 1 | 1 | 0% | 2,258 | 3,336 | +48% | 0 | 0 | — |
▸case-09 Our SaaS company wants to add a $499/month Enterprise plan with custom role-based access control (RBAC). Sales wants to write full code for custom roles before announcing the tier. How can we validate enterprise demand for this tier with minimal effort? | fail→pass | 19,515 | 17,372 | -11% | 1 | 1 | 0% | 2,384 | 3,520 | +48% | 0 | 0 | — |
▸case-10 We want to validate adding advanced genre and mood filters to a streaming video web app. Generate a structured experiment proposal for this feature. We want to make sure the response covers all required experiment definition components. | fail→pass | 16,447 | 16,868 | +3% | 1 | 1 | 0% | 3,173 | 3,393 | +7% | 0 | 0 | — |
▸case-11 We plan to deploy an automated upsell recommendation widget on high-volume B2B client invoicing pages. We want to run an A/B test on live traffic, but engineering is worried about breaking invoice delivery or slowing down page load times. Provide an experiment plan. | fail→pass | 18,597 | 18,305 | -2% | 1 | 1 | 0% | 2,771 | 3,687 | +33% | 0 | 0 | — |
▸case-12 A desktop file manager product wants to introduce semantic AI search. The engineering team estimates 3 months to build vector embeddings and local indexers. How can we test if users actually prefer semantic search over keyword search using minimal effort? | pass→pass | 21,730 | 16,444 | -24% | 1 | 1 | 0% | 2,535 | 3,174 | +25% | 0 | 0 | — |
▸case-13 Our field service mobile app needs to support offline sync for remote technicians. We are unsure if SQLite local storage can sync large media files reliably over spotty 3G networks. Design an experiment plan to validate this technical assumption before full application refactoring. | pass→pass | 22,909 | 18,045 | -21% | 1 | 1 | 0% | 4,089 | 3,429 | -16% | 0 | 0 | — |
▸case-14 A fitness app team believes users are dropping off during onboarding because the biometric goal selection screen is confusing. The team wants to re-code the screen in Flutter and push an update to the App Store. Suggest a faster validation experiment. | fail→pass | 11,735 | 14,119 | +20% | 1 | 1 | 0% | 1,910 | 2,345 | +23% | 0 | 0 | — |
▸case-15 A fintech app wants to introduce an AI financial advisor that creates custom monthly budget plans. Building the underlying financial recommendation algorithm will take 4 months. Design a low-effort experiment to test if users value customized budget plans. | pass→pass | 14,807 | 12,040 | -19% | 1 | 1 | 0% | 2,391 | 2,530 | +6% | 0 | 0 | — |
▸case-16 An e-commerce team wants to run an A/B test on a new search relevance algorithm. If the new algorithm performs poorly, customer purchases could drop drastically. How should the experiment be structured to protect business revenue? | fail→pass | 16,747 | 12,875 | -23% | 1 | 1 | 0% | 2,113 | 2,764 | +31% | 0 | 0 | — |
▸case-17 A CRM platform is planning to build a 2-way sync integration with Salesforce. The development effort is estimated at 800 engineering hours. How can the product team validate customer demand for this integration before writing integration code? | pass→pass | 21,220 | 19,800 | -7% | 1 | 1 | 0% | 2,623 | 3,891 | +48% | 0 | 0 | — |
▸case-18 A developer desktop tool team wants to know if users want a dark mode UI. The PM created a survey asking 'Do you like dark mode?'. How should this assumption be tested effectively? | fail→pass | 11,451 | 16,526 | +44% | 1 | 1 | 0% | 1,906 | 2,670 | +40% | 0 | 0 | — |
▸case-19 An online grocery service wants to test a 'Reorder Last Week's Items in One Click' button on the homepage. Create a complete validation plan for this assumption. | fail→fail | 17,364 | 15,513 | -11% | 1 | 1 | 0% | 3,009 | 2,647 | -12% | 0 | 0 | — |
▸case-20 Our B2B SaaS platform wants to introduce an automated contract renewal reminder system for account managers. Before spending 2 months building automated email triggers and CRM syncing, we need to test if account managers will utilize these automated reminders. Detail an experiment plan containing assumptions, experiment designs, behavioral metrics, risk controls, and target success thresholds. | fail→fail | 19,818 | 16,447 | -17% | 1 | 1 | 0% | 2,797 | 2,880 | +3% | 0 | 0 | — |
▸case-21 Write a complete Product Requirements Document (PRD) for a user authentication system using OAuth 2.0 and JWTs, including DB schemas, API endpoint specifications, and error handling codes. | pass→pass | 28,234 | 25,541 | -10% | 1 | 1 | 0% | 5,476 | 4,875 | -11% | 0 | 0 | — |
▸case-22 Here is a list of 5 feature requests from existing customers: 1. Export to PDF, 2. Dark mode, 3. SSO integration, 4. Audit logs, 5. Bulk delete. Apply the RICE (Reach, Impact, Confidence, Effort) prioritization framework to rank these features. | pass→pass | 15,077 | 24,353 | +62% | 1 | 1 | 0% | 3,210 | 4,084 | +27% | 0 | 0 | — |
▸case-23 We launched an automated invoice reminder feature 6 months ago. Here is our monthly churn and usage retention data: Month 1 (80% active), Month 2 (65% active), Month 3 (50% active). Analyze this retention curve, compute the 90-day retention rate, and recommend retention optimization strategies based on this data. | pass→pass | 18,085 | 13,985 | -23% | 1 | 1 | 0% | 2,543 | 2,791 | +10% | 0 | 0 | — |