▸case-01 I'm launching an automated reporting module for our B2B SaaS platform and need to triage these product assumptions:
1. Users will pay $50/mo for scheduled PDF exports.
2. Engineering can build custom template rendering in under two weeks.
3. Enterprise clients prefer email delivery over Slack webhooks.
4. Adding automated reports will reduce monthly churn by 5%.
Please evaluate each item on impact and risk, assign a clear recommended action for each item, outline behavioral experiment ideas for any that need testing, and present the final results in a structured markdown table. | fail→pass | 19,971 | 15,281 | -23% | 1 | 1 | 0% | 2,938 | 3,307 | +13% | 0 | 0 | — |
▸case-02 We're redesigning our fintech mobile app onboarding flow and compiled 5 core user hypotheses:
1. Users will link their primary bank account during registration if offered a $5 bonus.
2. Simplifying the KYC form to two screens will boost conversion by 15%.
3. Users trust biometric login more than SMS OTPs.
4. First-time users prefer a video tutorial over an interactive walkthrough.
5. Showing real-time portfolio previews increases completion rates.
Can you assess these across impact and risk dimensions, categorize the proper next step for each, specify low-effort experiment designs with concrete success metrics where appropriate, and organize the output as a prioritized summary matrix? | fail→pass | 24,323 | 20,615 | -15% | 1 | 1 | 0% | 3,394 | 4,262 | +26% | 0 | 0 | — |
▸case-03 Our e-commerce team is planning to introduce a personalized AI recommendation widget at checkout. Here are the main assumptions we have listed:
1. Shoppers will add impulse items under $10 to their cart right before paying.
2. The recommendation engine API latency will remain under 200ms at peak volume.
3. Showing recommended items won't increase cart abandonment rates.
4. Free shipping thresholds drive higher average order value than item discounts.
Please run these through an impact-versus-risk evaluation framework, determine the appropriate action category for each, draft targeted behavioral tests for those needing validation, and present the final breakdown in a clean markdown table. | fail→pass | 15,274 | 16,945 | +11% | 1 | 1 | 0% | 2,939 | 3,548 | +21% | 0 | 0 | — |
▸case-04 We are triaging hypotheses for the Acme Analytics dashboard redesign. We classified our four core hypotheses into the four quadrants of our impact and risk matrix. Many product managers default to running A/B tests on every single quadrant or labeling low-impact items as 'low priority A/B test'. Please confirm the exact recommended action for each of the 4 matrix quadrants so our team knows when to build, test, defer, or discard. | pass→pass | 11,076 | 13,805 | +25% | 1 | 1 | 0% | 2,011 | 2,886 | +44% | 0 | 0 | — |
▸case-05 We are preparing a launch strategy for CloudSync Desktop v3. Our PM suggested labeling our risky assumptions as 'User Interview Needed' or 'Design Prototype'. We have 3 assumptions: 1. Users want automatic background syncing (High Impact, Low Risk). 2. Users will pay $15/yr for LAN sync (Low Impact, High Risk). 3. Syncing zero-byte files will corrupt local caches (High Impact, High Risk). Please classify these using standard impact-risk matrix actions and present them in a structured table. | fail→pass | 9,507 | 13,172 | +39% | 1 | 1 | 0% | 1,805 | 2,417 | +34% | 0 | 0 | — |
▸case-06 For the PayPulse POS tablet app, we need to evaluate whether merchants will adopt tap-to-pay on Android. Many teams present assumption triage as a loose narrative paragraph. Please process our technical and business assumptions using step-by-step risk-impact analysis and save the primary output structure in clean markdown format. | fail→pass | 16,889 | 22,757 | +35% | 1 | 1 | 0% | 2,969 | 4,527 | +52% | 0 | 0 | — |
▸case-07 We are prioritizing feature ideas for the HealthTrack wearable app. A team member suggested calculating Impact simply as subjective rating 1 to 10 without considering market size or customer dissatisfaction. Using Dan Olsen's opportunity scoring principles within ICE prioritization, how should Impact be derived when evaluating our user assumptions? | fail→pass | 14,303 | 15,685 | +10% | 1 | 1 | 0% | 2,554 | 3,376 | +32% | 0 | 0 | — |
▸case-08 In our SaaS platform dev team, engineers want Risk measured purely by story points, while product managers want it measured purely by lack of data. How should Risk be mathematically defined in an assumption triage framework combining confidence and effort? | pass→pass | 18,184 | 15,039 | -17% | 1 | 1 | 0% | 3,178 | 3,259 | +3% | 0 | 0 | — |
▸case-09 We are testing whether enterprise HR managers will use AI-generated job descriptions in Workforce Suite. The lead PM suggested conducting 10 user opinion surveys to validate this high-impact, high-risk assumption. What key design principles should govern the experiment suggested for validating this assumption? | pass→pass | 18,561 | 17,854 | -4% | 1 | 1 | 0% | 2,346 | 2,844 | +21% | 0 | 0 | — |
▸case-10 We are designing a concierge experiment for AutoFleet logistics software to test if dispatchers adopt automated route re-clustering. How should the success criterion for this experiment be defined? | pass→pass | 13,329 | 14,500 | +9% | 1 | 1 | 0% | 2,175 | 2,321 | +7% | 0 | 0 | — |
▸case-11 During our review of the ShopFast checkout flow, we identified an assumption: 'Changing the button radius from 4px to 6px will look slightly cleaner.' We determined this has low impact and low risk. Teammates suggest setting up a 2-week A/B test. What is the correct action for this item? | fail→pass | 8,242 | 10,057 | +22% | 1 | 1 | 0% | 1,354 | 1,766 | +30% | 0 | 0 | — |
▸case-12 For StreamBox video app, adding offline download capabilities for paid tier users has been verified as high impact and low risk. Engineers want to design a 3-stage fake-door test before writing code. What action should be taken? | pass→pass | 7,250 | 7,647 | +5% | 1 | 1 | 0% | 1,248 | 1,749 | +40% | 0 | 0 | — |
▸case-13 For the SecureVault password manager, supporting legacy Windows Phone 8 sync requires 3 months of refactoring (high effort/risk) but affects less than 0.01% of non-paying users (low impact). The team wants to run a smoke test. What action should be assigned? | fail→pass | 6,697 | 7,152 | +7% | 1 | 1 | 0% | 1,235 | 1,689 | +37% | 0 | 0 | — |
▸case-14 In FleetTracker GPS, our core hypothesis is that long-haul drivers will accept real-time driver fatigue monitoring via mobile camera. This affects all enterprise accounts (high impact) but drivers may reject camera surveillance (low confidence, high risk). Product managers want to immediately start full engineering implementation. What action must be assigned? | fail→pass | 7,866 | 8,615 | +10% | 1 | 1 | 0% | 1,320 | 1,977 | +50% | 0 | 0 | — |
▸case-15 We are launching explicit tiering in DataPulse BI platform. Evaluate these 2 assumptions: 1. Small business users will upgrade to Pro if exports are capped at 1,000 rows (High Impact, High Risk). 2. Enterprise admins want SSO SAML support (High Impact, Low Risk). The team usually lists actions as 'TBD' or 'Discuss in Sprint Planning'. What are the matrix-governed actions? | pass→pass | 14,097 | 11,656 | -17% | 1 | 1 | 0% | 2,107 | 2,458 | +17% | 0 | 0 | — |
▸case-16 For the FitTrack mobile app, we suspect offering 7-day trials will double trial-to-paid conversion (High Impact, High Risk). A junior PM suggested running a 50-person focus group to test this. How should this assumption be classified and tested under formal triage principles? | fail→pass | 18,051 | 13,645 | -24% | 1 | 1 | 0% | 2,266 | 2,785 | +23% | 0 | 0 | — |
▸case-17 At QuickCart Online, we are evaluating: 1. Adding Apple Pay will increase checkout completion by 8% (High Impact, Low Risk). 2. Changing footer copyright text to include 2025 will boost trust (Low Impact, Low Risk). What specific action categories apply to each? | fail→pass | 7,164 | 4,682 | -35% | 1 | 1 | 0% | 1,085 | 1,354 | +25% | 0 | 0 | — |
▸case-18 In BankSwift Mobile, we assume building a custom proprietary facial recognition engine for KYC will reduce fraud by 1% (Low Impact, High Risk due to 6-month dev time). Developers want to build a prototype. What is the triage framework's recommendation? | pass→pass | 9,552 | 8,591 | -10% | 1 | 1 | 0% | 1,653 | 1,944 | +18% | 0 | 0 | — |
▸case-19 For TaskHero service marketplace, we assume gig workers will accept 5% lower commission if paid daily instead of weekly (High Impact, High Risk). How should we structure the triage and follow-up test? | fail→pass | 17,131 | 19,720 | +15% | 1 | 1 | 0% | 2,772 | 3,124 | +13% | 0 | 0 | — |
▸case-20 Please write a Product Requirements Document (PRD) for adding Dark Mode to our Desktop Notes app. Include executive summary, user stories, functional requirements, and technical constraints. | pass→fail | 16,982 | 23,560 | +39% | 1 | 1 | 0% | 3,119 | 4,744 | +52% | 0 | 0 | — |
▸case-21 We are organizing a user story mapping workshop for a new ride-sharing app's passenger journey (Account Creation -> Ride Search -> Booking -> Payment -> Driver Rating). Please list user stories grouped by journey backbone activities. | pass→fail | 13,950 | 22,970 | +65% | 1 | 1 | 0% | 2,588 | 5,020 | +94% | 0 | 0 | — |
▸case-22 Draft Q3 Objectives and Key Results (OKRs) for our Customer Support team aiming to improve customer satisfaction and reduce response time for ticket escalations. | pass→pass | 9,865 | 21,106 | +114% | 1 | 1 | 0% | 1,855 | 3,333 | +80% | 0 | 0 | — |
▸case-23 Calculate the 12-month Customer Lifetime Value (CLV) for a SaaS business with $100 ARPU, 2% monthly churn, and 80% gross margin using standard SaaS financial formulas. | pass→pass | 14,001 | 17,401 | +24% | 1 | 1 | 0% | 2,282 | 3,388 | +48% | 0 | 0 | — |