▸case-03 Can you run a formal evidence evaluation on our claim that 'enterprise buyers require an AI-assisted reporting module to close deals'? The decision riding on this is prioritizing it for Q3 planning. Inputs: 2 lost deal notes citing missing AI features, 1 major client pitch where the buyer praised the prototype, an industry analyst report highlighting AI trends in B2B SaaS, and usage logs showing only 3% of current active users open our legacy reporting tab. Format your analysis with an inventory catalog including data quality, a summary paragraph of net strength, a sufficiency determination tied to the decision's risk level, and a cost-effective next step to bolster our data if needed. | fail→fail | 16,158 | 16,395 | +1% | 1 | 1 | 0% | 2,530 | 3,639 | +44% | 0 | 0 | — |
▸case-04 We are deciding whether to swap our ecommerce catalog search engine with a semantic vector search vendor. The contract costs $80,000 annually. Claim: 'Semantic search increases checkout conversion rate by at least 0.5%.' We ran a 14-day randomized split-traffic A/B test on 50,000 visitors showing a +0.7% conversion lift with p=0.02, but 3 senior engineers prefer our existing lexical index because of lower query latency. Grade this evidence pile and tell us if it meets sufficiency for signing the contract. | fail→pass | 16,334 | 15,248 | -7% | 1 | 1 | 0% | 2,708 | 3,667 | +35% | 0 | 0 | — |
▸case-01 I need an evidence audit for our proposed checkout redesign, which will cost around $150k and take 3 months to build. Our claim is that 'users want one-click guest checkout.' We have notes from 6 customer interviews, 15 support tickets requesting it, a feedback survey filled out by 200 users (though questions were worded somewhat leadingly), and internal churn metrics showing 4% drop-off at account creation. Please give me an itemized inventory breakdown with source quality notes and direction, a clear assessment of what this net evidence actually proves, a verdict on whether this is sufficient given our high cost and low reversibility, and the cheapest practical follow-up test if we need stronger proof. | fail→fail | 14,349 | 14,048 | -2% | 1 | 1 | 0% | 2,526 | 3,376 | +34% | 0 | 0 | — |
▸case-02 Please evaluate the evidence behind our plan to shift from a seat-based pricing model to usage-based pricing. Here is what we have: a competitor blog post claiming usage pricing increased their ARR, 4 sales calls where prospects complained about per-seat costs, intuition from our VP of Sales, and a recent churn analysis showing 12 clients canceled due to budget cuts. I need an inventory table cataloging all pieces of data (including opposing evidence), a honest synthesis of the net evidence weight, a decision verdict based on our financial stakes, and a ranked list of low-cost ways to upgrade our confidence if the current data isn't enough. | fail→pass | 16,999 | 17,167 | +1% | 1 | 1 | 0% | 2,618 | 3,814 | +46% | 0 | 0 | — |
▸case-05 Our iOS team wants to delay the Q2 release by two weeks to implement Dark Mode. Claim: 'Dark mode increases 30-day mobile retention.' Evidence: 45 tweets asking for dark mode, 3 App Store reviews complaining about bright backgrounds, and a tech blog post arguing dark mode reduces eye strain. Assess whether this evidence justifies a two-week ship delay. | fail→pass | 11,237 | 13,229 | +18% | 1 | 1 | 0% | 1,662 | 3,240 | +95% | 0 | 0 | — |
▸case-06 We want to reduce our annual subscription upfront discount from 20% to 10% across all SMB self-serve signups. Claim: 'Lowering the annual discount will not increase self-serve signup drop-off.' Evidence: 5 customer success managers saying buyers don't care about the discount percentage, and 1 enterprise customer who specifically demanded 20%. Evaluate the net strength and propose the cheapest experiment. | fail→pass | 15,199 | 14,221 | -6% | 1 | 1 | 0% | 2,297 | 3,426 | +49% | 0 | 0 | — |
▸case-07 We plan to restrict export capabilities on our free plan to 3 exports per month. Claim: 'Restricting free exports forces power users to upgrade to paid plans without increasing churn.' Evidence: Product analytics showing 12% of free users perform >3 exports/month, 8 customer service tickets threatening cancellation if exports are capped, and a survey of 100 free users where 65% said export limits would cause them to leave. Grade this evidence and assess sufficiency. | pass→pass | 12,395 | 11,853 | -4% | 1 | 1 | 0% | 2,136 | 2,950 | +38% | 0 | 0 | — |
▸case-08 Our developer relations team proposes rebuilding our developer documentation portal on a new headless platform ($40k spend). Claim: 'Developers abandon integration due to poor API docs.' Evidence: 12 forum posts complaining about broken code samples, 1 key partner who successfully integrated after receiving dedicated Slack support, and internal web analytics showing high bounce rates on the API reference page. Grade the evidence inventory. | fail→pass | 12,646 | 13,626 | +8% | 1 | 1 | 0% | 1,944 | 3,250 | +67% | 0 | 0 | — |
▸case-09 Leadership wants to switch account executive commission from total contract value to multi-year upfront collected revenue. Claim: 'TCV commission leads to high multi-year churn.' Evidence: VP of Finance spreadsheet model, 3 anecdotal case studies of customers who defaulted in year 2, and 1 sales rep who quit citing quota pressure. Evaluate this data for making a permanent compensation structure change. | fail→pass | 14,472 | 15,833 | +9% | 1 | 1 | 0% | 2,346 | 3,460 | +47% | 0 | 0 | — |
▸case-10 Our product team wants to completely remove the 5-step modal product walkthrough for new signups. Claim: 'Modal walkthroughs degrade user activation rates.' Evidence: Mixed results from a 500-user split test showing +1% activation with no statistical significance, 20 interview notes where users said they clicked 'Skip' immediately, and a UX design agency recommendation paper. Grade the pile. | pass→pass | 13,064 | 12,262 | -6% | 1 | 1 | 0% | 1,894 | 2,966 | +57% | 0 | 0 | — |
▸case-11 We are considering spending $100k to fully translate and localize our software for the Japanese market. Claim: 'Japanese mid-market firms will adopt our platform if localized.' Evidence: 3 inbound requests from Japanese domain emails last month, an analyst report estimating Japanese SaaS growth at 18% CAGR, and a competitor recently opening a Tokyo office. Grade this evidence. | fail→pass | 14,748 | 13,684 | -7% | 1 | 1 | 0% | 2,276 | 3,169 | +39% | 0 | 0 | — |
▸case-12 Our infrastructure team proposes migrating core databases to GCP to save 20% on cloud costs ($200k/year). Claim: 'GCP database managed services will perform identically to AWS under peak load.' Evidence: AWS bill analysis, GCP salesperson benchmark slides, and a 1-hour synthetic load test run on a single non-production microservice. Grade this evidence pile. | pass→pass | 12,529 | 13,343 | +6% | 1 | 1 | 0% | 1,917 | 3,096 | +62% | 0 | 0 | — |
▸case-13 We are deciding whether to spend $50,000 on HIPAA compliance certification. Claim: 'Lack of HIPAA compliance is blocking $500k in enterprise healthcare pipeline deals.' Evidence: CRM pipeline logs with 4 deals marked 'Lost - Security/Compliance', notes from 2 sales reps, and a survey of 15 prospect security officers. Evaluate this evidence. | fail→pass | 14,780 | 12,473 | -16% | 1 | 1 | 0% | 2,408 | 2,984 | +24% | 0 | 0 | — |
▸case-14 Operations wants to hire 2 full-time support agents to offer 24/7 live chat support on our web app. Claim: 'Live chat availability reduces trial user churn.' Evidence: 100 survey responses where 70% said they prefer live chat over email support, 5 customer quotes praising chat support during a 3-day trial, and a case study from an unrelated B2C app. Grade this data. | fail→fail | 13,140 | 13,132 | -0% | 1 | 1 | 0% | 2,077 | 3,345 | +61% | 0 | 0 | — |
▸case-15 Engineering wants to deprecate our legacy Android app to save engineering maintenance hours. Claim: 'Fewer than 1% of active paid users rely on the legacy Android app.' Evidence: Telemetry data from the last 90 days showing 0.8% of daily active users on legacy Android builds, 3 passionate support tickets begging us not to deprecate, and an email from a major enterprise client stating their staff uses Android devices exclusively. Evaluate the evidence. | fail→pass | 12,055 | 13,683 | +14% | 1 | 1 | 0% | 2,118 | 3,272 | +54% | 0 | 0 | — |
▸case-16 Marketing wants to pivot our core product messaging to highlight AI text generation capabilities. Claim: 'Positioning as an AI-first tool increases landing page visitor-to-lead conversion.' Evidence: 5 customer interviews praising the AI feature, internal team excitement, and a 50-person survey where respondents ranked AI highly. Evaluate sufficiency for an irreversible domain and brand repositioning. | fail→pass | 13,291 | 8,929 | -33% | 1 | 1 | 0% | 1,977 | 2,541 | +29% | 0 | 0 | — |
▸case-17 Sales wants to offer a guaranteed 99.99% uptime SLA with financial penalties to close enterprise prospects. Claim: 'Our current infrastructure reliably achieves four nines uptime.' Evidence: Past 12 months status page uptime logs showing 99.96% actual availability, engineering lead opinion that 'four nines is easily doable', and 0 penalty payouts last year under a non-binding SLA. Grade this pile. | pass→pass | 11,364 | 12,948 | +14% | 1 | 1 | 0% | 1,969 | 3,248 | +65% | 0 | 0 | — |
▸case-18 Product management wants to add Google and Github OAuth login to the web sign-up flow. Claim: 'Adding OAuth will reduce sign-up drop-off rate by 15%.' Evidence: 12 user interviews where participants mentioned forgetting passwords, a competitor analysis showing 8/10 top competitors offer Google OAuth, and internal product manager hypothesis. Grade this evidence and recommend the cheapest next step. | pass→pass | 13,449 | 13,810 | +3% | 1 | 1 | 0% | 2,100 | 3,185 | +52% | 0 | 0 | — |
▸case-19 Executive team wants to change our company brand name and logo ($250k total cost). Claim: 'Our current company name causes target enterprise clients to view us as a cheap SMB utility.' Evidence: 3 comments from advisory board members, 1 lost enterprise deal feedback note mentioning the brand name sounded consumer-grade, and a brand awareness survey of 40 CMOs. Grade the evidence. | fail→pass | 13,352 | 13,354 | +0% | 1 | 1 | 0% | 2,099 | 3,221 | +53% | 0 | 0 | — |
▸case-20 Write a 5-question qualitative user interview script to gather feedback on our project management dashboard without asking leading questions. | pass→pass | 10,363 | 12,856 | +24% | 1 | 1 | 0% | 1,570 | 3,152 | +101% | 0 | 0 | — |
▸case-21 Create a decision journal record documenting our choice to select PostgreSQL over MongoDB for our core transactional database, including context, alternatives considered, and expected outcomes. | pass→fail | 16,674 | 31,153 | +87% | 1 | 1 | 0% | 2,564 | 5,963 | +133% | 0 | 0 | — |
▸case-22 Synthesize the following 3 interview snippets into thematic clusters regarding why users abandon onboarding: 1) 'I couldn't find where to upload my CSV', 2) 'The CSV format error message was unclear', 3) 'I didn't have my CSV file ready when signing up'. | pass→pass | 6,197 | 9,703 | +57% | 1 | 1 | 0% | 1,066 | 2,841 | +167% | 0 | 0 | — |