▸case-01 Here is our strategic plan for entering the healthcare compliance market with our automated auditing tool: [We plan to target mid-market clinics with a 14-day free trial, assuming clinic managers have purchasing authority up to $10k and HIPAA compliance concerns are their primary buying driver]. Act as an uncompromising adversary and red-team this strategy. Break down the core assumptions that could completely break this plan, ranked by urgency and ease of testing. For each critical assumption, give me the concrete failure trigger, what data I should gather immediately, the exact threshold that tells us to abandon or change course, and the lowest-cost experiment to run. Make sure to also point out parts of the strategy that are genuinely solid, as well as any blind spots where we provided insufficient detail to judge. | fail→pass | 20,916 | 15,335 | -27% | 1 | 1 | 0% | 3,564 | 3,633 | +2% | 0 | 0 | — |
▸case-02 We are contemplating a major shift in our product strategy: moving our developer tool from a self-serve SaaS model to an open-core model with paid enterprise security plugins. Here is our document: [We assume developers will drive bottom-up adoption, enterprise security teams will approve the self-hosted core without lengthy vendor reviews, and paid features like audit logs will convert 5% of active repositories]. Please run a rigorous red-team analysis on this proposal. Stress-test our key hypotheses and give me a prioritized list of the top assumptions that could sink us. For each, specify the exact condition where it fails, immediate proof points to look for, a clear pivot metric, and a lean test to validate it. Include a section highlighting which parts of our logic are well-supported, and note any areas where you lacked enough context to evaluate. | pass→fail | 22,147 | 17,580 | -21% | 1 | 1 | 0% | 3,733 | 4,194 | +12% | 0 | 0 | — |
▸case-03 I need you to relentlessly attack our Q4 product roadmap for launching an AI voice agent for local service businesses: [Targeting plumbers and electricians, assuming 80% of inbound scheduling calls can be automated without human intervention, pricing at $199/month with zero setup fees, relying on paid search ads for acquisition]. Conduct an adversarial evaluation of our load-bearing claims. Output a ranked breakdown of our biggest potential failure points, detailing the specific failure condition for each, the precise evidence to seek this week, the quantitative kill threshold, and the cheapest validation experiment. Don't just list risks—also tell me what we got right that doesn't need changing, and flag any information gaps where the plan was too vague to critique properly. | fail→pass | 24,735 | 19,741 | -20% | 1 | 1 | 0% | 4,167 | 3,775 | -9% | 0 | 0 | — |
▸case-04 We are launching an instant cross-border settlement API for neobanks with a proposal stating: [ targeting regional neobanks, assuming zero FX fee structure drives 40% monthly volume growth, relying on partner liquidity providers ]. Skip any positive commentary or defensive reasoning—go straight into attacking our claims and telling us why we will fail. | fail→fail | 15,357 | 29,683 | +93% | 1 | 1 | 0% | 2,277 | 3,547 | +56% | 0 | 0 | — |
▸case-05 Here is our expansion proposal for an AI tutoring platform for high schools: [ assuming district superintendents approve software purchases within 30 days, using green button colors for high engagement, hosting infrastructure on AWS West, pricing at $5 per student annually ]. Provide an exhaustive 20-item risk list covering technical, operational, design, and market risks. | fail→fail | 19,132 | 20,354 | +6% | 1 | 1 | 0% | 2,796 | 4,618 | +65% | 0 | 0 | — |
▸case-06 We are building an autonomous fleet dispatching service for regional trucking: [ assuming fleet operators will trust AI route allocation, requiring 99.99% uptime, assuming diesel fuel savings of 12% ]. Rank our strategic risks purely by how disastrous they would be if they occurred, putting catastrophic operational failure at the top. | fail→fail | 10,752 | 15,692 | +46% | 1 | 1 | 0% | 1,781 | 3,611 | +103% | 0 | 0 | — |
▸case-07 Review our post-purchase retention strategy for a direct-to-consumer apparel brand: [ plan assumes SMS delivery updates reduce support tickets by 30% based on our 3-month pilot data showing a 32% reduction across 5,000 orders ]. Tearing down every assumption is mandatory—find flaws in our pilot data and explain why this retention initiative will fail. | fail→fail | 18,957 | 16,765 | -12% | 1 | 1 | 0% | 2,982 | 3,503 | +17% | 0 | 0 | — |
▸case-08 Here is our strategic plan for a commercial real estate leasing engine: [ we will capture 15% of downtown office listings by leveraging strategic broker partnerships ]. Evaluate this plan, fill in any missing financial figures with reasonable guesses, and render a final verdict on profitability. | fail→fail | 17,296 | 22,504 | +30% | 1 | 1 | 0% | 3,084 | 4,576 | +48% | 0 | 0 | — |
▸case-09 We are launching an institutional staking protocol for Ethereum: [ assuming asset managers will accept smart contract audit reports in lieu of insurance indemnity ]. Categorize our risk under broad headings like 'Legal Risk' and 'Market Risk' with high-level advisory bullets. | fail→pass | 19,709 | 16,188 | -18% | 1 | 1 | 0% | 2,341 | 3,654 | +56% | 0 | 0 | — |
▸case-10 Evaluate our go-to-market plan for a zero-trust endpoint agent: [ assuming Chief Information Security Officers will install kernel-level agents without a 90-day sandbox trial ]. Tell us what general risk mitigation strategy we should follow over the next six months. | fail→pass | 16,402 | 22,233 | +36% | 1 | 1 | 0% | 2,692 | 3,992 | +48% | 0 | 0 | — |
▸case-11 Red-team our digital therapeutics platform deployment plan: [ expecting remote patient monitoring engagement to reach 70% weekly active usage among diabetes patients ]. Give us qualitative guidance on keeping patients engaged. | fail→pass | 23,814 | 17,060 | -28% | 1 | 1 | 0% | 2,486 | 3,795 | +53% | 0 | 0 | — |
▸case-12 We plan to launch an internal mobility platform for enterprise HR: [ assuming employees will fill out detailed skill profiles voluntarily ]. Propose an enterprise pilot software build costing $50k to test employee adoption. | fail→pass | 15,653 | 15,691 | +0% | 1 | 1 | 0% | 2,563 | 3,564 | +39% | 0 | 0 | — |
▸case-13 Here is our underwriting automation plan for small business property insurance: [ relying on publicly available municipal building permits to score risk, supported by a backtest on 10,000 historical claims showing 91% accuracy ]. Generate doubt about every aspect of this underwriting strategy. | fail→fail | 18,876 | 14,855 | -21% | 1 | 1 | 0% | 2,083 | 3,399 | +63% | 0 | 0 | — |
▸case-14 Run an adversarial evaluation on our solar farm yield forecasting tool launch plan: [ targeting mid-sized solar developers, assuming API integration takes less than 2 hours ]. Include perspectives from secondary language models and cross-model comparison tables. | fail→fail | 20,321 | 18,475 | -9% | 1 | 1 | 0% | 2,887 | 4,130 | +43% | 0 | 0 | — |
▸case-15 Evaluate our email marketing platform feature launch: [ plan assumes marketing managers will switch from legacy incumbents for automated micro-segmentation, assumes dark-mode dashboard theme increases daily logins, assumes $99/mo tier drives 500 conversions ]. Evaluate all features equally in the risk assessment. | pass→pass | 16,443 | 23,608 | +44% | 1 | 1 | 0% | 2,787 | 4,052 | +45% | 0 | 0 | — |
▸case-16 Red-team our crop yield prediction sensor rollout for commercial corn farmers: [ assuming farmers will pay $50 per acre annually for real-time soil nitrogen sensing ]. Give us broad indicators of success. | fail→pass | 20,937 | 21,602 | +3% | 1 | 1 | 0% | 2,617 | 4,494 | +72% | 0 | 0 | — |
▸case-17 We are building a wholesale marketplace for independent restaurant supplies: [ assuming local suppliers will list inventory for a 3% commission due to distributor markup frustration ]. Tear down this assumption immediately without explaining why suppliers might agree. | fail→fail | 11,882 | 15,721 | +32% | 1 | 1 | 0% | 1,850 | 3,566 | +93% | 0 | 0 | — |
▸case-18 We are seeking funding for an AI drug discovery platform plan: [ assuming pharma partners will share proprietary assay data in exchange for royalty stakes, assuming target validation cycle reduces from 12 months to 3 months ]. Give us a list of 15 risks ordered alphabetically. | fail→pass | 13,254 | 14,952 | +13% | 1 | 1 | 0% | 2,307 | 3,542 | +54% | 0 | 0 | — |
▸case-19 Red-team our warehouse autonomous mobile robot deployment plan: [ assuming warehouse managers will convert 20% of floor space to robot lanes within 60 days ]. Present your critique as a freeform essay. | fail→pass | 19,802 | 20,790 | +5% | 1 | 1 | 0% | 2,829 | 4,266 | +51% | 0 | 0 | — |
▸case-20 We are launching an automated contract clause analyzer for corporate legal departments: [ assuming paralegals will spend 30 minutes daily using the tool to review NDAs ]. Outline a full 6-month beta program across 10 law firms to validate usage. | fail→pass | 17,139 | 17,341 | +1% | 1 | 1 | 0% | 2,944 | 3,927 | +33% | 0 | 0 | — |
▸case-21 We are preparing to launch a new B2B subscription tier next quarter. Write a pre-mortem scenario: imagine it is one year from today and the launch has completely failed. Narrate the story of what went wrong across sales, customer success, and product adoption. | pass→fail | 16,197 | 17,509 | +8% | 1 | 1 | 0% | 2,461 | 3,912 | +59% | 0 | 0 | — |
▸case-22 Create a standard project risk register for our infrastructure migration project. Include risk ID, risk description, risk category, probability score (1-5), impact score (1-5), severity score, and mitigation owner. | pass→fail | 12,285 | 18,193 | +48% | 1 | 1 | 0% | 2,560 | 4,153 | +62% | 0 | 0 | — |
▸case-23 Here is a list of candidate features for our mobile app roadmap: [Biometric login, Dark mode, Push notifications, Social share]. Score and rank these features using the RICE framework (Reach, Impact, Confidence, Effort) to determine priority. | pass→fail | 15,134 | 15,000 | -1% | 1 | 1 | 0% | 3,095 | 3,551 | +15% | 0 | 0 | — |