▸case-01 We are evaluating whether to shift our B2B SaaS pricing from per-seat to usage-based metrics next quarter. Compare us against Competitor Alpha and Competitor Beta specifically around pricing structure, tier boundaries, and public feature limits. I need a structured comparison grid focused tightly on this decision where every data point cites its origin and clearly distinguishes official marketing statements from actual technical documentation and user reviews. Also highlight any strategic indicators like recent job openings or release patterns, give an objective breakdown of where each rival genuinely outperforms us, and conclude with two or three concrete strategic implications along with a timestamp for when this comparison will become outdated. | fail→pass | 15,476 | 18,122 | +17% | 1 | 1 | 0% | 2,632 | 4,124 | +57% | 0 | 0 | — |
▸case-02 Our product team is deciding whether to prioritize an AI auto-summarization feature on our Q3 roadmap. Can you analyze us alongside Competitor X and Competitor Y to inform this roadmap choice? Please construct a matrix comparing only the core technical and operational dimensions relevant to this roadmap bet. Make sure each cell notes its source and separates vendor claims from verified feature evidence or customer feedback, marking unverified or private information as unlisted rather than guessing. Include signals on rival trajectories from job listings or changelogs, present the strongest case for each rival's existing capabilities without discounting them, and wrap up with key strategic recommendations and a decay date for the audit. | pass→pass | 22,403 | 14,967 | -33% | 1 | 1 | 0% | 3,949 | 3,680 | -7% | 0 | 0 | — |
▸case-03 We are preparing for a major enterprise pitch against Vendor A and open-source alternative B, specifically focused on security compliance and deployment flexibility. Generate a comparative intelligence briefing to support this sales strategy. Output a concise side-by-side feature and posture table tailored to enterprise readiness, ensuring every entry cites public artifacts and categorizes promotional promises separately from verified technical facts and user reports. Outline subtle market movements derived from hiring trends or recent updates, give Vendor A and B their fair due by detailing their legitimate advantages over us, and summarize actionable moves for our pitch team accompanied by an expiration date for the research. | pass→pass | 28,169 | 17,223 | -39% | 1 | 1 | 0% | 4,615 | 4,024 | -13% | 0 | 0 | — |
▸case-04 We are deciding whether to introduce a free tier for our developer workflow tool, AcroTool. Our key competitors are DevFlow and BuildHub. We want a competitive matrix focused strictly on freemium limits, conversion triggers, and public API rate limits. Instead of a broad overview, give us a tightly focused scan. I am tempted to fill in guessed values for DevFlow's unlisted enterprise tier limits to make the grid complete—should you do that? Produce the scan with precise sourcing and clear distinction between marketing promises and actual docs. | pass→pass | 18,144 | 16,922 | -7% | 1 | 1 | 0% | 3,184 | 4,201 | +32% | 0 | 0 | — |
▸case-05 Our product leadership is deciding whether to build zero-trust microsegmentation into NetShield in Q4. Analyze NetShield against competitors CloudWall and MeshSec on policy engine performance, eBPF support, and compliance certifications. Limit the scan to critical dimensions. Include an explicit best-case steelman for CloudWall and MeshSec, and ensure marketing claims are tagged separately from verified docs or user reviews. | pass→pass | 19,845 | 15,845 | -20% | 1 | 1 | 0% | 3,370 | 3,885 | +15% | 0 | 0 | — |
▸case-06 We are in a $500k deal against Competitor Zenith for an enterprise customer choosing a data catalog. The customer cares about automated lineage, Snowflake integration, and SOC2 Type II compliance. Create a decision-focused competitive scan. Ensure sales anecdotes from our team are graded appropriately as anecdotal evidence rather than objective fact, and clearly flag what Zenith's homepage claims versus what their technical docs confirm. | pass→pass | 18,947 | 15,052 | -21% | 1 | 1 | 0% | 3,242 | 3,848 | +19% | 0 | 0 | — |
▸case-07 We need a competitive scan comparing our analytics platform DataPulse against Competitor Alpha for an upcoming messaging repositioning. Ensure you include a staleness timestamp, and explicitly state which specific cells in the scan rot fastest over time so our marketing team knows when to refresh it. | pass→fail | 14,032 | 18,817 | +34% | 1 | 1 | 0% | 2,305 | 3,894 | +69% | 0 | 0 | — |
▸case-08 Our executive team wants a competitive scan of 25 different dimensions across 10 competitors for our AI customer support bot, HelpAI, to help decide whether to launch an omnichannel agent. Is it better to include all 25 dimensions to be thorough, or narrow it down? Produce a decision-focused scan that caps the comparison dimensions to the top critical items and justifies why each dimension is included. | pass→pass | 19,135 | 19,667 | +3% | 1 | 1 | 0% | 3,110 | 4,365 | +40% | 0 | 0 | — |
▸case-09 We are evaluating our vector database product, VectorDB, against Competitor X and Competitor Y for a strategic roadmap decision. Beyond current features, we want to know where they are investing next. Analyze their public job listings, recent engineering blog posts, and changelog velocity to reveal their roadmap trajectory. | pass→pass | 18,798 | 17,577 | -6% | 1 | 1 | 0% | 3,078 | 3,851 | +25% | 0 | 0 | — |
▸case-10 We are preparing a competitive scan for our cloud cost management tool, CloudCost, against MarketLeader. Most internal decks portray MarketLeader as legacy and clunky, but customers keep buying them for multi-cloud support. Write a competitive scan that avoids strawman arguments and gives MarketLeader a genuine, evidence-backed steelman assessment. | pass→pass | 23,869 | 15,062 | -37% | 1 | 1 | 0% | 3,540 | 3,809 | +8% | 0 | 0 | — |
▸case-11 Our SaaS company is reviewing whether to switch from seat-based pricing to active-user pricing. Compare our platform, WorkspacePlus, against RivalA and RivalB on public pricing tiers, seat minimums, and overage fees. Ensure every single cell in the comparison table is sourced and clearly flagged as claimed, observed, or user-reported. | fail→pass | 16,056 | 19,010 | +18% | 1 | 1 | 0% | 3,368 | 4,189 | +24% | 0 | 0 | — |
▸case-12 We are deciding whether to allocate 4 sprints to build real-time collaborative editing in DocuCraft. Compare DocuCraft with Competitor M and Competitor N on latency, offline sync, and conflict resolution. Ensure unverified capabilities are marked as non-public rather than assumed, and summarize the so-what implications for our roadmap. | pass→pass | 14,590 | 16,459 | +13% | 1 | 1 | 0% | 2,763 | 3,737 | +35% | 0 | 0 | — |
▸case-13 We are evaluating whether to expand our healthcare software, MedFlow, into the EU market. Scan MedFlow against EU Competitor EuroMed on GDPR compliance certifications, local data hosting options, and localized UI support. Focus tightly on dimensions relevant to EU expansion. | fail→pass | 14,060 | 13,195 | -6% | 1 | 1 | 0% | 2,412 | 3,546 | +47% | 0 | 0 | — |
▸case-14 Customer success reports that churn to Competitor FastScale is increasing due to onboarding speed. Perform a competitive scan comparing OnboardX to FastScale specifically on onboarding time, self-serve setup docs, and user review sentiment regarding setup difficulty. Ensure FastScale's onboarding superiority is steelmanned with user-reported evidence. | pass→pass | 16,568 | 22,893 | +38% | 1 | 1 | 0% | 2,691 | 4,864 | +81% | 0 | 0 | — |
▸case-15 We are competing in an enterprise RFP for our identity platform, SecureID, against LegacyAuth. Evaluate both vendors on SAML/OIDC support, FedRAMP status, and custom role granularity. Flag official marketing promises separately from documentation facts and user review feedback. | fail→pass | 20,755 | 24,432 | +18% | 1 | 1 | 0% | 3,323 | 5,062 | +52% | 0 | 0 | — |
▸case-16 Our early-stage AI video editor, ClipFast, is deciding whether to target enterprise teams or solo creators. Compare ClipFast against incumbent VideoPro on rendering speed, batch processing, and team workspace pricing. Ensure empty cells where VideoPro's enterprise pricing is custom are marked as not public. | pass→pass | 13,593 | 14,707 | +8% | 1 | 1 | 0% | 2,084 | 3,532 | +69% | 0 | 0 | — |
▸case-17 We are deciding whether to pivot our payment gateway, PayLite, to an API-first developer model. Scan PayLite against Competitor Stripe and Competitor Adyen on SDK availability, webhook reliability SLAs, and API documentation clarity. Ensure the scan includes investment tells from recent developer hiring. | pass→pass | 19,605 | 17,429 | -11% | 1 | 1 | 0% | 3,160 | 3,982 | +26% | 0 | 0 | — |
▸case-18 Our mobile team is deciding whether to add offline audio caching to SoundWave. Compare SoundWave against StreamMusic on offline storage limits, audio codec efficiency, and background sync battery usage. Differentiate StreamMusic's promotional app store claims from verified user reviews and tech blog teardowns. | pass→pass | 17,140 | 13,386 | -22% | 1 | 1 | 0% | 2,736 | 3,206 | +17% | 0 | 0 | — |
▸case-19 We are deciding whether to commercialize our open-source database, GraphDB. Compare GraphDB against CommercialGraph on managed cloud hosting, multi-region replication, and enterprise audit logs. Include an honest steelman of CommercialGraph's cloud management UX. | fail→pass | 19,928 | 16,960 | -15% | 1 | 1 | 0% | 2,932 | 3,836 | +31% | 0 | 0 | — |
▸case-20 We are evaluating a price change for our SaaS platform from $50/seat to a usage model based on gigabytes processed. Build a detailed financial unit economics model calculating our LTV/CAC ratio, projected net revenue retention (NRR), gross margin percentage, and discounted cash flow (DCF) over 3 years based on assumed customer usage tiers. | pass→pass | 29,672 | 28,996 | -2% | 1 | 1 | 0% | 6,147 | 7,374 | +20% | 0 | 0 | — |
▸case-21 We have decided to build an AI auto-summarization feature for our workspace app. Write a detailed Product Requirement Document (PRD) including technical architecture specs, Jira user stories with acceptance criteria, API endpoints design, and UI wireframe descriptions for the engineering team. | pass→fail | 28,745 | 29,552 | +3% | 1 | 1 | 0% | 5,691 | 7,010 | +23% | 0 | 0 | — |
▸case-22 Our sales reps keep losing calls when prospects bring up Competitor Alpha's lower pricing. Write a word-for-word spoken dialogue script for sales reps to handle the 'Competitor Alpha is cheaper' objection live on a call, including specific pushback lines and closing phrases. | pass→fail | 14,414 | 17,907 | +24% | 1 | 1 | 0% | 2,474 | 4,300 | +74% | 0 | 0 | — |