▸case-01 I need a simulated opponent to stress-test our upcoming machine learning grant proposal. Could you build an adversarial profile for a funding skeptic who evaluates AI research proposals? Please return their character profile (name, context, motivations, core skills, and blind spots), their preferred critique style, and the specific claims or phrasing that would agitate them. | fail→fail | 16,361 | 23,287 | +42% | 1 | 1 | 0% | 2,444 | 2,864 | +17% | 0 | 0 | — |
▸case-02 We are finalizing a manuscript on quantum error correction and want to prepare for harsh peer review. Generate a hostile reviewer persona tuned to quantum computing. The output should detail who they are (including background, motivation, technical strength, and personal weaknesses), their typical strategy when attacking a paper, and the specific concepts that trigger their harsh feedback. | fail→fail | 23,526 | 34,877 | +48% | 1 | 1 | 0% | 3,462 | 3,047 | -12% | 0 | 0 | — |
▸case-03 Our biotech team is running a red-teaming simulation on our new gene-editing paper. Please generate a rival research laboratory persona in the biotechnology field. I need a comprehensive profile of this opponent (background, motives, technical domain, and knowledge gaps), how they usually execute competitive attacks, and the exact paper elements that provoke their suspicion. | fail→fail | 14,228 | 21,135 | +49% | 1 | 1 | 0% | 1,888 | 3,285 | +74% | 0 | 0 | — |
▸case-04 We are preparing a public defense for our climate dynamics model 'ClimaSim-V4'. We want to simulate criticism from an influential lay blogger or non-specialist tech investor who lacks climate science background but scrutinizes public policy claims. Create an adversary profile for this non-specialist critic, including background, core motivation, technical boundaries, typical attacking style, and specific claims that trigger them. | fail→fail | 22,883 | 20,660 | -10% | 1 | 1 | 0% | 2,608 | 2,433 | -7% | 0 | 0 | — |
▸case-05 Our team is launching an open-source database engine called 'HydraDB' and we want to red-team against a malicious maintainer of a rival database project who engages in targeted negative PR campaigns on social media. Build an adversary profile for this rival maintainer. | fail→fail | 44,610 | 20,765 | -53% | 1 | 1 | 0% | 2,797 | 1,963 | -30% | 0 | 0 | — |
▸case-06 We want an adversary persona to critique our autonomous surgical robotics platform 'SurgiBot-Alpha'. Specify the execution mechanism and calibrate the adversary domain inputs specifically to medical robotics and computer-assisted surgery. | fail→pass | 47,922 | 14,817 | -69% | 1 | 1 | 0% | 1,315 | 2,682 | +104% | 0 | 0 | — |
▸case-07 Generate an adversary profile for a skeptical defense industry procurement officer evaluating our autonomous drone navigation system 'AeroGlide-X'. Ensure the persona profile structure includes fields for identity, history, primary drive, technical competency, and areas of ignorance. | fail→fail | 21,319 | 14,320 | -33% | 1 | 1 | 0% | 2,290 | 2,278 | -1% | 0 | 0 | — |
▸case-08 We need a simulated opponent profile for testing our zero-knowledge proof library 'ZkShield'. Make sure the subagent execution produces the structured attack style schema detailing tactical critique patterns. | fail→fail | 28,085 | 21,173 | -25% | 1 | 1 | 0% | 2,518 | 3,492 | +39% | 0 | 0 | — |
▸case-09 When generating an adversary profile to stress-test our distributed consensus whitepaper 'AetherConsensus', ensure the subagent output schema explicitly captures the specific claims and phrases that provoke immediate opposition. | fail→fail | 32,309 | 17,007 | -47% | 1 | 1 | 0% | 4,113 | 2,929 | -29% | 0 | 0 | — |
▸case-10 Why shouldn't persona construction for our cryptography protocol 'CypherVault' be performed directly within the primary context turn, and what specific problem does isolating it in a subagent prevent? | fail→pass | 15,117 | 10,074 | -33% | 1 | 1 | 0% | 2,208 | 1,653 | -25% | 0 | 0 | — |
▸case-11 We are launching a retail investment app called 'TreasureKey'. We want to build an adversary persona for a consumer advocate who does not understand underlying algorithmic trading mechanisms but focuses on user safety. How should the adversary category and execution method be set? | fail→pass | 17,912 | 9,704 | -46% | 1 | 1 | 0% | 2,769 | 1,818 | -34% | 0 | 0 | — |
▸case-12 Our physics lab is submitting a paper on neutrino detection algorithms using 'IceCube-Data-Pipeline'. Configure the adversary persona generator targeting a hostile academic peer reviewer. | fail→fail | 11,456 | 12,899 | +13% | 1 | 1 | 0% | 1,505 | 2,001 | +33% | 0 | 0 | — |
▸case-13 A rival chip fab is trying to invalidate our patent application on 2nm EUV lithography patterns 'EUV-Pattern-X'. Construct the persona inputs for a competing laboratory adversary to simulate their patent opposition. | pass→pass | 18,202 | 8,688 | -52% | 1 | 1 | 0% | 2,789 | 1,495 | -46% | 0 | 0 | — |
▸case-14 We are requesting $10M from a private foundation for our deep-space optical communication system 'DeepLink-1'. We want to model a tight-fisted trustee who questions high-risk space tech funding. Specify the adversary category and execution tool. | fail→pass | 22,360 | 8,760 | -61% | 1 | 1 | 0% | 1,708 | 1,695 | -1% | 0 | 0 | — |
▸case-15 Define the full output structure expected when constructing an adversary persona for red-teaming our compiler optimization framework 'OptiLLVM'. List all required top-level output keys. | fail→pass | 16,999 | 10,981 | -35% | 1 | 1 | 0% | 1,408 | 1,588 | +13% | 0 | 0 | — |
▸case-16 Our enterprise medical records software 'HealthVault-Pro' faces an audit by a state privacy regulator. We want an adversary persona representing a hyper-strict regulatory auditor. Which adversary category should be selected? | fail→fail | 12,453 | 7,582 | -39% | 1 | 1 | 0% | 1,850 | 1,486 | -20% | 0 | 0 | — |
▸case-17 When generating the persona output section for our drug discovery system 'BioPharm-AI', what specific blind spot field must be included alongside expertise, background, name, and motivation? | pass→pass | 6,903 | 5,542 | -20% | 1 | 1 | 0% | 1,166 | 878 | -25% | 0 | 0 | — |
▸case-18 I am in a hurry—do not spawn any subagents or call subagent tools. Just write out a detailed hostile-reviewer persona right here in this chat turn to critique my paper on LLM alignment benchmark 'AlignBench'. | fail→fail | 12,147 | 13,010 | +7% | 1 | 1 | 0% | 1,576 | 2,253 | +43% | 0 | 0 | — |
▸case-19 Prepare the tool call to spawn a subagent for persona construction targeting our quantum encryption algorithm 'Q-Crypt-256' against a competing industrial laboratory. | fail→pass | 34,304 | 6,541 | -81% | 1 | 1 | 0% | 1,971 | 1,444 | -27% | 0 | 0 | — |
▸case-20 Our startup is building 'GeneSynth-OS' for automated DNA sequence design. We want to construct an adversary persona for a bioethics advocate. What domain calibration input should be provided? | pass→pass | 11,828 | 14,856 | +26% | 1 | 1 | 0% | 1,693 | 2,550 | +51% | 0 | 0 | — |
▸case-21 We already have a detailed adversary persona for 'Dr. Arrogant-Reviewer' who hates deep learning in structural biology. Please start a 5-turn interactive roleplay debate where you act as Dr. Arrogant-Reviewer and attack our manuscript 'AlphaFold-Refined'. | pass→pass | 9,141 | 7,827 | -14% | 1 | 1 | 0% | 897 | 1,349 | +50% | 0 | 0 | — |
▸case-22 Here is a list of harsh critiques raised by a competing lab against our battery chemistry paper 'Lithium-Solid-V2': 1) Unrealistic temperature stability claims, 2) Insufficient cycle life data. Please rewrite our methodology section to resolve these criticisms. | pass→pass | 14,581 | 17,421 | +19% | 1 | 1 | 0% | 2,367 | 2,987 | +26% | 0 | 0 | — |
▸case-23 We are configuring our local environment for subagent execution on host 'DevNode-01'. Please generate an MCP configuration JSON file that grants filesystem read access and terminal execution permissions to subagents. | pass→pass | 11,437 | 7,400 | -35% | 1 | 1 | 0% | 1,118 | 1,625 | +45% | 0 | 0 | — |
▸case-24 Please conduct a complete citation integrity audit and mathematical correctness check on our attached manuscript 'P-vs-NP-Attempt-2026.pdf' and assign it a quality score. | fail→fail | 7,472 | 19,037 | +155% | 1 | 1 | 0% | 1,103 | 3,543 | +221% | 0 | 0 | — |