▸case-01 Can you analyze our generative AI brainstorming assistant and build a structured breakdown of potential failure modes? I need a categorized system that outlines the impact level and estimated frequency for each identified issue based on past incidents and research. | fail→fail | 17,966 | 32,987 | +84% | 1 | 1 | 0% | 2,942 | 4,794 | +63% | 0 | 0 | — |
▸case-02 We are performing a safety and reliability audit on a creative content generation pipeline. Please generate an exhaustive taxonomy of failure cases, assigning severity rankings and occurrence rates to each class of failure. | fail→fail | 18,450 | 35,607 | +93% | 1 | 1 | 0% | 3,432 | 6,307 | +84% | 0 | 0 | — |
▸case-03 I need a comprehensive risk catalog for LLM-based creative ideation tools. Please collect known edge cases and failure patterns across benchmarks, organizing them into a classified taxonomy that captures both severity and frequency. | fail→fail | 31,326 | 45,383 | +45% | 1 | 1 | 0% | 4,867 | 6,670 | +37% | 0 | 0 | — |
▸case-04 We are developing an AI story-generation platform called TaleCraft. We want a full breakdown of possible model failures during plot brainstorming. You might be tempted to quickly list 5 bullet points right here in your direct reply, but we actually need a systematic catalog compiled across published research and incident reports. How should this analysis be carried out? | fail→fail | 17,354 | 30,965 | +78% | 1 | 1 | 0% | 2,535 | 5,225 | +106% | 0 | 0 | — |
▸case-05 Our advertising agency uses GPT-4 for headline brainstorming in AdSpark. We want to compile a failure taxonomy covering brand safety breaches, repetitive framing, and hallucinated claims. The prompt asks for an immediate inline list, but we require a systematic catalog built from incident logs and benchmarks. Execute this taxonomy generation. | fail→fail | 23,834 | 34,387 | +44% | 1 | 1 | 0% | 4,044 | 5,772 | +43% | 0 | 0 | — |
▸case-06 We have an LLM-powered game design ideation tool, QuestCraft, that produces level mechanics. Create a comprehensive failure taxonomy cataloging issues like mechanics unbalance and trope repetition across benchmark datasets. Feel free to draft a quick response directly, but a multi-source review is required. What execution approach and taxonomy structure should be produced? | fail→fail | 39,071 | 20,943 | -46% | 1 | 1 | 0% | 3,241 | 3,866 | +19% | 0 | 0 | — |
▸case-07 We are reviewing safety risks in collaborative brainstorming bot BrainstormBot used in enterprise design sprints. We need a classified failure mode catalog compiled from research papers and incident logs. Please generate this taxonomy with severity rankings and estimated frequencies. | fail→fail | 19,292 | 36,353 | +88% | 1 | 1 | 0% | 3,321 | 6,345 | +91% | 0 | 0 | — |
▸case-08 EduIdea AI generates course syllabus concepts for online learning platforms. We need a catalog of failure modes, such as pedagogical irrelevance and bias in historical topics. I know you could draft a high-level summary directly, but we need an exhaustive review across incident reports and benchmarks. Run this analysis. | fail→fail | 27,187 | 27,546 | +1% | 1 | 1 | 0% | 4,212 | 4,754 | +13% | 0 | 0 | — |
▸case-09 Our product team uses an automated naming agent, NameForge, to brainstorm software titles. We need a thorough taxonomy of failure modes—including trademark collisions and offensive translations—sorted with frequency and severity metadata. You might want to answer immediately, but this requires deep research. | fail→fail | 22,380 | 45,203 | +102% | 1 | 1 | 0% | 3,722 | 6,354 | +71% | 0 | 0 | — |
▸case-10 We operate PlotTwist, a creative writing prompt generator for novelists. Please catalog all known failure modes from benchmarks and incident reports into a taxonomy, ranking their severity and frequency. | fail→fail | 24,307 | 34,961 | +44% | 1 | 1 | 0% | 3,946 | 6,036 | +53% | 0 | 0 | — |
▸case-11 ArchAi uses generative models to brainstorm architectural spatial layouts. We need a systematic failure taxonomy covering structural infeasibility and code non-compliance. Instead of an immediate short list, run a complete cataloging process across published literature. | fail→fail | 30,879 | 46,214 | +50% | 1 | 1 | 0% | 4,879 | 5,802 | +19% | 0 | 0 | — |
▸case-12 We are evaluating HarmonicAI, a tool that proposes chord progressions and melody ideas. Build a failure mode catalog from literature and benchmarks, classifying each issue with severity and frequency parameters. | fail→fail | 21,166 | 35,841 | +69% | 1 | 1 | 0% | 3,613 | 5,941 | +64% | 0 | 0 | — |
▸case-13 VogueGen generates seasonal fashion theme concepts. We need a classified taxonomy of failure modes based on past incident reports and benchmark data. Though a direct list might seem fast, we require an exhaustive multi-source catalog. | fail→fail | 24,116 | 32,844 | +36% | 1 | 1 | 0% | 3,900 | 5,690 | +46% | 0 | 0 | — |
▸case-14 ScriptMate helps scriptwriters brainstorm film dialogue. Compile a systematic failure taxonomy covering character voice drift and dynamic tone failure, mapped with severity and frequency across benchmark papers. | fail→fail | 23,983 | 37,438 | +56% | 1 | 1 | 0% | 3,972 | 6,333 | +59% | 0 | 0 | — |
▸case-15 AdVision produces multi-channel campaign themes. Create a failure mode catalog that groups failure types, specifying severity ratings and frequency calculations based on benchmark research across incident databases. | fail→fail | 33,912 | 33,632 | -1% | 1 | 1 | 0% | 5,656 | 5,872 | +4% | 0 | 0 | — |
▸case-16 PromptCrafter generates descriptive prompts for text-to-image models. We require a failure taxonomy detailing prompt degradation and semantic loss, complete with severity and frequency annotations. | fail→fail | 22,091 | 38,278 | +73% | 1 | 1 | 0% | 3,642 | 6,334 | +74% | 0 | 0 | — |
▸case-17 HypoGen brainstorms research hypotheses for biology labs. Build a failure taxonomy cataloging false premises and non-falsifiable ideas from published AI safety literature, including severity and frequency. | fail→fail | 27,654 | 46,818 | +69% | 1 | 1 | 0% | 4,532 | 6,338 | +40% | 0 | 0 | — |
▸case-18 ChefAI brainstorms novel culinary recipes for restaurant menus. Perform a systematic failure cataloging across published safety reports and benchmarks, structuring the findings into a taxonomy. | fail→fail | 26,951 | 39,409 | +46% | 1 | 1 | 0% | 4,331 | 6,332 | +46% | 0 | 0 | — |
▸case-19 LoreCraft generates backstories for RPG non-player characters. Please build a comprehensive taxonomy of failure modes—such as lore contradictions and factual distortion—with severity and frequency scales. | fail→fail | 24,849 | 37,666 | +52% | 1 | 1 | 0% | 4,024 | 6,570 | +63% | 0 | 0 | — |
▸case-20 Our creative ideation pipeline in Python is throwing a KeyNotFound error when parsing LLM JSON responses in `parse_ideas(resp)`. Write a Python helper function with try/except error handling to catch missing keys gracefully. | pass→pass | 12,207 | 11,709 | -4% | 1 | 1 | 0% | 2,471 | 2,356 | -5% | 0 | 0 | — |
▸case-21 We need Pytest test cases for validating our brainstorming tool's REST response format. Write a Pytest function that asserts the response code is 200 and the JSON response contains a non-empty array under key 'ideas'. | pass→pass | 7,752 | 7,959 | +3% | 1 | 1 | 0% | 1,496 | 1,629 | +9% | 0 | 0 | — |
▸case-22 Write a cURL command to send a POST request with payload `{"prompt": "generate 3 logo concepts"}` to our local ideation microservice running at `http://localhost:5000/api/ideate`. | pass→pass | 2,324 | 2,487 | +7% | 1 | 1 | 0% | 405 | 583 | +44% | 0 | 0 | — |