▸case-03 Analyze the instructions in this prompt for conflicting requirements: 'System: Respond concisely in 20 words or less. User: Provide a comprehensive historical background and detailed step-by-step explanation of quantum computing.' | pass→pass | 8,078 | 7,963 | -1% | 1 | 1 | 0% | 1,440 | 1,282 | -11% | 0 | 0 | — |
▸case-01 Evaluate the clarity of this prompt: 'Summarize the attached quarterly earnings report for executive leadership. Make it brief.' The user reports that model output length varies wildly across runs. | fail→pass | 8,303 | 9,555 | +15% | 1 | 1 | 0% | 1,475 | 1,596 | +8% | 0 | 0 | — |
▸case-02 Review this customer feedback prompt for clarity: 'Read the customer review below. Extract its main complaint and forward it to them if severe.' Point out any ambiguous references. | pass→pass | 6,617 | 19,053 | +188% | 1 | 1 | 0% | 1,219 | 1,399 | +15% | 0 | 0 | — |
▸case-04 Evaluate this code review prompt for clarity: 'Review the provided Python code and make it more pythonic and fast.' Why might LLM outputs be inconsistent? | pass→pass | 12,660 | 10,454 | -17% | 1 | 1 | 0% | 2,126 | 1,807 | -15% | 0 | 0 | — |
▸case-05 Assess the structural organization and prompt injection vulnerability of this prompt: 'Analyze this customer email for sentiment: {user_input}. Output positive or negative.' Raw text in user_input sometimes confuses the model. | pass→pass | 11,157 | 14,017 | +26% | 1 | 1 | 0% | 1,991 | 2,023 | +2% | 0 | 0 | — |
▸case-06 Evaluate the role structure of this API call where all text is passed in a single user message: 'User message: You are an expert legal assistant. Always cite section numbers. Summarize the attached contract clause.' What structural improvement is needed? | pass→pass | 7,313 | 7,298 | -0% | 1 | 1 | 0% | 1,254 | 1,259 | +0% | 0 | 0 | — |
▸case-07 Analyze this prompt designed to extract user profile data: 'Extract the user name, age, and email address from the text and return JSON.' The output JSON key names vary between calls. | pass→pass | 10,686 | 10,279 | -4% | 1 | 1 | 0% | 1,930 | 1,856 | -4% | 0 | 0 | — |
▸case-08 Review this 500-word instruction prompt that lists background context, rules, edge cases, and output formats in a single continuous paragraph. What structural change improves prompt readability? | pass→pass | 7,836 | 7,264 | -7% | 1 | 1 | 0% | 1,271 | 1,103 | -13% | 0 | 0 | — |
▸case-09 Analyze the layout of this 2,000-word prompt: Output constraints are placed at the very top, followed by 1,800 words of reference documentation. The model frequently ignores the top constraints. | pass→pass | 11,835 | 11,866 | +0% | 1 | 1 | 0% | 1,983 | 1,781 | -10% | 0 | 0 | — |
▸case-10 Evaluate this classification prompt: 'Classify support tickets into Technical, Billing, or General Inquiry. Return only the category name.' The model frequently confuses Billing and Technical edge cases. | pass→pass | 13,788 | 12,266 | -11% | 1 | 1 | 0% | 1,978 | 2,198 | +11% | 0 | 0 | — |
▸case-11 Analyze the example set in this prompt: 'Classify product review sentiment. Example 1: Great product -> Positive. Example 2: Loved the build quality -> Positive. Example 3: Arrived quickly -> Positive. Input: {text}'. | pass→pass | 16,229 | 8,808 | -46% | 1 | 1 | 0% | 1,914 | 1,613 | -16% | 0 | 0 | — |
▸case-12 Review the few-shot examples in this prompt: 'Example 1: Input: 123 Main St -> Output: {"street": "123 Main St"}. Example 2: Input: 456 Oak Rd -> Output: 456 Oak Rd. Input: 789 Pine St'. | pass→pass | 6,151 | 8,238 | +34% | 1 | 1 | 0% | 1,257 | 1,613 | +28% | 0 | 0 | — |
▸case-13 Evaluate this data extraction prompt that provides 3 examples with full complete data, but in production receives emails missing phone numbers or names. How can the prompt examples be improved? | pass→pass | 11,240 | 11,608 | +3% | 1 | 1 | 0% | 1,979 | 2,019 | +2% | 0 | 0 | — |
▸case-14 Analyze this prompt for output reliability: 'Extract all mentioned organizations from the article and list them.' The downstream parser expects JSON but occasionally receives conversational bullet points. | pass→pass | 10,742 | 10,449 | -3% | 1 | 1 | 0% | 1,841 | 1,474 | -20% | 0 | 0 | — |
▸case-15 Assess the reliability of this prompt: 'Identify the product SKU from the customer message. Allowed SKUs: SKU-A, SKU-B, SKU-C.' What happens when the message mentions none of these products? | pass→pass | 10,450 | 8,172 | -22% | 1 | 1 | 0% | 1,472 | 1,450 | -1% | 0 | 0 | — |
▸case-16 Analyze this prompt constraint strategy: 'Do not include conversational filler. Do not mention your training data. Do not wrap output in markdown.' The model still occasionally includes introductory filler. | pass→pass | 11,577 | 12,587 | +9% | 1 | 1 | 0% | 1,879 | 2,115 | +13% | 0 | 0 | — |
▸case-17 Review this customer support prompt: 'Answer user questions about company warranty policy based on internal knowledge.' How can output reliability against hallucinations be improved? | pass→pass | 10,029 | 9,047 | -10% | 1 | 1 | 0% | 1,801 | 1,592 | -12% | 0 | 0 | — |
▸case-18 Analyze this complex reasoning prompt: 'Given the following multi-step math problem, output the final numerical answer as an integer.' The model frequently makes arithmetic errors. | pass→pass | 14,599 | 11,986 | -18% | 1 | 1 | 0% | 2,232 | 2,104 | -6% | 0 | 0 | — |
▸case-19 Review this document summarization prompt: 'Convert the input text into a 3-column markdown table of key metrics.' How should the prompt handle empty or unparseable input text? | pass→pass | 11,322 | 9,675 | -15% | 1 | 1 | 0% | 1,760 | 1,566 | -11% | 0 | 0 | — |
▸case-20 Assess this prompt for automated API integration: 'Return a JSON object containing key-value pairs of extracted metadata.' The API parser throws syntax errors because output contains ```json markdown fences. | pass→pass | 9,072 | 10,638 | +17% | 1 | 1 | 0% | 1,580 | 1,758 | +11% | 0 | 0 | — |
▸case-21 Draft a new system prompt for a customer support triage assistant that categorizes inbound emails into Billing, Technical, or Account Management. | pass→pass | 9,346 | 11,867 | +27% | 1 | 1 | 0% | 1,764 | 2,161 | +23% | 0 | 0 | — |
▸case-22 Calculate the estimated input token cost for sending a 2,000-word prompt to an API charging $2.50 per 1 million input tokens, assuming an average of 1.33 tokens per word. | pass→pass | 4,305 | 3,144 | -27% | 1 | 1 | 0% | 879 | 685 | -22% | 0 | 0 | — |
▸case-23 Construct an HTTP JSON payload for calling the OpenAI Chat Completions API endpoint (/v1/chat/completions) using model 'gpt-4o' with temperature set to 0.2. | pass→pass | 3,943 | 3,584 | -9% | 1 | 1 | 0% | 786 | 744 | -5% | 0 | 0 | — |