▸case-01 We just modified our customer support agent's system prompt and need to run our regression suite against the updated version. Please evaluate the agent across our standard benchmark test cases, generate individual quality scores, perform a before-and-after regression check, and calculate the final composite pass status. | fail→pass | 10,574 | 15,681 | +48% | 1 | 1 | 0% | 1,807 | 3,706 | +105% | 0 | 0 | — |
▸case-02 Can you perform a full quality evaluation on our newly created data extraction skill? I need you to assess its output correctness, completeness against required fields, and check how consistently it performs across repeated runs, then provide a summary report showing whether it meets our production readiness standard. | fail→pass | 10,589 | 15,185 | +43% | 1 | 1 | 0% | 1,881 | 3,348 | +78% | 0 | 0 | — |
▸case-03 Please run a process verification benchmark for our multi-step workflow agent. Execute the end-to-end task with our known test inputs, check that each execution phase met its expected criteria, score the output dimensions, and give us a consolidated overall score with pass/fail determination. | fail→pass | 17,337 | 15,196 | -12% | 1 | 1 | 0% | 2,740 | 3,338 | +22% | 0 | 0 | — |
▸case-04 We are setting up automated regression testing for our updated document summarizer agent in our CI/CD pipeline. Developers want to know what pass rate percentage across our previously-passing test cases is required for the pipeline build to succeed. Should we pass if 90% of cases succeed? | fail→pass | 14,720 | 4,498 | -69% | 1 | 1 | 0% | 2,198 | 1,136 | -48% | 0 | 0 | — |
▸case-05 We want to measure the consistency of our Python code reviewer agent's output. A team member suggested running the prompt twice and measuring character overlap. How should consistency be measured in our evaluation setup? | fail→fail | 12,854 | 11,536 | -10% | 1 | 1 | 0% | 2,306 | 2,440 | +6% | 0 | 0 | — |
▸case-06 Our quality evaluator needs to score our SQL generator skill across output correctness, required element coverage, and run-to-run repeatability. Should each dimension be weighted equally at 33.3% when calculating the overall grade? | fail→pass | 11,765 | 7,058 | -40% | 1 | 1 | 0% | 1,985 | 1,664 | -16% | 0 | 0 | — |
▸case-07 When evaluating our customer FAQ bot against ground truth answers, one candidate response included the exact correct answer but also invented two fictional refund policy rules. How should accuracy be scored for this response on a 0-100 scale? | fail→fail | 11,974 | 8,145 | -32% | 1 | 1 | 0% | 1,964 | 1,765 | -10% | 0 | 0 | — |
▸case-08 Our quarterly financial report generator agent produced an output that answered all required prompt sections and added a section detailing unexpected macroeconomic risk factors. How should completeness be evaluated on a 0-100 scale? | fail→pass | 10,242 | 6,354 | -38% | 1 | 1 | 0% | 1,779 | 1,474 | -17% | 0 | 0 | — |
▸case-09 An AI contract reviewer agent achieved an accuracy score of 85, a completeness score of 70, and a consistency score of 80. Calculate the composite score and determine whether this agent passes or fails the quality evaluation gate. | fail→pass | 7,433 | 2,891 | -61% | 1 | 1 | 0% | 1,451 | 976 | -33% | 0 | 0 | — |
▸case-10 We are running process verification testing for a 4-stage data ETL pipeline agent. What operational execution parameters must be verified besides checking individual phase outputs? | fail→pass | 16,753 | 12,456 | -26% | 1 | 1 | 0% | 2,767 | 2,497 | -10% | 0 | 0 | — |
▸case-11 We just wrote a complex agent skill for multi-currency accounting. What specific testing steps should be included in our skill quality testing harness beyond standard happy-path inputs? | fail→fail | 20,863 | 17,026 | -18% | 1 | 1 | 0% | 3,328 | 3,245 | -2% | 0 | 0 | — |
▸case-12 Our engineering team is debating when we should trigger our formal evaluation harness for our customer service agent. One developer suggests running it only once per quarter. When should quality evaluation and regression suites be executed? | fail→pass | 14,721 | 11,457 | -22% | 1 | 1 | 0% | 2,417 | 2,394 | -1% | 0 | 0 | — |
▸case-13 When evaluating our clinical trial entity extraction skill, how should the evaluation harness validate extracted medical entities to ensure safety and factual correctness? | fail→fail | 21,400 | 13,815 | -35% | 1 | 1 | 0% | 3,014 | 2,853 | -5% | 0 | 0 | — |
▸case-14 How should we assemble our regression test collection for an evolving document parsing agent, and what specific action should be taken when test results change post-modification? | fail→fail | 16,213 | 12,481 | -23% | 1 | 1 | 0% | 2,645 | 2,438 | -8% | 0 | 0 | — |
▸case-15 A developer wants to score accuracy, completeness, and consistency for our translation agent on a 1 to 5 scale. What specific numerical scale standard should be used for individual quality metrics in our evaluation framework? | fail→pass | 14,713 | 5,173 | -65% | 1 | 1 | 0% | 2,623 | 1,430 | -45% | 0 | 0 | — |
▸case-16 We are setting up benchmark evaluation for our legal translation agent. Beyond calculating single-run quality scores, what tracking mechanism should our evaluation harness implement over time? | fail→fail | 17,217 | 14,668 | -15% | 1 | 1 | 0% | 2,515 | 2,687 | +7% | 0 | 0 | — |
▸case-17 In our technical document summary evaluation, an agent correctly identified 3 out of 4 required key points. Should this response receive zero points for failing exact matching? | fail→fail | 9,869 | 6,960 | -29% | 1 | 1 | 0% | 1,600 | 1,511 | -6% | 0 | 0 | — |
▸case-18 For our multi-agent customer onboarding pipeline, how should process verification be executed during automated evaluation runs? | fail→fail | 20,366 | 14,105 | -31% | 1 | 1 | 0% | 3,451 | 2,727 | -21% | 0 | 0 | — |
▸case-19 When evaluating a generated software architecture document against required specification headers, the agent omitted the Security Requirements section. How does completeness scoring handle this missing content? | fail→fail | 11,495 | 3,777 | -67% | 1 | 1 | 0% | 1,864 | 1,016 | -45% | 0 | 0 | — |
▸case-20 We are designing system prompts for a new technical support chatbot. How should we write system prompt instructions to make the bot sound polite and empathetic when dealing with angry customers? | fail→fail | 14,143 | 15,610 | +10% | 1 | 1 | 0% | 2,343 | 2,937 | +25% | 0 | 0 | — |
▸case-21 We need to configure Prometheus and Grafana alerts for our production agent service to notify our on-call team when API latency spikes or HTTP 500 rate exceeds 1%. How should we configure these production infrastructure metrics? | fail→fail | 15,649 | 12,943 | -17% | 1 | 1 | 0% | 3,302 | 3,096 | -6% | 0 | 0 | — |
▸case-22 We are building a FastAPI backend endpoint that processes user subscriptions. How should we write pytest fixtures and mock database calls to unit test our subscription endpoint logic? | fail→fail | 14,995 | 21,784 | +45% | 1 | 1 | 0% | 3,217 | 3,996 | +24% | 0 | 0 | — |