▸case-01 Could you do a quick quality audit on the answers you generated in our last few turns? Point out any inaccuracies or logical inconsistencies you find. | fail→fail | 2,554 | 3,284 | +29% | 1 | 1 | 0% | 454 | 493 | +9% | 0 | 0 | — |
▸case-02 Review the following previous agent calculation output for quality and accuracy: 'To calculate the monthly payment on a $200,000 loan at 6% annual interest over 30 years, monthly interest rate is 0.5%. Formula is P * r / (1 - (1+r)^-n). Here 200000 * 0.005 / (1 - (1.005)^-360) = $1,199.10. Adding $200 property tax gives $1,499.10 total.' Audit this output for accuracy. | fail→pass | 7,837 | 7,351 | -6% | 1 | 1 | 0% | 1,809 | 1,627 | -10% | 0 | 0 | — |
▸case-03 Audit this recent Python API response snippet generated by an AI assistant: 'To convert JSON to YAML using PyYAML, use `import yaml; output = yaml.dump_json(data)`. This produces valid YAML.' Evaluate this snippet for quality and factual accuracy. | fail→fail | 8,396 | 6,052 | -28% | 1 | 1 | 0% | 1,609 | 1,212 | -25% | 0 | 0 | — |
▸case-04 Perform a consistency audit on this generated summary: 'Section 1 states that all servers must be patched within 24 hours of a critical CVE release. Section 4 states that critical vulnerability patches must undergo a 72-hour staging bake period before deployment.' Evaluate this for output consistency. | pass→pass | 7,617 | 14,209 | +87% | 1 | 1 | 0% | 1,462 | 1,061 | -27% | 0 | 0 | — |
▸case-05 Review the following AI-generated code snippet intended to handle SQL connection cleanup: 'def fetch_users(db_pool):
conn = db_pool.get_connection()
cursor = conn.cursor()
cursor.execute("SELECT * FROM users")
return cursor.fetchall()' Evaluate this snippet for quality and resource safety. | pass→pass | 8,908 | 7,138 | -20% | 1 | 1 | 0% | 1,780 | 1,430 | -20% | 0 | 0 | — |
▸case-06 Audit this generated report text for chronological consistency: 'Company X was founded in 2012 in Austin, Texas. By 2010, the startup had already expanded its operations into three European markets and secured Series A funding.' | pass→pass | 5,825 | 3,777 | -35% | 1 | 1 | 0% | 986 | 738 | -25% | 0 | 0 | — |
▸case-07 Audit the quality and security of this generated backend route handler: 'app.get("/user", (req, res) => { const query = `SELECT * FROM users WHERE id = ${req.query.id}`; db.query(query); });'. Is this output safe and high quality? | pass→pass | 10,410 | 19,536 | +88% | 1 | 1 | 0% | 2,191 | 1,537 | -30% | 0 | 0 | — |
▸case-08 Review this generated unit test suite output: 'def test_divide(): assert divide(10, 2) == 5
assert divide(9, 3) == 3'. Evaluate the accuracy and thoroughness of this agent output. | pass→pass | 9,440 | 9,176 | -3% | 1 | 1 | 0% | 1,898 | 1,902 | +0% | 0 | 0 | — |
▸case-09 Audit the formatting quality of this Markdown output generated by an agent: '| Item | Quantity | Price |
| --- | --- |
| Widget A | 10 | $5.00 |'. Does this output maintain structural quality and consistency? | fail→fail | 7,037 | 3,614 | -49% | 1 | 1 | 0% | 1,186 | 728 | -39% | 0 | 0 | — |
▸case-10 Audit this agent-generated API response documentation: 'When a user submits valid JSON but the specified user ID does not exist in the database, the endpoint returns HTTP 400 Bad Request.' Evaluate this for technical accuracy. | pass→pass | 10,169 | 7,911 | -22% | 1 | 1 | 0% | 2,002 | 1,538 | -23% | 0 | 0 | — |
▸case-11 Audit the accuracy and validity of this generated JSON snippet: '{"name": "Alice", "roles": ["admin", "user",]}'. Evaluate whether this output meets standards. | fail→fail | 5,780 | 5,338 | -8% | 1 | 1 | 0% | 1,119 | 779 | -30% | 0 | 0 | — |
▸case-12 Review this agent's generated Git tutorial steps: 'To undo the last local commit while keeping changes in your working directory, run `git reset --hard HEAD~1`.' Audit this recommendation for quality and accuracy. | pass→pass | 6,309 | 5,762 | -9% | 1 | 1 | 0% | 1,211 | 1,087 | -10% | 0 | 0 | — |
▸case-13 Audit this agent-generated IAM policy statement for quality and security: '{"Effect": "Allow", "Action": "s3:*", "Resource": "*"}'. Is this output appropriate for a read-only service account? | pass→fail | 9,269 | 5,897 | -36% | 1 | 1 | 0% | 1,917 | 1,176 | -39% | 0 | 0 | — |
▸case-14 Audit this agent-generated regular expression for validating email addresses: `^[a-zA-Z0-9]+@[a-zA-Z0-9]+\.[a-zA-Z0-9]+$`. Evaluate its quality and edge-case accuracy. | pass→pass | 14,612 | 11,389 | -22% | 1 | 1 | 0% | 2,541 | 2,302 | -9% | 0 | 0 | — |
▸case-15 Review this TypeScript function generated in a previous step: 'function processAge(age: number): string { return age.toUpperCase(); }'. Evaluate this for quality and type accuracy. | fail→fail | 4,733 | 4,949 | +5% | 1 | 1 | 0% | 963 | 791 | -18% | 0 | 0 | — |
▸case-16 Audit this generated CSS for centering an item horizontally and vertically in a container: `.container { display: flex; align-items: center; }`. Evaluate if this output accomplishes the requested objective. | fail→fail | 8,086 | 5,327 | -34% | 1 | 1 | 0% | 1,289 | 993 | -23% | 0 | 0 | — |
▸case-17 Audit this Python output for quality and consistency: 'def append_to_list(element, target_list=[]): target_list.append(element); return target_list'. Is this clean and error-free code? | pass→pass | 7,898 | 6,329 | -20% | 1 | 1 | 0% | 1,526 | 1,242 | -19% | 0 | 0 | — |
▸case-18 Audit this agent-generated Dockerfile quality: 'FROM node:18
COPY . .
RUN npm install
CMD ["npm", "start"]'. Is this output optimized for caching and efficiency? | pass→pass | 11,955 | 8,731 | -27% | 1 | 1 | 0% | 2,367 | 1,736 | -27% | 0 | 0 | — |
▸case-19 Audit this list of API endpoints generated by an assistant: 'GET /getUsers', 'POST /create_user', 'DELETE /delete-user/id'. Evaluate this for consistency and quality. | pass→pass | 12,479 | 7,327 | -41% | 1 | 1 | 0% | 2,048 | 1,531 | -25% | 0 | 0 | — |
▸case-20 Audit this agent output detailing JWT implementation: 'To verify an incoming JWT in Node.js, decode the base64 payload and check if `payload.exp < Date.now()`. If true, the token is valid.' Is this token validation logic accurate? | pass→fail | 10,548 | 6,840 | -35% | 1 | 1 | 0% | 1,903 | 1,402 | -26% | 0 | 0 | — |
▸case-21 Audit this JavaScript output for quality and runtime accuracy: 'async function getUserData(userId) { const user = fetchUser(userId); return user.name; }'. Evaluate whether this code functions correctly. | fail→fail | 8,811 | 7,351 | -17% | 1 | 1 | 0% | 1,775 | 1,139 | -36% | 0 | 0 | — |
▸case-22 Review this human-authored PR submission for a React component: `function Header({ title }) { return <h1>{title}</h1>; }`. Provide a code review focusing on prop types and React best practices. | pass→pass | 7,732 | 6,541 | -15% | 1 | 1 | 0% | 1,539 | 1,326 | -14% | 0 | 0 | — |
▸case-23 Write a new Python function from scratch that calculates the Fibonacci sequence up to n terms using memoization. | pass→pass | 9,586 | 6,403 | -33% | 1 | 1 | 0% | 1,622 | 1,379 | -15% | 0 | 0 | — |
▸case-24 Based on official PostgreSQL documentation, explain the difference between `VARCHAR(n)` and `TEXT` data types in terms of storage and performance. | pass→pass | 11,539 | 9,800 | -15% | 1 | 1 | 0% | 2,095 | 1,661 | -21% | 0 | 0 | — |