▸case-01 I have a pull request for our payment gateway module that touches microservice communication, database locking, and token validation. Could you run a multi-angle code review on this pull request using separate specialist roles for security, database performance, and API design? Please provide a unified final output that includes an overview of each expert's findings, combined recommendations, and a prioritized checklist of fixes needed before deployment. | pass→pass | 22,522 | 48,039 | +113% | 1 | 1 | 0% | 3,606 | 2,863 | -21% | 0 | 0 | — |
▸case-02 Our authentication service is undergoing an architectural and security assessment. Can you coordinate concurrent expert evaluations covering threat modeling, structural design, and automated test coverage? Please format the final output to highlight the task summary, individual specialist insights, combined recommendations, and immediate remediation steps. | pass→pass | 33,849 | 15,886 | -53% | 1 | 1 | 0% | 5,326 | 2,931 | -45% | 0 | 0 | — |
▸case-03 We have a missing semicolon and broken import on line 12 of `utils/logger.js`. Can you fix this syntax error in the file? | fail→fail | 10,134 | 3,294 | -67% | 1 | 1 | 0% | 1,637 | 841 | -49% | 0 | 0 | — |
▸case-04 Can you explain how `git rebase -i` works when squashing two commits? | pass→pass | 14,422 | 13,032 | -10% | 1 | 1 | 0% | 2,520 | 2,848 | +13% | 0 | 0 | — |
▸case-05 Please change the `MAX_CONNECTIONS` parameter in `database.yaml` from 50 to 100. | fail→fail | 2,622 | 2,525 | -4% | 1 | 1 | 0% | 422 | 695 | +65% | 0 | 0 | — |
▸case-06 We are preparing our legacy monolithic e-commerce application in `src/legacy` for a cloud migration. We need a comprehensive review evaluating architecture, security vulnerabilities, and database query efficiency. How should this multi-domain assessment be executed across expert domains? | fail→pass | 23,503 | 20,761 | -12% | 1 | 1 | 0% | 3,827 | 2,075 | -46% | 0 | 0 | — |
▸case-07 We built a new file upload feature touching Node.js backend controllers, React UI upload widgets, and S3 bucket storage policies. We want a full feature review across backend, frontend, and cloud infrastructure. | fail→pass | 21,910 | 13,502 | -38% | 1 | 1 | 0% | 3,692 | 2,707 | -27% | 0 | 0 | — |
▸case-08 We need an in-depth security audit of our OAuth2 authentication service, focusing on token handling, session management, and RBAC policy enforcement. How should this security evaluation be structured using specialist roles? | fail→pass | 22,901 | 16,110 | -30% | 1 | 1 | 0% | 3,843 | 3,118 | -19% | 0 | 0 | — |
▸case-09 We need to refactor our database layer in PostgreSQL first, then update our REST API endpoints to consume the new queries, and finally update the OpenAPI specification. The API work depends directly on the database schema output. How should this multi-step task sequence be handled? | fail→pass | 15,949 | 8,448 | -47% | 1 | 1 | 0% | 2,291 | 1,829 | -20% | 0 | 0 | — |
▸case-10 An initial architecture agent analyzed our GraphQL schema and produced structural findings. Now a performance agent needs to evaluate query complexity based on those exact architectural findings. How should data pass from the first analysis to the second? | pass→pass | 14,488 | 7,989 | -45% | 1 | 1 | 0% | 2,576 | 1,754 | -32% | 0 | 0 | — |
▸case-11 A previous agent session was performing a security scan on `auth/jwt.py` and halted after finding token expiration issues. We need to continue analyzing that same context to evaluate refresh token rotation without starting over from scratch. | fail→pass | 16,144 | 4,640 | -71% | 1 | 1 | 0% | 2,235 | 1,148 | -49% | 0 | 0 | — |
▸case-12 We are running a multi-agent review of our GraphQL microservice covering schema safety and query performance. What high-level introductory section must be present at the top of the combined report to outline the objective and scope? | pass→fail | 12,584 | 3,346 | -73% | 1 | 1 | 0% | 2,025 | 1,024 | -49% | 0 | 0 | — |
▸case-13 When collecting results from individual security, performance, and frontend agents evaluating a web portal, how should each individual specialist's raw findings be presented in the final deliverable? | pass→pass | 17,044 | 9,179 | -46% | 1 | 1 | 0% | 2,759 | 1,897 | -31% | 0 | 0 | — |
▸case-14 After individual sub-agents finish reviewing backend API design, database indexing, and client caching for our search feature, how should overlapping architectural guidance be presented? | pass→pass | 14,165 | 11,728 | -17% | 1 | 1 | 0% | 2,213 | 2,278 | +3% | 0 | 0 | — |
▸case-15 Our multi-expert evaluation of the checkout pipeline generated dozens of observations across security, latency, and test coverage. How should executable next steps be organized at the conclusion of the report? | pass→pass | 13,099 | 10,282 | -22% | 1 | 1 | 0% | 2,104 | 2,067 | -2% | 0 | 0 | — |
▸case-16 We have two tasks: Task A is adding a missing `import os` line in `script.py`. Task B is evaluating a new payment checkout system across PCI-DSS compliance, database transaction isolation, and React UI state management. Which task requires multi-agent orchestration? | pass→pass | 4,798 | 3,341 | -30% | 1 | 1 | 0% | 762 | 1,060 | +39% | 0 | 0 | — |
▸case-17 We need to run a standalone, isolated code quality scan on a single Golang utility file `pkg/str/utils.go` without needing any other expertise domains. How should this isolated agent invocation be configured? | pass→pass | 14,691 | 5,007 | -66% | 1 | 1 | 0% | 2,145 | 1,241 | -42% | 0 | 0 | — |
▸case-18 We are implementing a user notification system spanning PostgreSQL migration scripts, a NestJS webhooks service, and a Vue.js frontend settings page. How should these three component domains be coordinated during review? | fail→pass | 19,136 | 10,481 | -45% | 1 | 1 | 0% | 3,029 | 2,090 | -31% | 0 | 0 | — |
▸case-19 Our REST API endpoint `/api/v1/orders` needs review from security, performance, and code quality perspectives simultaneously. What pattern should be used for this review? | fail→pass | 17,634 | 5,668 | -68% | 1 | 1 | 0% | 2,799 | 1,357 | -52% | 0 | 0 | — |
▸case-20 Our platform team requires a comprehensive review of our new Kubernetes deployment manifests covering infrastructure architecture, secret security, and integration testing strategy. How should this review be orchestrated? | fail→pass | 20,848 | 10,855 | -48% | 1 | 1 | 0% | 3,007 | 2,229 | -26% | 0 | 0 | — |
▸case-21 We need to change a static landing page headline string from 'Welcome' to 'Welcome Back' in `Header.tsx`. Should we launch specialized security, performance, and architecture sub-agents for this edit? | pass→pass | 6,650 | 2,921 | -56% | 1 | 1 | 0% | 1,036 | 899 | -13% | 0 | 0 | — |
▸case-22 We are executing a full multi-agent review of our core billing microservice. What are the essential structural components that must compose the final consolidated synthesis report? | fail→pass | 15,957 | 9,028 | -43% | 1 | 1 | 0% | 2,556 | 1,691 | -34% | 0 | 0 | — |