▸case-04 I need to fix a null pointer bug in the authentication middleware file auth/jwt.ts inside a monorepo containing over 400 service files. Should I systematically load and read all 400 service files into context before making the fix? Explain the recommended approach. | pass→pass | 11,234 | 5,962 | -47% | 1 | 1 | 0% | 1,852 | 1,269 | -31% | 0 | 0 | — |
▸case-10 We want an automated agent to refactor 30 database repository files in one continuous, unconstrained execution loop without setting turn limits or phase boundaries. Evaluate this execution approach and suggest improvements. | pass→pass | 15,847 | 14,295 | -10% | 1 | 1 | 0% | 2,572 | 2,466 | -4% | 0 | 0 | — |
▸case-01 I have a massive 50MB crash log file and a dozen related trace files in my project directory. I need you to investigate the root cause of the recent system outage and give me a concise diagnosis report with fix recommendations. Please manage the review process cleanly so we don't overwhelm our chat history. | fail→fail | 10,278 | 2,855 | -72% | 1 | 1 | 0% | 1,990 | 687 | -65% | 0 | 0 | — |
▸case-02 I'm reviewing PR #402, which includes extensive design docs, technical specs, and several UI screenshots. Provide a high-level summary of the architectural changes and potential risks, structuring the output as a brief executive summary with action items. | fail→fail | 10,688 | 4,474 | -58% | 1 | 1 | 0% | 1,518 | 868 | -43% | 0 | 0 | — |
▸case-03 We need to calculate statistical metrics across a 100,000-row CSV telemetry export in our analytics repository. Should we load the full raw CSV dataset into our primary agent context window before running calculations, or handle this differently? Provide a strategy recommendation. | fail→fail | 13,879 | 8,976 | -35% | 1 | 1 | 0% | 2,364 | 1,696 | -28% | 0 | 0 | — |
▸case-05 We are designing a Cursor skill for SQL query formatting that runs on every prompt invocation. We have a standard SQL style checklist schema. Should we store this checklist in a separate workspace file and perform a tool read call on every invocation, or embed it directly inside the skill definition? | pass→pass | 11,637 | 3,827 | -67% | 1 | 1 | 0% | 1,871 | 845 | -55% | 0 | 0 | — |
▸case-06 We are planning an automated agent workflow to migrate 50 microservices from REST to gRPC in a single continuous session. What context management controls should we implement to prevent reasoning degradation during this multi-service task? | fail→fail | 15,604 | 16,180 | +4% | 1 | 1 | 0% | 2,743 | 2,585 | -6% | 0 | 0 | — |
▸case-07 Our visual regression test suite generated 20 high-resolution PNG failure screenshots in the artifacts directory. We need our primary coding assistant to evaluate these visual failures. How should these visual assets be processed during our session? | fail→pass | 11,829 | 8,076 | -32% | 1 | 1 | 0% | 1,809 | 1,422 | -21% | 0 | 0 | — |
▸case-08 We encounter an unhandled rejection in node_modules/express-rate-limit/lib/express-rate-limit.js. Should our debugging process begin by reading every package file inside node_modules into the current chat session? | pass→pass | 9,391 | 4,068 | -57% | 1 | 1 | 0% | 1,572 | 897 | -43% | 0 | 0 | — |
▸case-09 A nightly integration pipeline generated a 200MB text log containing output from 1,200 test suites. We need a summary of broken test suites. Should we upload the entire 200MB log directly into the current chat session? | fail→fail | 10,329 | 5,398 | -48% | 1 | 1 | 0% | 1,880 | 1,028 | -45% | 0 | 0 | — |
▸case-11 Our agent skill validates OpenAPI spec formatting on every generated route. We have a 10-line JSON validation schema used on every single turn. Should we fetch this file from disk using a read tool every turn, or keep it inside the skill prompt? | pass→pass | 7,806 | 3,418 | -56% | 1 | 1 | 0% | 1,331 | 801 | -40% | 0 | 0 | — |
▸case-12 We need to update a single payload field in the POST /v1/billing endpoint. The project repository includes a 1,500-page PDF specification document covering the entire enterprise suite. Should we load the whole PDF into our prompt context? | pass→pass | 9,946 | 3,771 | -62% | 1 | 1 | 0% | 1,711 | 773 | -55% | 0 | 0 | — |
▸case-13 We have three technical whitepapers totaling 500 pages in the docs folder that explain our distributed consensus protocol. We need a summary of consensus timeouts for our current coding session. How should we process these documents? | fail→pass | 12,796 | 5,630 | -56% | 1 | 1 | 0% | 2,187 | 1,143 | -48% | 0 | 0 | — |
▸case-14 We are upgrading our React frontend from Class components to Functional components across 80 component files. Should we open and inspect all 80 component files in a single prompt phase before writing any code? | pass→pass | 13,335 | 7,386 | -45% | 1 | 1 | 0% | 2,081 | 1,308 | -37% | 0 | 0 | — |
▸case-15 Our custom agent tool returns 500KB of raw JSON metadata on every database query tool invocation. How should we handle these large tool responses before returning results to the primary assistant? | pass→pass | 14,371 | 8,040 | -44% | 1 | 1 | 0% | 2,551 | 1,614 | -37% | 0 | 0 | — |
▸case-16 PR #108 changes 5 lines in src/utils/format.ts out of a repository containing 600 total source files. Should we read all 600 files to build context before reviewing the PR? | pass→pass | 7,318 | 2,950 | -60% | 1 | 1 | 0% | 1,239 | 700 | -44% | 0 | 0 | — |
▸case-17 We are configuring an agent skill that formats JSON error messages. The formatting rule is 3 lines of markdown template used on every user turn. Should we require a file-read tool call to fetch this template on each request? | pass→pass | 7,472 | 3,112 | -58% | 1 | 1 | 0% | 1,355 | 752 | -45% | 0 | 0 | — |
▸case-18 We need to analyze 80MB of active server telemetry logs stored in /var/log/telemetry.log to find memory spike events. How should we configure our main agent session to handle this payload? | fail→pass | 13,780 | 7,804 | -43% | 1 | 1 | 0% | 2,489 | 1,424 | -43% | 0 | 0 | — |
▸case-19 An automated linter identified 120 syntax errors across 40 files in our TypeScript project. Should we attempt to resolve all 120 errors in a single long prompt execution without turn limits? | fail→pass | 11,699 | 5,916 | -49% | 1 | 1 | 0% | 1,954 | 1,214 | -38% | 0 | 0 | — |
▸case-20 We are experiencing slow write throughput on our PostgreSQL orders table, which processes 5,000 inserts per second. We need to optimize index creation for order_timestamp and customer_id columns. Should we create a composite B-tree index or a BRIN index for time-series range queries? | pass→pass | 16,907 | 11,997 | -29% | 1 | 1 | 0% | 2,960 | 2,346 | -21% | 0 | 0 | — |
▸case-21 We are choosing between HS256 (symmetric HMAC) and RS256 (asymmetric RSA) for signing JWT tokens in our microservices authentication architecture. Which algorithm should we choose when external third-party services need to verify tokens without access to our signing secret? | pass→pass | 9,578 | 5,250 | -45% | 1 | 1 | 0% | 1,817 | 1,142 | -37% | 0 | 0 | — |
▸case-22 In our React interface, flex items inside a container with display: flex and flex-wrap: wrap are collapsing to zero width instead of growing to fill space equally. How do we configure CSS flex properties so that each child item grows to fill available horizontal space? | pass→pass | 9,957 | 8,122 | -18% | 1 | 1 | 0% | 2,174 | 1,760 | -19% | 0 | 0 | — |