▸case-16 For the C++ game engine module plan, we want to use Batch Execution. Does this mean we group all tasks by file type (e.g., all header files first, then all cpp files)? | pass→pass | 12,699 | 5,769 | -55% | 1 | 1 | 0% | 2,052 | 1,271 | -38% | 0 | 0 | — |
▸case-01 I need to build a user authentication module for my React/Node app and I have the requirements document ready. Before we write any actual features, please generate a comprehensive step-by-step development roadmap. Make sure the document includes a top-level summary of the objective and technologies, detailed file locations with the exact code to be written, testing instructions showing what the console should output, a way to save our progress locally, and a choice on how we should hand off the actual coding work to the AI. | fail→fail | 24,783 | 23,509 | -5% | 1 | 1 | 0% | 6,243 | 6,414 | +3% | 0 | 0 | — |
▸case-02 Here are the specifications for a new Python data ingestion pipeline. I want to map out the entire workflow before we start building. Please draft a blueprint that outlines the main objective and system design at the top. Break the work down into individual assignments containing the full code and exact paths. Also, include the test-driven steps with their anticipated results, a mechanism to track the assignments persistently, and give me the options for how the AI will execute these steps once the blueprint is finalized. | fail→fail | 26,281 | 23,985 | -9% | 1 | 1 | 0% | 6,237 | 5,942 | -5% | 0 | 0 | — |
▸case-03 We are migrating our iOS app's networking layer and I have the full spec sheet. Let's create a detailed action plan prior to doing any real coding. The output needs to establish the overall goal and stack first. Then, define each coding chore with precise file names and complete snippets, along with the testing phases and what the terminal should show. Finally, set up a local file to track our chore status and present the available strategies for delegating the implementation phase. | fail→pass | 24,232 | 25,598 | +6% | 1 | 1 | 0% | 5,343 | 6,165 | +15% | 0 | 0 | — |
▸case-04 I have the specs for a new GraphQL API service. Let's draft the initial planning document before we start coding. Should I output this plan as a standard markdown file with checkboxes, or is there a specific tracking file format I should use? | fail→pass | 9,481 | 4,244 | -55% | 1 | 1 | 0% | 1,635 | 943 | -42% | 0 | 0 | — |
▸case-05 We are planning the implementation of a Redis caching layer based on the provided requirements. Once the plan is ready, I was thinking we could just have one agent execute the whole thing end-to-end. Is that the recommended approach? | pass→pass | 13,027 | 5,767 | -56% | 1 | 1 | 0% | 2,030 | 1,133 | -44% | 0 | 0 | — |
▸case-06 I'm mapping out the tasks for a new Vue.js component library before we begin development. Once the plan is ready, I'll just run `npm start` to trigger the AI coding agent. Is there a specific script I need to invoke instead? | fail→fail | 8,944 | 6,158 | -31% | 1 | 1 | 0% | 1,608 | 1,238 | -23% | 0 | 0 | — |
▸case-07 We are creating a roadmap for a Django backend overhaul. We need a way to keep track of which steps are done across different sessions. Should we use a standard `.env` file or a `.progress.yaml` file for this? | fail→fail | 10,299 | 6,928 | -33% | 1 | 1 | 0% | 1,923 | 1,314 | -32% | 0 | 0 | — |
▸case-08 The planning phase for the microservices split is complete. We need to decide how to execute the plan. One option involves using individual workers for each piece. What is this specific execution strategy called and how does it handle reviews? | fail→pass | 13,207 | 4,213 | -68% | 1 | 1 | 0% | 2,020 | 856 | -58% | 0 | 0 | — |
▸case-09 We've finished writing the blueprint for the new payment gateway integration. We want to execute multiple steps at once rather than one by one. What is this execution mode called and how do we ensure quality control during it? | fail→pass | 12,810 | 3,639 | -72% | 1 | 1 | 0% | 2,018 | 796 | -61% | 0 | 0 | — |
▸case-10 I'm setting up the environment to generate a plan for our Flutter app redesign. Which specific file defines the process workers that will be utilized for this planning phase? | fail→pass | 9,582 | 2,326 | -76% | 1 | 1 | 0% | 1,680 | 475 | -72% | 0 | 0 | — |
▸case-11 We have the requirements for the new analytics dashboard and need to trigger the planning process. What is the exact path to the babysitter process that we need to invoke to start this? | fail→pass | 15,256 | 2,009 | -87% | 1 | 1 | 0% | 687 | 477 | -31% | 0 | 0 | — |
▸case-12 I'm drafting the task sequence for a Rust CLI tool. I need to configure the AI workers that will execute this plan. Should I define the process workers in the `package.json`? | fail→pass | 11,938 | 5,923 | -50% | 1 | 1 | 0% | 2,074 | 1,245 | -40% | 0 | 0 | — |
▸case-13 For the Go microservice plan, we will use the Subagent-Driven handoff strategy. Does this strategy allow the subagent to merge code directly to main without any oversight? | fail→pass | 7,285 | 3,793 | -48% | 1 | 1 | 0% | 1,230 | 809 | -34% | 0 | 0 | — |
▸case-14 In the roadmap for the Angular frontend update, we are using the Batch Execution strategy. Since it's batched, does the AI just run through all 50 tasks and only show me the final result at the very end? | pass→pass | 9,541 | 6,152 | -36% | 1 | 1 | 0% | 1,624 | 1,224 | -25% | 0 | 0 | — |
▸case-15 We are outlining the steps for a Ruby on Rails API endpoint and plan to use the Subagent-Driven strategy. To save on token costs, does this strategy reuse the exact same agent instance for every single task in the plan? | pass→pass | 9,804 | 3,293 | -66% | 1 | 1 | 0% | 1,649 | 740 | -55% | 0 | 0 | — |
▸case-17 We are choosing the Subagent-Driven handoff method for our Next.js project plan. Does this method reuse the same worker for all tasks to save context, or how does it assign workers to tasks? | fail→pass | 12,057 | 3,614 | -70% | 1 | 1 | 0% | 1,977 | 732 | -63% | 0 | 0 | — |
▸case-18 We opted for the Batch Execution strategy for implementing the Docker orchestration plan. Does this mean we run the entire project end-to-end without stopping, or how is the work grouped? | pass→pass | 9,871 | 5,518 | -44% | 1 | 1 | 0% | 1,765 | 1,103 | -38% | 0 | 0 | — |
▸case-19 I have the specs for a new Spring Boot service. Before coding, generate a plan outline. For the execution handoff, should we just assign the whole plan to a single long-running agent session? | pass→pass | 13,361 | 7,397 | -45% | 1 | 1 | 0% | 2,355 | 1,642 | -30% | 0 | 0 | — |
▸case-20 I have a bug report for a null pointer exception in the user profile view. It's a one-line fix in `Profile.js`. Please provide the steps to fix this. | fail→fail | 7,806 | 10,612 | +36% | 1 | 1 | 0% | 1,478 | 2,198 | +49% | 0 | 0 | — |
▸case-21 We are halfway through implementing the shopping cart feature. I've written the frontend components but need help connecting them to the existing Redux store. How should I wire up the dispatch calls? | pass→pass | 11,417 | 8,660 | -24% | 1 | 1 | 0% | 2,387 | 1,840 | -23% | 0 | 0 | — |
▸case-22 My colleague already wrote a project plan for the database migration and saved it as `migration_plan.md`. Can you review it and tell me if the timeline looks realistic? | fail→pass | 4,159 | 1,887 | -55% | 1 | 1 | 0% | 708 | 409 | -42% | 0 | 0 | — |
▸case-23 I am putting together the task list for a new WebGL rendering pipeline. I will create a `.tasks.yaml` file to track the progress of each step. Is YAML the correct format for this? | fail→pass | 11,920 | 4,651 | -61% | 1 | 1 | 0% | 2,030 | 991 | -51% | 0 | 0 | — |