▸case-01 We are launching a complex customer portal with auth, billing, and analytics services. The scope is too large for a single agent run, so please break this goal down into discrete, parallelizable subtasks. Map out the dependencies between each work item, estimate the required effort, assign priorities, and classify which items represent persistent core deliverables versus temporary setup scripts. | fail→pass | 22,570 | 14,726 | -35% | 1 | 1 | 0% | 3,847 | 3,452 | -10% | 0 | 0 | — |
▸case-02 I need to orchestrate a major refactoring of our payment processing pipeline using multiple concurrent worker agents. Please analyze the architecture project and divide it into clear, single-agent work units. Include a dependency structure showing what blocks what, assign priority levels, and distinguish throwaway exploratory tasks from permanent code additions. | fail→pass | 23,188 | 20,643 | -11% | 1 | 1 | 0% | 4,106 | 4,381 | +7% | 0 | 0 | — |
▸case-03 Our team is starting a massive legacy system migration to cloud microservices. Help us decompose this objective into manageable, trackable work modules that can be handled independently. Provide a mapped sequence of dependencies, priority and effort estimates for each package, and identify which tasks create lasting system artifacts versus ephemeral scaffolding. | fail→fail | 19,739 | 14,599 | -26% | 1 | 1 | 0% | 3,866 | 2,798 | -28% | 0 | 0 | — |
▸case-04 We are breaking down a multi-service data pipeline setup. The team wants to allow feedback loops where downstream tasks block upstream tasks if testing fails. How should dependencies between decomposed work units be structured? | fail→fail | 13,839 | 11,374 | -18% | 1 | 1 | 0% | 2,419 | 2,385 | -1% | 0 | 0 | — |
▸case-05 When splitting a large database refactoring goal, should we create two large multi-agent persistent task packages or several smaller single-agent persistent items? We want to minimize management overhead by keeping packages large. | fail→fail | 16,069 | 6,059 | -62% | 1 | 1 | 0% | 2,459 | 1,181 | -52% | 0 | 0 | — |
▸case-06 We have a work item that spans both front-end UI components and back-end database schema creation. Can we assign this single work unit to two concurrent agents to work on simultaneously? | fail→fail | 10,385 | 4,548 | -56% | 1 | 1 | 0% | 1,620 | 1,016 | -37% | 0 | 0 | — |
▸case-07 When generating task definitions for a distributed agent swarm building a CLI tool, is attribution tracking optional for minor setup tasks to reduce metadata clutter? | fail→pass | 12,401 | 4,221 | -66% | 1 | 1 | 0% | 1,987 | 863 | -57% | 0 | 0 | — |
▸case-08 We need to create a temporary test runner script and an exploratory benchmark script before writing the production caching layer. How should these temporary items be classified versus the production caching code? | fail→pass | 12,099 | 6,679 | -45% | 1 | 1 | 0% | 2,062 | 1,344 | -35% | 0 | 0 | — |
▸case-09 Which specific tool and step should be invoked via the babysitter process to run the automated work analysis phase for project decomposition? | fail→pass | 6,827 | 1,412 | -79% | 1 | 1 | 0% | 1,279 | 439 | -66% | 0 | 0 | — |
▸case-10 We are outlining a process for decomposing a large feature request. Should we estimate effort and assign priorities before or after identifying natural seams and creating bounded work items? | fail→fail | 10,553 | 6,825 | -35% | 1 | 1 | 0% | 1,745 | 1,406 | -19% | 0 | 0 | — |
▸case-11 A team lead proposes labeling persistent database models, core API endpoints, and user documentation as 'permanent tickets'. What specific classification term should be used instead? | fail→pass | 9,589 | 2,522 | -74% | 1 | 1 | 0% | 1,599 | 625 | -61% | 0 | 0 | — |
▸case-12 In our proposed task decomposition graph, Task A generates test data for Task B, and Task B updates the schema definition that Task A depends on. Is this circular blocking relationship acceptable? | fail→fail | 9,587 | 4,709 | -51% | 1 | 1 | 0% | 1,627 | 991 | -39% | 0 | 0 | — |
▸case-13 During initial project spike work, an engineer creates a throwaway prototype script to test API throughput. What specific term should be assigned to this throwaway task? | fail→pass | 6,531 | 1,956 | -70% | 1 | 1 | 0% | 1,161 | 511 | -56% | 0 | 0 | — |
▸case-14 An engineer drafted a task called 'Build Full E-commerce Stack' requiring 5 concurrent agents working across 3 repositories. How should this task definition be adjusted? | fail→fail | 11,480 | 9,100 | -21% | 1 | 1 | 0% | 1,878 | 1,876 | -0% | 0 | 0 | — |
▸case-15 When applying work decomposition to a monolithic app rewrite, what step immediately follows analyzing the goal and project context? | fail→pass | 5,313 | 1,921 | -64% | 1 | 1 | 0% | 845 | 491 | -42% | 0 | 0 | — |
▸case-16 An engineer suggests that wisps do not need attribution tracking because they are ephemeral throwaway work. Is this suggestion compliant with decomposition principles? | fail→fail | 13,648 | 2,839 | -79% | 1 | 1 | 0% | 2,009 | 705 | -65% | 0 | 0 | — |
▸case-17 In an automated deployment pipeline breakdown, we have persistent database migration files and ephemeral build container cleanup scripts. Categorize both according to the decomposition framework. | fail→pass | 9,201 | 2,890 | -69% | 1 | 1 | 0% | 1,580 | 666 | -58% | 0 | 0 | — |
▸case-18 What is the full process invocation string required to trigger the work decomposition analysis step in the babysitter framework? | fail→pass | 11,207 | 3,478 | -69% | 1 | 1 | 0% | 1,566 | 644 | -59% | 0 | 0 | — |
▸case-19 We have an objective that exceeds what a single agent can complete in a single execution context. What strategy and task structures should be employed? | fail→pass | 15,439 | 7,400 | -52% | 1 | 1 | 0% | 2,487 | 1,520 | -39% | 0 | 0 | — |
▸case-20 We are defining work items for an infrastructure project. Since cloud infrastructure can be destroyed and redeployed at any time, an engineer proposes classifying the persistent infrastructure-as-code modules as wisps. Is this classification correct? | fail→pass | 11,883 | 7,465 | -37% | 1 | 1 | 0% | 1,761 | 1,385 | -21% | 0 | 0 | — |
▸case-21 During a backend refactoring project, a team lead wants to assign a single high-priority MEOW simultaneously to three concurrent worker agents to complete it faster. Is assigning one MEOW to multiple concurrent agents permitted? | pass→pass | 9,501 | 4,308 | -55% | 1 | 1 | 0% | 1,572 | 907 | -42% | 0 | 0 | — |
▸case-22 When creating ephemeral wisp tasks for single-use exploratory scripts, a developer suggests skipping attribution tracking to reduce file header clutter. Is skipping attribution tracking allowed for ephemeral wisps? | fail→pass | 11,365 | 2,796 | -75% | 1 | 1 | 0% | 1,741 | 672 | -61% | 0 | 0 | — |