▸case-01 We need to plan a new integration feature for Stripe webhook handling in our Express application. Please research how to securely verify webhook signatures using standard libraries and recent 2024-2025 best practices. Scrape official doc pages if needed, query library details and common security patterns, then compile all your findings into a markdown handoff document inside the `./handoffs` directory and summarize your results back to me. | fail→fail | 35,548 | 17,502 | -51% | 1 | 1 | 0% | 8,404 | 1,229 | -85% | 0 | 0 | — |
▸case-02 I am preparing to refactor our frontend state management to use React Server Components alongside React hooks. Can you conduct research on current best practices, query library documentation, collect code snippets and potential pitfalls, save the complete handoff file under `./handoffs/`, and reply with a brief summary report so our planning team can take over? | fail→fail | 18,740 | 14,975 | -20% | 1 | 1 | 0% | 2,426 | 1,153 | -52% | 0 | 0 | — |
▸case-03 Our backend service needs an update to support OAuth2 authorization code flow with PKCE in FastAPI. Please research standard implementation patterns, search for best practice guidelines, gather API documentation for Python OAuth tools, write up the findings with code examples and pitfalls into a handoff markdown report in the `./docs/handoffs` directory, and present a final completion message. | fail→fail | 11,159 | 6,074 | -46% | 1 | 1 | 0% | 249 | 1,236 | +396% | 0 | 0 | — |
▸case-04 We already completed the research phase for adding OAuth2 PKCE to our FastAPI service and have the research findings ready. Please draft the step-by-step technical implementation plan file plan-01-oauth-pkce.md detailing the file modifications, database schema updates, and execution phases for our engineering team. | pass→fail | 54,359 | 8,060 | -85% | 1 | 1 | 0% | 7,713 | 1,427 | -81% | 0 | 0 | — |
▸case-23 We are researching how to implement rate limiting middleware in Express using express-rate-limit. Rather than describing the code logic purely in prose text, how should code patterns found during research be presented in the handoff markdown document? | pass→pass | 20,014 | 15,727 | -21% | 1 | 1 | 0% | 2,662 | 2,592 | -3% | 0 | 0 | — |
▸case-05 We have finished reviewing best practices for React Server Components state management. Please implement the React component state code directly in src/components/UserProfile.tsx with full TypeScript types and hook calls. | fail→fail | 21,539 | 28,669 | +33% | 1 | 1 | 0% | 3,929 | 1,189 | -70% | 0 | 0 | — |
▸case-06 Our research on Stripe webhook verification is complete. Please write PyTest unit tests in tests/test_webhooks.py to verify our signature checking logic against invalid and valid signatures. | pass→fail | 39,439 | 34,775 | -12% | 1 | 1 | 0% | 3,095 | 6,488 | +110% | 0 | 0 | — |
▸case-07 We need to look up API documentation and usage patterns for the Zustand state management library in React. You might be tempted to run a general web search query or web page scraper. Which specific research script and harness invocation should be used specifically for fetching library documentation? | fail→pass | 21,311 | 8,183 | -62% | 1 | 1 | 0% | 2,507 | 1,471 | -41% | 0 | 0 | — |
▸case-08 We are researching general best practices and architectural patterns for rate limiting in Node.js microservices. Rather than scraping a single specific web page or querying a single library's API docs, which research script and mode option should be invoked to gather general best practices? | fail→pass | 21,365 | 8,194 | -62% | 1 | 1 | 0% | 3,343 | 1,423 | -57% | 0 | 0 | — |
▸case-09 We have a specific documentation URL https://docs.celeryq.dev/en/stable/userguide/tasks.html that contains exact configuration details for Celery task queues in Python. You might be tempted to perform a search query. Which research script should be called to extract content directly from this specific web page? | fail→pass | 7,595 | 2,611 | -66% | 1 | 1 | 0% | 413 | 1,364 | +230% | 0 | 0 | — |
▸case-10 We need to research how to implement JWT authentication in FastAPI using PyJWT. You might be inclined to run python scripts/nia_docs.py directly using standard python. Show the precise Bash command prefix required to run research scripts in this environment. | fail→pass | 19,424 | 3,367 | -83% | 1 | 1 | 0% | 333 | 1,443 | +333% | 0 | 0 | — |
▸case-11 We are wrapping up research on Redis caching strategies for Postgres queries. We need to save the findings in ./handoffs. Instead of naming the file redis-caching.md or notes.md, what exact filename convention must be used for saving research handoffs? | fail→pass | 12,015 | 2,013 | -83% | 1 | 1 | 0% | 1,251 | 1,244 | -1% | 0 | 0 | — |
▸case-12 We are compiling research on GraphQL subscription scaling in Apollo Server. Rather than using standard markdown headers or JSON block headers, what metadata section structure must be included at the very top of the research handoff markdown file? | fail→fail | 4,603 | 7,795 | +69% | 1 | 1 | 0% | 766 | 1,407 | +84% | 0 | 0 | — |
▸case-13 We need to research background context for migrating our database ORM to Prisma in a Node.js TypeScript project. Instead of creating a custom free-form summary file, save the research handoff in ./handoffs/research-01-prisma-migration.md. Make sure the markdown document contains all required handoff sections for research question, findings, code examples, recommendations, pitfalls, and agent handoff. | fail→fail | 18,572 | 10,520 | -43% | 1 | 1 | 0% | 979 | 1,220 | +25% | 0 | 0 | — |
▸case-14 We need to analyze micro-frontend architectures in Next.js 14. We want to synthesize findings from API docs, general community patterns, and specific vendor documentation pages into a single handoff document rather than separate notes files. How should findings from multiple sources be structured in the handoff file? | pass→fail | 23,811 | 15,646 | -34% | 1 | 1 | 0% | 3,192 | 3,131 | -2% | 0 | 0 | — |
▸case-15 We are researching building a real-time collaborative editor with Yjs and WebSockets in 2025. You might be tempted to just run a single web search script. Instead, use the specialized harness scripts to query Yjs library documentation, perform general best practice research for WebSockets, and scrape the official Yjs website docs, then compile the findings. | fail→fail | 34,276 | 15,655 | -54% | 1 | 1 | 0% | 5,951 | 1,158 | -81% | 0 | 0 | — |
▸case-16 While researching gRPC streaming in Go, the Nia docs script failed due to an API timeout, but Perplexity search succeeded. Instead of aborting the process or failing the task completely, how should the agent proceed with generating the handoff? | pass→pass | 13,445 | 7,474 | -44% | 1 | 1 | 0% | 2,066 | 2,105 | +2% | 0 | 0 | — |
▸case-17 During research on Tailwind CSS v4 setup, the Firecrawl scraping script failed due to a missing API key, but library docs and best practice search completed successfully. What exact value should be set in the frontmatter status field of the handoff markdown document? | fail→fail | 15,576 | 2,061 | -87% | 1 | 1 | 0% | 2,495 | 1,286 | -48% | 0 | 0 | — |
▸case-18 We completed research on Docker multi-stage builds for Rust applications and saved the handoff to ./handoffs/research-02-docker-rust.md. Instead of returning a plain chat message or unstructured notes, what exact response format must be returned to the calling process? | fail→fail | 15,950 | 7,884 | -51% | 1 | 1 | 0% | 1,826 | 1,464 | -20% | 0 | 0 | — |
▸case-19 We are finishing a research task on Kubernetes operator development in Go. Which specific agent should be explicitly referenced as ready to continue in the final line of the completion output? | fail→fail | 13,411 | 6,545 | -51% | 1 | 1 | 0% | 1,323 | 1,167 | -12% | 0 | 0 | — |
▸case-20 We need to research current best practices for state management in Next.js App Router applications. When conducting this research, what timeframe should be used as the reference period for contemporary best practices? | fail→fail | 11,747 | 10,064 | -14% | 1 | 1 | 0% | 1,699 | 1,705 | +0% | 0 | 0 | — |
▸case-21 We are researching standard Redis caching patterns in Go for high-throughput APIs. You might be inclined to present only positive recommendations. Ensure both actionable guidelines and common anti-patterns or mistakes are explicitly captured in dedicated sections in the output handoff document. | fail→fail | 48,851 | 12,849 | -74% | 1 | 1 | 0% | 8,238 | 1,249 | -85% | 0 | 0 | — |
▸case-22 When querying documentation for a newly released Python library, the research tools returned only basic setup steps and no details on advanced batching parameters. Should you speculate or generate plausible batching configuration flags to fill in the gaps in the handoff document? | pass→pass | 10,031 | 10,240 | +2% | 1 | 1 | 0% | 1,584 | 1,652 | +4% | 0 | 0 | — |