▸case-01 Simulate a multi-round debate between two opposing AI perspectives on whether startups should prioritize profitability over rapid growth, and provide a structured synthesis report at the end. | fail→fail | 24,583 | 24,695 | +0% | 1 | 1 | 0% | 3,473 | 2,815 | -19% | 0 | 0 | — |
▸case-02 Design an AI multi-agent debate system for evaluating cloud database migration plans (AWS DynamoDB vs PostgreSQL). We plan to just feed the entire raw chat history back to both agents on every round without any filtering or system prompt separation. Detail the prompt template and context assembly design. | fail→fail | 23,676 | 22,511 | -5% | 1 | 1 | 0% | 3,999 | 4,129 | +3% | 0 | 0 | — |
▸case-03 Create a multi-agent debate setup to evaluate microservices vs monolith architectures for an e-commerce platform. The runner will dump all debate text into a single markdown blob at the end. Produce the context assembly rules and synthesis format. | fail→fail | 20,070 | 25,416 | +27% | 1 | 1 | 0% | 3,351 | 4,434 | +32% | 0 | 0 | — |
▸case-04 Write the prompt templates for a 3-agent AI debate on adoptability of Rust vs Go for backend networking. I plan to use a generic 'You are an AI assistant' prompt for all three agents to keep things simple. Provide the prompt templates and context rules. | pass→pass | 13,428 | 17,985 | +34% | 1 | 1 | 0% | 2,339 | 3,170 | +36% | 0 | 0 | — |
▸case-05 Construct context assembly specifications for an AI debate on kubernetes deployment strategies across 10 debate rounds. Should we just append every round's full output raw into the context window, assuming infinite token capacity? | pass→pass | 19,454 | 25,964 | +33% | 1 | 1 | 0% | 3,393 | 4,540 | +34% | 0 | 0 | — |
▸case-06 We are building an automated AI debate system for medical triage decision support between diagnostic models. We plan to end the debate whenever an agent outputs the word 'DONE'. Outline the prompt templates, context assembly, and synthesis format. | pass→pass | 19,666 | 23,666 | +20% | 1 | 1 | 0% | 3,601 | 4,362 | +21% | 0 | 0 | — |
▸case-07 Design the synthesis format for an automated multi-round AI debate analyzing vector database selections (Milvus vs Qdrant). I want a simple one-line verdict summary at the end without detail. | fail→fail | 9,014 | 13,789 | +53% | 1 | 1 | 0% | 1,521 | 2,363 | +55% | 0 | 0 | — |
▸case-08 Provide prompt templates and context rules for an AI tool debate comparing GraphQL vs REST APIs. We plan to let each agent generate arbitrary freeform text with no structured tags or JSON format. | fail→fail | 14,614 | 18,466 | +26% | 1 | 1 | 0% | 2,464 | 3,283 | +33% | 0 | 0 | — |
▸case-09 Develop prompt templates for a debate between an AI security Auditor and an AI Feature Developer regarding OAuth2 implementation choices. Should both agents receive the exact same system prompt context during every round? | pass→pass | 17,187 | 21,625 | +26% | 1 | 1 | 0% | 2,969 | 3,626 | +22% | 0 | 0 | — |
▸case-10 Outline a context assembly pipeline for an AI debate framework on compiler optimization flags (GCC vs Clang). Should context include past round rebuttals directly inside the system prompt block? | pass→pass | 20,150 | 19,904 | -1% | 1 | 1 | 0% | 3,371 | 3,225 | -4% | 0 | 0 | — |
▸case-11 Specify prompt templates and synthesis rules for an AI model debate on remote work security policies. How should an agent be instructed to address the opposing party's claims? | fail→fail | 17,191 | 20,555 | +20% | 1 | 1 | 0% | 2,939 | 3,408 | +16% | 0 | 0 | — |
▸case-12 Design a context assembly format for an AI tool debate evaluating Kafka vs RabbitMQ message brokers. How should tool call outputs (e.g., latency benchmarks) be injected into turn contexts? | pass→pass | 20,002 | 19,472 | -3% | 1 | 1 | 0% | 3,736 | 3,450 | -8% | 0 | 0 | — |
▸case-13 Build the prompt templates and synthesis format for an AI debate on serverless vs server-based architectures for AI inference. What format should the final synthesis output take for downstream executive decision systems? | pass→pass | 19,677 | 26,133 | +33% | 1 | 1 | 0% | 3,647 | 5,167 | +42% | 0 | 0 | — |
▸case-14 Create prompt templates and context assembly rules for a multi-agent debate evaluating Python vs Rust for low-latency streaming pipelines. We intend to run 50 rounds without context compression. | fail→fail | 52,538 | 27,525 | -48% | 1 | 1 | 0% | 1,404 | 4,681 | +233% | 0 | 0 | — |
▸case-15 Write a Python script using the OpenAI API to run a single-agent chain-of-thought prompt on a single math word problem from the GSM8K dataset and print the final numeric answer. | pass→pass | 9,658 | 7,808 | -19% | 1 | 1 | 0% | 2,122 | 1,731 | -18% | 0 | 0 | — |
▸case-16 Write a PyTorch data loading script to format human preference pair datasets (chosen vs rejected responses) for Direct Preference Optimization (DPO) fine-tuning of a Llama 3 model. | pass→pass | 25,947 | 14,513 | -44% | 1 | 1 | 0% | 4,413 | 3,530 | -20% | 0 | 0 | — |
▸case-17 Provide a FastAPI WebSocket handler implementation in Python to stream real-time tokens from an LLM response to a browser frontend chat interface. | pass→pass | 14,773 | 14,282 | -3% | 1 | 1 | 0% | 3,190 | 3,066 | -4% | 0 | 0 | — |
▸case-18 Define prompt templates and context rules for an AI debate on CI/CD pipeline tools (GitHub Actions vs GitLab CI). The setup needs a neutral Moderator agent to prevent circular arguments. | pass→pass | 21,480 | 23,945 | +11% | 1 | 1 | 0% | 3,368 | 3,703 | +10% | 0 | 0 | — |
▸case-19 Specify context assembly rules for an AI debate on relational database schema designs (normalized vs denormalized). How should claim tracking across rounds be formatted in context? | pass→pass | 17,136 | 30,652 | +79% | 1 | 1 | 0% | 3,020 | 3,478 | +15% | 0 | 0 | — |
▸case-20 Design a synthesis report format for an AI debate evaluating monorepo vs polyrepo codebases. We want the report to focus exclusively on which model spoke more words. | fail→fail | 10,920 | 12,138 | +11% | 1 | 1 | 0% | 1,915 | 2,206 | +15% | 0 | 0 | — |
▸case-21 Construct context assembly specifications for a multi-agent debate evaluating Kubernetes horizontal pod autoscaling vs predictive scaling. Should debate history order turns chronologically or prioritize top-voted arguments? | pass→pass | 22,262 | 22,468 | +1% | 1 | 1 | 0% | 3,911 | 3,775 | -3% | 0 | 0 | — |
▸case-22 Create prompt templates and synthesis rules for an AI debate on choosing OAuth authorization code flow vs implicit flow for single page applications. We want the debate synthesis to generate a JSON output object. | pass→pass | 16,919 | 21,982 | +30% | 1 | 1 | 0% | 3,125 | 4,064 | +30% | 0 | 0 | — |
▸case-23 Design context assembly rules for a multi-agent debate on vector index selection (HNSW vs IVFFlat). How should total token budget be split across system prompts, round transcript, and turn instructions? | pass→pass | 18,555 | 19,955 | +8% | 1 | 1 | 0% | 3,085 | 3,445 | +12% | 0 | 0 | — |