▸case-13 I am attaching a built-in search tool and a built-in code execution tool alongside custom Python tools to a single agent runner instance, but the framework throws an incompatibility error. What tool constraint explains this behavior? | pass→pass | 13,345 | 5,057 | -62% | 1 | 1 | 0% | 2,104 | 1,341 | -36% | 0 | 0 | — |
▸case-01 I'm building a customer support bot with the Google Agent Development Kit. Right now, I am trying to store generated PDF receipts directly inside the session state dictionary so the user can download them later. Can you review this approach and explain the proper way to handle session data versus heavy binary files in ADK? | fail→pass | 13,132 | 9,234 | -30% | 1 | 1 | 0% | 2,392 | 2,163 | -10% | 0 | 0 | — |
▸case-02 My Google ADK backend relies on a single master agent that handles query parsing, backend API execution, PDF report generation, and email delivery all at once. How should I restructure this setup following official ADK architectural guidelines? | fail→pass | 17,528 | 15,387 | -12% | 1 | 1 | 0% | 2,960 | 3,258 | +10% | 0 | 0 | — |
▸case-03 I am designing a financial transaction approval pipeline in an agentic framework. Step 1 evaluates risk scoring with a prompt, but Step 2 must strictly execute sequential compliance checks and database commits without any LLM non-determinism or step skipping. Should I use a single LLM agent with instructions to follow steps 1 and 2, or mix LLM and non-LLM workflow agents? | pass→pass | 19,504 | 6,969 | -64% | 1 | 1 | 0% | 1,941 | 1,744 | -10% | 0 | 0 | — |
▸case-04 I am writing a diagnostic tool for a server management agent. The tool function takes a string command parameter directly from the model prompt and executes shell subprocess calls with it. Is this tool design acceptable if I instruct the LLM in the system prompt to only output safe diagnostic commands? | pass→pass | 11,672 | 8,747 | -25% | 1 | 1 | 0% | 1,977 | 2,103 | +6% | 0 | 0 | — |
▸case-05 To let an agent query an internal API, I am adding the API authorization bearer token into the agent's system instruction text so the LLM can construct HTTP request headers when invoking tools. What issue does this present and how should credentials be managed? | pass→pass | 14,276 | 8,099 | -43% | 1 | 1 | 0% | 2,396 | 1,923 | -20% | 0 | 0 | — |
▸case-06 In an e-commerce agent, I want the agent to remember user preferences across multiple separate chat sessions over several weeks. Should I store this data in the active session state or somewhere else? | pass→pass | 11,651 | 4,936 | -58% | 1 | 1 | 0% | 2,029 | 1,476 | -27% | 0 | 0 | — |
▸case-07 My agent relies on generating audio summaries and chart images using artifact operations. However, during execution, the artifact creation fails with unconfigured service errors. What prerequisite runner setup is required before performing artifact operations? | pass→pass | 12,375 | 7,401 | -40% | 1 | 1 | 0% | 2,006 | 1,747 | -13% | 0 | 0 | — |
▸case-08 An agent periodically generates an updated CSV financial report and saves it using a static filename latest_report.csv every hour. What problem can arise with this approach according to artifact management practices and how should filenames be handled? | pass→pass | 10,388 | 4,907 | -53% | 1 | 1 | 0% | 1,913 | 1,363 | -29% | 0 | 0 | — |
▸case-09 We are preparing a test suite for a multi-agent customer service system before releasing to production. Beyond basic model output quality benchmarks, what agent system components should be explicitly tested? | pass→pass | 17,175 | 10,523 | -39% | 1 | 1 | 0% | 2,719 | 2,137 | -21% | 0 | 0 | — |
▸case-10 We are deploying an agent application to production and need to set up monitoring dashboards. What specific metrics and failure modes should be tracked for multi-agent applications? | pass→pass | 17,574 | 16,088 | -8% | 1 | 1 | 0% | 2,886 | 3,134 | +9% | 0 | 0 | — |
▸case-11 Our team is storing production database connections, staging API endpoints, and local mock server URLs in a single hardcoded configuration dictionary inside the agent initialization script. How should environment settings be structured? | pass→pass | 12,195 | 9,053 | -26% | 1 | 1 | 0% | 2,131 | 1,967 | -8% | 0 | 0 | — |
▸case-12 When a backend tool in our agent framework encounters an unexpected HTTP 500 error, it catches the exception and returns generic false boolean values or raises generic runtime errors. How should tool errors be handled so the calling agent can recover? | pass→pass | 15,600 | 13,669 | -12% | 1 | 1 | 0% | 2,756 | 3,056 | +11% | 0 | 0 | — |
▸case-14 An agent unexpectedly routed a user request to a refund agent instead of the technical support agent. The raw completion text does not reveal why the decision was made. What debugging mechanism should be enabled to investigate agent decisions? | pass→pass | 9,043 | 3,045 | -66% | 1 | 1 | 0% | 1,439 | 1,012 | -30% | 0 | 0 | — |
▸case-15 A developer is storing complex custom class instances with open network handles and dynamically generated keys directly inside the conversation session dictionary. What principles should govern state key naming and data types? | pass→pass | 15,232 | 7,533 | -51% | 1 | 1 | 0% | 2,392 | 1,845 | -23% | 0 | 0 | — |
▸case-16 An agent has a tool named process_data() that performs data formatting, sends emails, and deletes expired database records depending on an internal boolean flag. Why is this tool design problematic for agent reasoning? | pass→pass | 11,900 | 8,569 | -28% | 1 | 1 | 0% | 1,945 | 1,901 | -2% | 0 | 0 | — |
▸case-17 To prevent non-admin users from triggering administrative database wipes, we added instructions in the system prompt telling the agent to check user roles before running destructive actions. Is this sufficient for security enforcement? | pass→pass | 11,462 | 6,245 | -46% | 1 | 1 | 0% | 1,872 | 1,509 | -19% | 0 | 0 | — |
▸case-18 When writing system instructions for a technical support agent, what key operational rules must be defined in the prompt to ensure predictable behavior when handling out-of-scope requests? | pass→pass | 16,167 | 11,234 | -31% | 1 | 1 | 0% | 2,557 | 2,381 | -7% | 0 | 0 | — |
▸case-19 We have two microservices written in Node.js and Python. Our team wants to split the agent system into a NodeAgent and a PythonAgent strictly because the code is in two different microservices. Is this the recommended way to divide multi-agent systems? | pass→pass | 13,841 | 7,595 | -45% | 1 | 1 | 0% | 2,360 | 1,894 | -20% | 0 | 0 | — |
▸case-20 Our agent code hardcodes model string references directly across ten different agent class constructors. What design pattern should be applied to model selection across agents? | pass→pass | 15,088 | 9,350 | -38% | 1 | 1 | 0% | 2,299 | 1,668 | -27% | 0 | 0 | — |
▸case-21 I need to deploy a Python web application container to Google Cloud Run using the gcloud CLI. Can you provide a Dockerfile using python:3.11-slim and the gcloud run deploy command syntax with unauthenticated access enabled? | pass→pass | 6,875 | 12,142 | +77% | 1 | 1 | 0% | 1,331 | 1,999 | +50% | 0 | 0 | — |
▸case-22 I am designing a PostgreSQL database schema to store user accounts, organization memberships, and hashed passwords with foreign key constraints. What DDL SQL statements should I write for this schema? | pass→pass | 14,519 | 11,248 | -23% | 1 | 1 | 0% | 3,141 | 2,738 | -13% | 0 | 0 | — |
▸case-23 Write a PyTorch script using Hugging Face PEFT to apply LoRA fine-tuning to a causal language model with target modules q_proj and v_proj. | pass→pass | 13,012 | 8,623 | -34% | 1 | 1 | 0% | 2,902 | 2,308 | -20% | 0 | 0 | — |