▸case-01 Our customer support bot built on Claude 3.5 Sonnet (200k limit) is ignoring crucial instructions, and our latency and cost are exploding. Average requests reach ~50k tokens because we dump 20 full knowledge base articles, 15 tool schemas, and un-truncated chat history, while dynamic user timestamps sit inside the top prompt header. Here is a recent full API request payload: [system prompt + tools + retrieval + chat history]. Please perform a context window audit. I need an inventory breakdown table listing each prompt element, token footprint, static/dynamic status, and a keep/cut/restructure verdict with reasoning. Highlight instruction contradictions, propose a cache-friendly ordering, and set a component token budget with specific code enforcement points. | fail→pass | 22,424 | 17,120 | -24% | 1 | 1 | 0% | 4,716 | 3,847 | -18% | 0 | 0 | — |
▸case-02 We are running an automated code review agent using GPT-4o (128k limit). Request logs show context sizes hovering around 40k tokens. Users report the model frequently forgets persona rules and tool selection accuracy is low because all 30 repository tools are loaded every turn. Here is the complete payload of a logged API request: [JSON payload with system prompt, tool definitions, retrieved files, and message stack]. Could you conduct a review of this assembled window? I need a component table with token estimates, static vs dynamic flags, and proposed actions (keep, cut, or restructure). Please call out conflicting instructions across sections, show an optimized ordering to maximize cache hits, and establish a token budget specifying where in the pipeline to enforce limits. | fail→pass | 26,336 | 15,575 | -41% | 1 | 1 | 0% | 4,135 | 3,471 | -16% | 0 | 0 | — |
▸case-03 Our legal document QA assistant on Claude 3 Haiku is running into token limit spikes and failing quality test cases. The assembly pipeline combines system instructions, raw PDF text extracts, user session metadata, and multi-turn history. We suspect the prompt layout breaks caching and buries critical rules. Please audit this logged context payload: [logged API payload]. I need a report containing: 1) A structured inventory table listing every component, estimated token count, dynamic/static classification, and actionable verdict. 2) An analysis of conflicting directives. 3) Recommended layout re-ordering for cache alignment. 4) A token budget allocation per section along with the enforcement location in code. | fail→pass | 14,784 | 23,817 | +61% | 1 | 1 | 0% | 2,424 | 4,260 | +76% | 0 | 0 | — |
▸case-04 Our customer support model on OpenAI is suffering from hallucination and high latency. We are preparing to fine-tune a Llama 3 8B model instead of engineering the context window. What learning rate, epoch count, batch size, and LoRA rank hyperparameters should we use for fine-tuning on our 10,000 domain Q&A pairs? | pass→pass | 14,292 | 17,277 | +21% | 1 | 1 | 0% | 2,699 | 4,116 | +53% | 0 | 0 | — |
▸case-05 We are starting a new project and need to write a zero-shot system prompt for a copywriter assistant that generates SaaS landing page headlines. We don't have any logged requests or context window logs yet. Can you author a clear, concise system prompt text with tone guidelines and output constraints? | pass→fail | 10,431 | 11,338 | +9% | 1 | 1 | 0% | 1,851 | 2,874 | +55% | 0 | 0 | — |
▸case-06 We are building the ingestion pipeline for our RAG search system. We need to split large PDF documents before storing them in Pinecone. Should we use recursive character text splitting with a 512 chunk size and 50 overlap, or parent document retriever architecture? How should we set up the vector database index parameters? | pass→pass | 18,434 | 18,323 | -1% | 1 | 1 | 0% | 2,969 | 3,903 | +31% | 0 | 0 | — |
▸case-07 In our current AI agent setup, we inject `Current Time: ISO-TIMESTAMP` and `User ID: 9482` right at the top of our system prompt section, followed by 10,000 tokens of immutable system rules and database schemas. Every API request incurs full prompt processing fees and latency without cache hits on Anthropic prompt caching. Should we leave the timestamp at the top so the model knows the time immediately, or change its position? How does this affect caching? | pass→pass | 13,056 | 12,194 | -7% | 1 | 1 | 0% | 2,283 | 3,079 | +35% | 0 | 0 | — |
▸case-08 Our SQL generation assistant loads all 45 database table schemas and 20 custom reporting tool definitions into the prompt on every single turn, even when the user just says 'hello' or asks a simple date clarification. Model accuracy for tool selection is dropping. Should we keep all tool definitions in the prompt prefix to ensure the model always has access to full functionality? | pass→pass | 72,882 | 11,797 | -84% | 1 | 1 | 0% | 2,025 | 2,707 | +34% | 0 | 0 | — |
▸case-09 Our long-running multi-turn sales agent passes the entire 80-turn raw message history in every request payload. Token counts reach 90,000 tokens per call, costing $0.15 per turn. The developer suggests leaving history unbounded because 'Claude has a 200k window anyway'. Audit this policy choice. | pass→pass | 17,538 | 17,557 | +0% | 1 | 1 | 0% | 2,795 | 3,609 | +29% | 0 | 0 | — |
▸case-10 Our search pipeline always retrieves top k=25 document chunks from Elasticsearch and pastes them verbatim into the LLM context, regardless of relevance score. When queries match only 1 relevant document, 24 irrelevant chunks are still stuffed into the prompt, causing the model to cite false facts. What is the context engineering fix? | pass→pass | 13,841 | 12,572 | -9% | 1 | 1 | 0% | 2,181 | 2,891 | +33% | 0 | 0 | — |
▸case-11 Our web research agent scrapes target web pages and dumps raw HTML markup (50,000 tokens of tags, scripts, and navigation bars) directly into the user message context before asking the model to summarize key product features. The model often hits token limits and misses content. How should this raw dump be restructured? | pass→pass | 15,114 | 21,769 | +44% | 1 | 1 | 0% | 2,399 | 2,881 | +20% | 0 | 0 | — |
▸case-12 In our context window assembly, the system prompt states: 'Be extremely concise. Never use polite fluff or introductory phrases. Respond in under 30 words.' However, the 5 few-shot examples embedded in the prompt all show friendly, 200-word responses beginning with 'Hello! I would be delighted to help you with that today.' The model keeps generating long, polite responses. What is the root cause and fix? | pass→pass | 5,119 | 9,470 | +85% | 1 | 1 | 0% | 882 | 2,652 | +201% | 0 | 0 | — |
▸case-13 We have a critical compliance directive: 'Never reveal customer PII under any circumstances'. Currently, this rule is placed in the middle of a 30,000-token retrieved document dump in the prompt payload. During testing, the model occasionally leaks PII when asked directly by adversarial prompts. Where should this directive be positioned? | pass→pass | 10,556 | 11,633 | +10% | 1 | 1 | 0% | 1,724 | 2,793 | +62% | 0 | 0 | — |
▸case-14 We are designing a customer agent on an 8,000 token context limit model. At p95 traffic, retrieved docs hit 4,000 tokens, chat history hits 3,500 tokens, system prompt is 1,500 tokens, and tool definitions are 1,000 tokens, totaling 10,000 tokens and causing frequent context length errors. How should we set up a component token budget? | pass→pass | 17,010 | 34,243 | +101% | 1 | 1 | 0% | 3,041 | 3,926 | +29% | 0 | 0 | — |
▸case-15 Here is our prompt template file: `System: You are an expert analyst. {{retrieved_docs}} {{chat_history}} {{user_query}}`. We haven't logged any actual API requests yet. Can you review this prompt template and confirm our average request token count and dynamic bloat risks? | pass→pass | 11,851 | 13,327 | +12% | 1 | 1 | 0% | 2,033 | 3,321 | +63% | 0 | 0 | — |
▸case-16 Our team lead wants to delete a 400-token system prompt section titled 'Edge Case Rules' because they feel it looks too long and intuitive cuts save money. However, no automated regression tests were run to check if accuracy drops on complex queries. What does context engineering require before shipping context cuts? | pass→pass | 11,149 | 13,131 | +18% | 1 | 1 | 0% | 1,949 | 3,082 | +58% | 0 | 0 | — |
▸case-17 We defined a nice token budget on paper: System Prompt = 1k, History = 2k, Retrieval = 3k. But in production, the retrieval module occasionally injects 8,000 tokens because there is no truncation logic in the assembly code. The team relies on the LLM 'figuring it out'. How must token budgets be enforced? | pass→pass | 14,088 | 11,592 | -18% | 1 | 1 | 0% | 2,548 | 3,107 | +22% | 0 | 0 | — |
▸case-18 Our RAG assistant receives 10 retrieved text snippets concatenated into a single unbroken string in the prompt payload. The model frequently mixes up information from older internal policy docs with newer public help articles. How should retrieved chunks be formatted within the window? | pass→pass | 11,665 | 10,645 | -9% | 1 | 1 | 0% | 1,989 | 2,827 | +42% | 0 | 0 | — |
▸case-19 In a logged session, user message turn 1 says 'My account balance is $500'. Turn 10 retrieves fresh database state showing 'Account balance: $50'. The model output incorrectly tells the user 'Your balance is $500' based on turn 1 history. What contradiction occurred and how should the context engineering review address it? | pass→pass | 29,554 | 8,789 | -70% | 1 | 1 | 0% | 1,771 | 2,486 | +40% | 0 | 0 | — |
▸case-20 Here is a logged API request context payload containing system rules, tool definitions, vector search results, session variables, and conversation messages. Audit this window by creating a context inventory that maps every component to its token size, static vs dynamic classification, and owning component/service. | fail→fail | 7,318 | 4,660 | -36% | 1 | 1 | 0% | 1,457 | 1,905 | +31% | 0 | 0 | — |
▸case-21 Our engineering manager recommends a blanket 50% character truncation across all system prompt guidelines, tool descriptions, and retrieved knowledge chunks to reduce token costs across the board. Is this indiscriminate truncation good context engineering practice? | pass→pass | 74,286 | 14,936 | -80% | 1 | 1 | 0% | 2,192 | 3,053 | +39% | 0 | 0 | — |
▸case-22 We want to optimize prompt caching for our Claude 3.5 Sonnet pipeline. We currently assemble requests in this order: `[User Session State (dyn)] -> [System Persona (static)] -> [Tool Schemas (static)] -> [Retrieved Chunks (dyn)] -> [Messages (dyn)]`. Why is prompt caching failing and what is the exact reordering required? | pass→pass | 14,650 | 10,894 | -26% | 1 | 1 | 0% | 2,479 | 3,063 | +24% | 0 | 0 | — |