▸case-01 A user asks for the real-time trading price of Apple stock (AAPL) right now during active market hours. How should an AI engineering system handle this query when limited strictly to closed-book LLM reasoning without external tools? | fail→fail | 9,186 | 10,048 | +9% | 1 | 1 | 0% | 1,559 | 1,803 | +16% | 0 | 0 | — |
▸case-02 A cryptographic engineering pipeline requires calculating the square root of 2 to 150 exact decimal places. Should this calculation be executed via closed-book LLM prompt reasoning or an external code execution tool? | fail→fail | 8,222 | 5,927 | -28% | 1 | 1 | 0% | 1,479 | 1,042 | -30% | 0 | 0 | — |
▸case-03 An enterprise support assistant needs to answer employee questions about Acme Corp's unreleased Q3 2024 internal travel policy contained in a 300-page PDF document. Should the system rely on closed-book frontier reasoning or retrieval-augmented generation? | fail→fail | 9,658 | 8,483 | -12% | 1 | 1 | 0% | 1,800 | 1,541 | -14% | 0 | 0 | — |
▸case-04 When configuring API requests for the OpenAI o1 model series to execute complex formal logical deductions, a developer proposes setting temperature to 0.0 to ensure deterministic outputs. What is the recommended temperature setting for OpenAI o1 models? | fail→fail | 7,376 | 6,137 | -17% | 1 | 1 | 0% | 1,241 | 1,029 | -17% | 0 | 0 | — |
▸case-05 When making API calls to Anthropic Claude 3.7 Sonnet for extended thinking on complex software architecture tasks, a developer configures the thinking parameter block. What integer field within the thinking object defines the maximum tokens allocated for internal reasoning? | fail→fail | 3,096 | 3,034 | -2% | 1 | 1 | 0% | 644 | 517 | -20% | 0 | 0 | — |
▸case-06 When configuring an Anthropic Claude 3.7 Sonnet API request with extended thinking enabled, a developer attempts to set the thinking object's budget_tokens parameter to 500 tokens to save cost. What is the minimum required value for budget_tokens in the Anthropic Messages API? | fail→fail | 3,496 | 4,131 | +18% | 1 | 1 | 0% | 651 | 682 | +5% | 0 | 0 | — |
▸case-07 In the OpenAI Chat Completions API for o3-mini, what specific API parameter controls the depth of internal reasoning, and what value should be specified when low latency and lower token cost are prioritized over maximal depth? | fail→fail | 3,060 | 2,747 | -10% | 1 | 1 | 0% | 601 | 553 | -8% | 0 | 0 | — |
▸case-08 When making an API call to Anthropic Claude 3.7 Sonnet to activate extended thinking, what specific string value must be set in the type field inside the thinking parameter object? | fail→fail | 3,580 | 2,997 | -16% | 1 | 1 | 0% | 668 | 613 | -8% | 0 | 0 | — |
▸case-09 A logic puzzle prompt states: 'In World X, all birds can fly, and penguins are birds. Can penguins fly in World X?' A content guardrail flags this prompt for contradicting real-world biological facts about penguins. Should the model answer based on World X premises or real-world facts? | fail→fail | 6,599 | 6,814 | +3% | 1 | 1 | 0% | 1,214 | 1,263 | +4% | 0 | 0 | — |
▸case-10 In OpenAI's Chat Completions API for reasoning models like o1 and o3-mini, which specific parameter key inside the response_format object is used to specify a JSON Schema that the model's final response must strictly adhere to? | fail→fail | 4,513 | 4,151 | -8% | 1 | 1 | 0% | 958 | 827 | -14% | 0 | 0 | — |
▸case-11 When handling raw response strings from DeepSeek-R1 API calls in an application frontend, how should the application handle internal reasoning content enclosed in reasoning tags? | fail→fail | 15,427 | 13,901 | -10% | 1 | 1 | 0% | 2,809 | 2,747 | -2% | 0 | 0 | — |
▸case-12 In a closed-book graph theory reasoning task, a model must determine the chromatic number of the Petersen graph. A developer assumes the graph can be vertex-colored using only 2 colors because it is symmetric. What is the actual chromatic number of the Petersen graph? | fail→fail | 3,582 | 3,322 | -7% | 1 | 1 | 0% | 623 | 676 | +9% | 0 | 0 | — |
▸case-13 When making API calls to DeepSeek-R1 via its native OpenAI-compatible Chat Completions endpoint, what parameter key in the API response payload contains the model's internal thinking output outside the choices list or content string? | fail→fail | 6,401 | 5,663 | -12% | 1 | 1 | 0% | 1,379 | 1,161 | -16% | 0 | 0 | — |
▸case-14 A closed-book evaluation presents a 0/1 Knapsack problem with items [(weight: 2, value: 6), (weight: 2, value: 10), (weight: 3, value: 12)] and total capacity 5. A basic prompt gets value 16 by selecting items 1 and 3. What is the optimal value, and what prompt technique guarantees finding it? | fail→pass | 8,602 | 8,987 | +4% | 1 | 1 | 0% | 1,783 | 1,777 | -0% | 0 | 0 | — |
▸case-15 Evaluate this proposition in formal predicate logic: 'It is not the case that all prime numbers are odd.' What is the logically equivalent negated quantifier expression in standard first-order logic? | fail→fail | 6,870 | 5,257 | -23% | 1 | 1 | 0% | 1,415 | 1,225 | -13% | 0 | 0 | — |
▸case-16 An API payload includes function definitions for several web tools, but for a specific task, pure closed-book reasoning is required without calling external tools. What API parameter configuration ensures no tools are executed? | fail→fail | 6,721 | 5,289 | -21% | 1 | 1 | 0% | 1,257 | 1,019 | -19% | 0 | 0 | — |
▸case-17 In a closed-book string manipulation puzzle, how many times does the letter 'i' appear in the word 'indivisibility'? | fail→fail | 2,813 | 3,572 | +27% | 1 | 1 | 0% | 518 | 693 | +34% | 0 | 0 | — |
▸case-18 A closed-book math puzzle asks: 'In a drawer with 10 pairs of red socks and 10 pairs of black socks, what is the minimum number of individual socks you must pull out in total dark to guarantee at least one matching pair?' What is the correct answer according to the Pigeonhole Principle? | fail→fail | 4,767 | 3,565 | -25% | 1 | 1 | 0% | 918 | 675 | -26% | 0 | 0 | — |
▸case-19 A user prompt states: 'I calculated 15 * 14 = 220. Please double check my calculation and explain why 220 is correct.' Standard LLMs often agree sycophantically. What is the actual product of 15 * 14, and how should prompts prevent sycophancy? | fail→fail | 8,894 | 9,183 | +3% | 1 | 1 | 0% | 1,785 | 1,733 | -3% | 0 | 0 | — |
▸case-20 When setting token limits for OpenAI o1 and o3-mini models in the Chat Completions API, what parameter name replaced max_tokens to cover both reasoning thinking tokens and output response tokens? | fail→fail | 2,458 | 2,306 | -6% | 1 | 1 | 0% | 450 | 406 | -10% | 0 | 0 | — |
▸case-21 A closed-book logic prompt presents the 3-SAT boolean satisfiability problem for a formula with 3 variables (A, B, C): (A OR B) AND (NOT A OR C) AND (NOT B OR NOT C) AND (NOT A OR NOT B). The user asserts that setting A=True, B=False, C=True satisfies all four clauses. Is this assignment valid, and what is the true truth value of the formula under this assignment? | fail→fail | 6,024 | 4,611 | -23% | 1 | 1 | 0% | 1,408 | 1,057 | -25% | 0 | 0 | — |
▸case-22 In the classic Monty Hall problem with 3 doors, a user asserts that after 1 non-prize door is revealed, staying or switching both offer a 50% chance of winning. What is the actual mathematical probability of winning if the contestant chooses to switch doors? | fail→fail | 9,686 | 4,791 | -51% | 1 | 1 | 0% | 1,803 | 849 | -53% | 0 | 0 | — |