▸case-01 An outage occurred because senior engineer Alex accidentally ran a TRUNCATE query against the production database instead of staging. Management wants to identify who approved the query and assign a performance reprimand. How should the postmortem frame this incident? | pass→pass | 12,753 | 19,611 | +54% | 1 | 1 | 0% | 2,032 | 3,611 | +78% | 0 | 0 | — |
▸case-02 We are scheduling a 60-minute postmortem meeting for a recent P1 incident. The team lead wants to spend 35 minutes reviewing the detailed event timeline and 5 minutes on root cause analysis and action items. How should the 60-minute agenda be structured? | pass→pass | 12,137 | 8,510 | -30% | 1 | 1 | 0% | 2,090 | 2,656 | +27% | 0 | 0 | — |
▸case-03 A production service went down due to database connection pool exhaustion caused by a new feature bypass. The engineering manager suggests stopping the postmortem analysis at 'Why did the connections exhaust?' since fixing the pool configuration solves the immediate bug. What depth of root cause analysis should be performed? | pass→pass | 12,625 | 13,303 | +5% | 1 | 1 | 0% | 1,958 | 3,146 | +61% | 0 | 0 | — |
▸case-04 Following an outage, the team generated action items like 'Improve database monitoring' and 'Make deployments safer' without specific assignees or target dates, arguing that team-wide ownership ensures faster completion. How should postmortem action items be defined? | fail→pass | 12,778 | 15,846 | +24% | 1 | 1 | 0% | 2,106 | 3,823 | +82% | 0 | 0 | — |
▸case-05 A minor internal staging environment deployment failed for 10 minutes with zero customer impact. The team is debating whether to write a full 10-page standard postmortem report with system diagrams and business impact metrics. How should minor incidents be documented? | pass→pass | 12,600 | 9,450 | -25% | 1 | 1 | 0% | 2,024 | 2,779 | +37% | 0 | 0 | — |
▸case-06 A critical microservice was returning 500 errors for 45 minutes before a customer submitted a support ticket, at which point alerts finally fired. The incident report draft focuses solely on fixing the underlying memory leak code. What incident postmortem section or analysis is missing? | pass→pass | 11,505 | 11,749 | +2% | 1 | 1 | 0% | 1,854 | 3,039 | +64% | 0 | 0 | — |
▸case-07 During a major network partition, automatic failover failed, but traffic was unexpectedly low due to a regional holiday, preventing total system collapse. A draft postmortem omits this detail as irrelevant to technical failure modes. Should this detail be included in the review? | pass→pass | 10,488 | 10,110 | -4% | 1 | 1 | 0% | 1,647 | 2,726 | +66% | 0 | 0 | — |
▸case-08 An incident timeline lists entries like 'In the morning', 'A few minutes later', and 'Around lunchtime'. Is this phrasing adequate for an incident postmortem timeline? | pass→pass | 11,627 | 8,818 | -24% | 1 | 1 | 0% | 1,801 | 2,642 | +47% | 0 | 0 | — |
▸case-09 The incident commander opens the postmortem meeting by immediately opening the code diff to debate the pull request changes. What should occur in the first 5 minutes of a postmortem meeting? | pass→pass | 11,876 | 8,803 | -26% | 1 | 1 | 0% | 2,074 | 2,588 | +25% | 0 | 0 | — |
▸case-10 An incident report summarizes impact as 'Service was down for 2 hours.' How should impact be categorized to provide full context to stakeholders? | pass→pass | 12,327 | 11,135 | -10% | 1 | 1 | 0% | 2,104 | 2,916 | +39% | 0 | 0 | — |
▸case-11 Six months ago, a similar caching failure occurred, but the team considers past incidents irrelevant since the codebase has been refactored. Should historical incidents be referenced in the new postmortem? | pass→pass | 11,538 | 10,378 | -10% | 1 | 1 | 0% | 1,760 | 2,593 | +47% | 0 | 0 | — |
▸case-12 An automated deployment had a minor transient retry spike that resolved automatically in 30 seconds without breach of SLA or customer impact. The team lead wants to mandate a mandatory formal postmortem. How should postmortem triggers be applied? | pass→pass | 12,632 | 13,777 | +9% | 1 | 1 | 0% | 2,038 | 3,287 | +61% | 0 | 0 | — |
▸case-13 To close out a postmortem quickly, the team proposes a single action item: 'Restart the API server manually if memory usage hits 90%.' How should action items address root causes? | pass→pass | 12,146 | 12,258 | +1% | 1 | 1 | 0% | 2,031 | 3,236 | +59% | 0 | 0 | — |
▸case-14 A complex cascade failure involved four microservices, two message queues, and a third-party payment gateway. The author wrote a purely textual narrative without visual context. What structural element should be included to clarify architecture failures? | fail→pass | 10,020 | 6,277 | -37% | 1 | 1 | 0% | 1,514 | 2,326 | +54% | 0 | 0 | — |
▸case-15 A postmortem draft contains only failure points, root causes, and corrective actions. An engineer asks if positive operational responses (such as rapid manual traffic shedding) should be removed to keep the document strictly focused on failure. How should effective incident handling be documented? | pass→pass | 12,288 | 12,612 | +3% | 1 | 1 | 0% | 1,927 | 3,128 | +62% | 0 | 0 | — |
▸case-16 As time runs out in a postmortem meeting, the meeting lead abruptly ends the call while participants are arguing over action item priorities. How should the final 5 minutes of a postmortem meeting be conducted? | pass→pass | 11,697 | 9,870 | -16% | 1 | 1 | 0% | 1,889 | 2,825 | +50% | 0 | 0 | — |
▸case-17 A draft postmortem states: 'The on-call engineer carelessly ignored the PagerDuty alert for 15 minutes because they were distracted.' How should this observation be rephrased? | pass→pass | 7,652 | 7,408 | -3% | 1 | 1 | 0% | 1,337 | 2,398 | +79% | 0 | 0 | — |
▸case-18 A postmortem claims that database latency spiked to 10 seconds, but provides no dashboard links, log snippets, or metric graphs. How should technical assertions in a postmortem be substantiated? | fail→pass | 13,621 | 15,961 | +17% | 1 | 1 | 0% | 2,221 | 3,721 | +68% | 0 | 0 | — |
▸case-19 When conducting a 5 Whys analysis on a bug that bypassed code review, the analysis ends at 'The reviewer missed the bug.' What deeper organizational question should be asked? | pass→pass | 9,823 | 9,800 | -0% | 1 | 1 | 0% | 1,577 | 2,764 | +75% | 0 | 0 | — |
▸case-20 A P1 outage is currently active; production database CPU is at 100% and checkout is completely down right now. Provide immediate, step-by-step incident command triage instructions to mitigate customer impact immediately. | pass→pass | 17,157 | 13,455 | -22% | 1 | 1 | 0% | 2,430 | 3,185 | +31% | 0 | 0 | — |
▸case-21 We are launching a new user authentication service and need to define our Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for latency and availability. Draft the SLO specification document. | pass→pass | 21,260 | 20,146 | -5% | 1 | 1 | 0% | 3,252 | 4,950 | +52% | 0 | 0 | — |
▸case-22 Write a Kubernetes Argo Rollouts YAML manifest that implements a canary deployment strategy with automated metric analysis for HTTP error rates. | pass→pass | 17,261 | 16,322 | -5% | 1 | 1 | 0% | 2,703 | 4,364 | +61% | 0 | 0 | — |