Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Blameless post-mortem expert for incidents, outages, regressions, customer escalations, missed launches, and failed experiments. Pairs Google SRE practice with Allspaw / Dekker / Perrow systems thinking.
.claude/skills/borghei-post-mortem/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 131% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 80% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 136% | 0% |
A post-mortem is a structured, blameless review held after an incident, outage, regression, missed launch, or failed experiment. The goal is not to assign fault but to learn how the system (people, process, code, and organization) produced the outcome, and to commit to durable changes that reduce the chance of recurrence.
This skill operationalizes the Google SRE blameless post-mortem template, the Etsy "morgue" tradition, John Allspaw's "How Complex Systems Fail" reading, Charles Perrow's Normal Accident Theory, and Sidney Dekker's Field Guide to Understanding "Human Error". Where the companion discovery/pre-mortem/ skill imagines failure before it happens, post-mortem learns from failure that already did.
If the incident is below the severity threshold and the team has seen the same class of failure recently, document the recurrence in the existing post-mortem rather than producing a new one.
Before authoring the post-mortem, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
assets/post_mortem_template.md.The blameless principles, severity thresholds, full template, workflow, and root-cause methods all live in the references below.
In Scope: Post-mortem authoring for incidents, outages, regressions, missed launches, failed experiments, customer escalations, and near misses. Severity classification, blameless facilitation, 5 Whys, Causal Tree analysis, action-item tracking, distribution and archival practice.
Out of Scope: Live incident command and on-call coordination (see delivery-manager/). Sprint retrospectives on team practice (see sprint-retrospective/). Risk surfacing before launch (see discovery/pre-mortem/). Quantitative reliability engineering and error-budget policy (see engineering/ skills if present).
Important Caveats: Blameless culture is a prerequisite, not an output — if management uses post-mortems as a performance signal, the documents become sanitized; establish the separation in writing. Regulated industries (medical devices, aviation, financial services) that require named accountability should run a parallel internal blameless post-mortem alongside the regulator-facing report. A post-mortem is a learning artifact, not a fixing artifact: the fix lives in the action items and their follow-through, so one with zero completed action items is a failed post-mortem. (Cook's 4-page "How Complex Systems Fail" is worth reading before facilitating a complex Sev 0/1 — linked in the blameless culture guide.)
| Integration | Direction | What flows | |---|---|---| | discovery/pre-mortem/ | Bidirectional | Post-mortem findings update next launch's pre-mortem risk register; pre-mortem mitigations become post-mortem-checked controls | | delivery-manager/ | Receives from | Incident response context, severity classification, on-call handoffs | | sprint-retrospective/ | Bidirectional | Team-practice retros surface incident patterns; post-mortems feed retro themes | | daci-framework/ | Feeds into | Action items use DACI to assign owner (D), accountable (A), consulted, informed | | execution/dependency-map/ | Receives from | Cross-team contributing factors map to dependency-graph nodes | | execution/status-update-generator/ | Feeds into | Sev 0/1 incidents surface in weekly executive status updates | | senior-pm/ | Feeds into | Repeated incident classes feed portfolio risk register via risk_matrix_analyzer.py | | scrum-master/ | Feeds into | Action items become sprint backlog items with mitigation-focused stories |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 22,141 | 35,057 | +58% | 1 | 1 | 0% | 3,340 | 7,239 | +117% | 0 | 0 | — |
case-02 | fail→fail | 20,157 | 19,169 | -5% | 1 | 1 | 0% | 2,866 | 4,904 | +71% | 0 | 0 | — |
case-03 | fail→fail | 13,285 | 11,950 | -10% | 1 | 1 | 0% | 2,089 | 3,779 | +81% | 0 | 0 | — |
case-04 | pass→pass | 18,360 | 21,201 | +15% | 1 | 1 | 0% | 2,874 | 5,171 | +80% | 0 | 0 | — |
case-05 | pass→pass | 10,637 | 11,725 | +10% | 1 | 1 | 0% | 1,528 | 3,602 | +136% | 0 | 0 | — |
case-06 | pass→pass | 8,433 | 9,773 | +16% | 1 | 1 | 0% | 1,242 | 3,255 | +162% | 0 | 0 | — |
case-07 | pass→pass | 12,038 | 11,935 | -1% | 1 | 1 | 0% | 1,759 | 3,559 | +102% | 0 | 0 | — |
case-08 | pass→pass | 8,682 | 4,578 | -47% | 1 | 1 | 0% | 1,290 | 2,588 | +101% | 0 | 0 | — |
case-09 | fail→pass | 8,775 | 5,238 | -40% | 1 | 1 | 0% | 1,323 | 2,734 | +107% | 0 | 0 | — |
case-10 | pass→pass | 14,302 | 14,362 | +0% | 1 | 1 | 0% | 2,213 | 4,240 | +92% | 0 | 0 | — |
case-11 | fail→fail | 7,418 | 6,148 | -17% | 1 | 1 | 0% | 1,107 | 2,767 | +150% | 0 | 0 | — |
case-12 | pass→pass | 9,632 | 8,555 | -11% | 1 | 1 | 0% | 1,455 | 3,197 | +120% | 0 | 0 | — |
case-13 | pass→pass | 18,036 | 18,133 | +1% | 1 | 1 | 0% | 2,611 | 4,455 | +71% | 0 | 0 | — |
case-18 | pass→pass | 17,749 | 14,254 | -20% | 1 | 1 | 0% | 2,486 | 4,002 | +61% | 0 | 0 | — |
case-14 | pass→pass | 11,090 | 13,210 | +19% | 1 | 1 | 0% | 1,576 | 4,069 | +158% | 0 | 0 | — |
case-15 | pass→pass | 15,140 | 14,909 | -2% | 1 | 1 | 0% | 2,208 | 4,277 | +94% | 0 | 0 | — |
case-16 | pass→pass | 11,551 | 9,698 | -16% | 1 | 1 | 0% | 1,643 | 3,326 | +102% | 0 | 0 | — |
case-17 | pass→pass | 12,890 | 11,358 | -12% | 1 | 1 | 0% | 1,967 | 3,625 | +84% | 0 | 0 | — |
case-19 | pass→pass | 13,837 | 10,833 | -22% | 1 | 1 | 0% | 2,061 | 3,523 | +71% | 0 | 0 | — |
case-20 | fail→pass | 10,672 | 11,589 | +9% | 1 | 1 | 0% | 1,739 | 3,754 | +116% | 0 | 0 | — |
case-21 | fail→pass | 8,857 | 5,917 | -33% | 1 | 1 | 0% | 1,254 | 2,894 | +131% | 0 | 0 | — |
case-22 | pass→pass | 21,243 | 10,985 | -48% | 1 | 1 | 0% | 3,052 | 3,658 | +20% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +14 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.