Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when analyzing failures, outages, incidents, or negative outcomes, conducting blameless postmortems, documenting root causes with 5 Whys or fishbone diagrams, identifying corrective actions with owners and timelines, learning from near-misses, establishing prevention strategies, or when user mentions postmortem, incident review, failure analysis, RCA, lessons learned, or after-action review.
.claude/skills/nicepkg-postmortem/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 210% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 165% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 135% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 298% | 0% |
Conduct blameless postmortems that transform failures into learning opportunities by documenting what happened, why it happened, impact quantification, root cause analysis, and actionable preventions with clear ownership.
Use this skill when:
Do NOT use when:
Postmortem is a structured, blameless analysis of failures that answers:
Key Principles:
Quick Example:
Incident: Database outage, 2-hour downtime, 50K users affected
Timeline:
Impact: 2-hour outage, 50K users unable to access, estimated $20K revenue loss
Root Cause (5 Whys):
Corrective Actions:
What Went Well: Alerts fired quickly, team responded within 5 minutes, good communication
Copy this checklist and track your progress:
Postmortem Progress:
- [ ] Step 1: Assemble timeline and quantify impact
- [ ] Step 2: Conduct root cause analysis
- [ ] Step 3: Define corrective and preventive actions
- [ ] Step 4: Document and share postmortem
- [ ] Step 5: Track action items to completionStep 1: Assemble timeline and quantify impact
Gather facts: when detected, when started, key events, when resolved. Quantify impact: users affected, duration, revenue/SLA impact, customer complaints. For straightforward incidents use resources/template.md. For complex incidents with multiple causes or cascading failures, study resources/methodology.md for advanced timeline reconstruction techniques.
Step 2: Conduct root cause analysis
Ask "Why?" 5 times to get from symptom to root cause, or use fishbone diagram for complex incidents with multiple contributing factors. See Root Cause Analysis Techniques for guidance. Focus on system failures (process gaps, missing safeguards) not human errors.
Step 3: Define corrective and preventive actions
For each root cause, identify actions to prevent recurrence. Must be specific (not "improve testing"), owned (named person), and time-bound (deadline). Categorize as immediate fixes vs. long-term improvements. See Corrective Actions for framework.
Step 4: Document and share postmortem
Create postmortem document using template. Include timeline, impact, root cause, actions, what went well. Share widely (engineering, product, leadership) to enable learning. Present in team meeting for discussion. Archive in knowledge base.
Step 5: Track action items to completion
Assign owners, set deadlines, add to project tracker. Review progress in standups or weekly meetings. Close postmortem only when all actions complete. Self-assess quality using resources/evaluators/rubric_postmortem.json. Minimum standard: ≥3.5 average score.
Production Outages (system failures, downtime):
Security Incidents (breaches, vulnerabilities):
Product/Project Failures (launches, deadlines):
Process Failures (operational, procedural):
Human Error (surface cause, dig deeper):
Process Gap (missing or unclear procedures):
Technical Debt (deferred maintenance):
External Dependencies (third-party failures):
Systemic Issues (organizational, cultural):
5 Whys:
Example: Database outage → Why? Bad config → Why? Wrong value → Why? Template error → Why? New team member unfamiliar → Why? No config review in onboarding
Fishbone Diagram (Ishikawa):
Fault Tree Analysis:
Types of Actions:
SMART Actions:
Prioritization:
Prevention Hierarchy (from most to least effective):
Blameless Culture:
Root Cause Depth:
Actionability:
Impact Quantification:
Timeliness:
Resources:
Success Criteria:
Common Mistakes:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 13,576 | 9,769 | -28% | 1 | 1 | 0% | 1,989 | 4,679 | +135% | 0 | 0 | — |
case-02 | pass→pass | 6,478 | 4,619 | -29% | 1 | 1 | 0% | 1,000 | 3,980 | +298% | 0 | 0 | — |
case-03 | fail→pass | 9,571 | 8,122 | -15% | 1 | 1 | 0% | 1,430 | 4,428 | +210% | 0 | 0 | — |
case-04 | pass→pass | 10,217 | 7,824 | -23% | 1 | 1 | 0% | 1,518 | 4,407 | +190% | 0 | 0 | — |
case-05 | pass→pass | 5,717 | 7,288 | +27% | 1 | 1 | 0% | 993 | 4,403 | +343% | 0 | 0 | — |
case-06 | pass→pass | 10,914 | 8,694 | -20% | 1 | 1 | 0% | 1,673 | 4,586 | +174% | 0 | 0 | — |
case-07 | fail→fail | 10,892 | 8,268 | -24% | 1 | 1 | 0% | 1,535 | 4,458 | +190% | 0 | 0 | — |
case-08 | pass→pass | 7,025 | 5,458 | -22% | 1 | 1 | 0% | 1,073 | 3,985 | +271% | 0 | 0 | — |
case-09 | pass→pass | 8,386 | 4,993 | -40% | 1 | 1 | 0% | 1,330 | 3,941 | +196% | 0 | 0 | — |
case-10 | pass→pass | 13,228 | 11,593 | -12% | 1 | 1 | 0% | 2,014 | 4,944 | +145% | 0 | 0 | — |
case-11 | pass→pass | 14,735 | 12,416 | -16% | 1 | 1 | 0% | 2,311 | 5,072 | +119% | 0 | 0 | — |
case-12 | pass→pass | 14,286 | 10,610 | -26% | 1 | 1 | 0% | 1,972 | 4,769 | +142% | 0 | 0 | — |
case-13 | pass→pass | 9,080 | 5,749 | -37% | 1 | 1 | 0% | 1,266 | 4,062 | +221% | 0 | 0 | — |
case-14 | pass→pass | 14,595 | 12,626 | -13% | 1 | 1 | 0% | 2,192 | 5,041 | +130% | 0 | 0 | — |
case-15 | pass→pass | 14,484 | 9,496 | -34% | 1 | 1 | 0% | 2,016 | 4,544 | +125% | 0 | 0 | — |
case-16 | fail→pass | 12,462 | 7,459 | -40% | 1 | 1 | 0% | 1,839 | 4,321 | +135% | 0 | 0 | — |
case-17 | pass→pass | 11,110 | 11,735 | +6% | 1 | 1 | 0% | 1,654 | 5,047 | +205% | 0 | 0 | — |
case-18 | pass→pass | 13,378 | 10,209 | -24% | 1 | 1 | 0% | 2,020 | 4,738 | +135% | 0 | 0 | — |
case-19 | pass→pass | 13,025 | 11,688 | -10% | 1 | 1 | 0% | 1,949 | 4,857 | +149% | 0 | 0 | — |
case-20 | pass→pass | 13,645 | 12,912 | -5% | 1 | 1 | 0% | 2,063 | 5,183 | +151% | 0 | 0 | — |
case-21 | pass→pass | 15,403 | 14,855 | -4% | 1 | 1 | 0% | 2,379 | 5,409 | +127% | 0 | 0 | — |
case-22 | pass→fail | 16,637 | 25,677 | +54% | 1 | 1 | 0% | 2,663 | 7,047 | +165% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.