Loading skill
Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use this skill when handling production incidents, outages, or critical bugs — from initial detection through resolution and post-mortem. Trigger on keywords: incident, outage, production down, on-call, P1, P2, critical bug, postmortem, root cause analysis, runbook, service disruption.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -4% | 0% |
DETECT → ASSESS → RESPOND → RESOLVE → LEARNEvery incident touches all five. Never skip LEARN — it's the only one that prevents recurrence.
1. What is broken? (specific service, feature, endpoint)
2. Who is affected? (all users, subset, specific region)
3. What is the impact? (data loss? revenue? user-facing?)
4. When did it start? (correlate with recent deploys/changes)
5. Severity: P1 (all users, data loss) / P2 (major feature) / P3 (minor)Communicate immediately — even if you don't know the cause yet:
"We're investigating an issue with [service].
Impact: [who is affected].
We'll update in 15 minutes."Roll back immediately if:
Document as you go:
Timeline:
[time] - Incident detected: [symptom]
[time] - Identified probable cause: [cause]
[time] - Applied fix: [action taken]
[time] - Monitoring for stability
[time] - Incident resolvedmarkdown## Incident: [title] **Date:** | **Duration:** | **Severity:** ### Summary [2-3 sentences: what happened and impact] ### Timeline [Chronological events from detection to resolution] ### Root Cause [The actual cause — not the symptom] ### Contributing Factors [Conditions that allowed this to happen] ### What Went Well [Things that helped during response] ### Action Items | Action | Owner | Due Date | |--------|-------|----------| | [preventive fix] | [name] | [date] | | [monitoring improvement] | [name] | [date] |
For recurring incident types, create a runbook:
markdown## Runbook: [incident type] ### Symptoms [How to recognize this incident] ### Immediate Actions 1. [First thing to check/do] 2. [Second thing] ### Investigation Steps 1. Check [X] for [Y] 2. Run [command] to verify [Z] ### Resolution [Steps to fix] ### Escalation If not resolved in [time]: escalate to [person/team]
Other measured skills in the registry, with their headline benchmark lift.