Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Manage active production incidents through detection, triage, mitigation, communication, and resolution with structured roles and decision-making. Use this skill whenever the user has an active incident, a production issue, a service outage, a security incident, or needs to plan incident response procedures. Triggers on incident response, production incident, outage, service down, site down, P0, P1, severity, downtime, on-call, incident commander, status page, postmortem prep. Also triggers when
.claude/skills/rampstackco-incident-response/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 159% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 108% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 99% | 0% |
Manage active production incidents from detection to resolution. Stack-agnostic. Tool-agnostic.
This skill is for active incidents and incident process. For after-the-fact analysis, use after-action-report. For planned launches, use launch-runbook.
after-action-report)launch-runbook)qa-testing)How the incident becomes known.
Detection sources:
On detection:
Establish severity and impact.
Severity rubric:
| Severity | Definition | Response | |---|---|---| | SEV-1 (Critical) | Major customer-facing functionality broken. Data integrity at risk. Security breach. | All-hands. Incident commander. Active war room. Public communication required. | | SEV-2 (Major) | Significant degradation. Some customers affected. Revenue impact. | Incident commander assigned. Active response. Internal communication. May or may not need public communication. | | SEV-3 (Minor) | Limited impact. Workaround available. Affecting a small group of users. | Standard on-call response. Single owner. | | SEV-4 (Low) | Cosmetic, edge-case, or low-frequency. No urgent action needed. | Tracked as bug. Addressed in normal queue. |
Severity can change. Re-evaluate as more info emerges.
Stop the bleeding before fixing the cause.
Mitigation patterns (faster than full fix):
Mitigation principle: Stop user impact first. Cause analysis second.
Three audiences during an incident:
Internal team:
Internal stakeholders:
External / customers:
Communication principles:
Verified fix, customers restored, incident closed.
Resolution criteria:
After closure:
| Role | Responsibility | |---|---| | Incident commander (IC) | Owns the response. Calls decisions. Assigns work. Not necessarily the most technical person; needs to coordinate. | | Communications lead | Owns internal and external messaging. Reduces IC's communication burden. | | Operations lead | Drives the technical investigation and mitigation. Often the most senior on-call engineer. | | Scribe | Captures the timeline as the incident unfolds. Critical for AAR. | | Subject matter experts | Pulled in as needed. Service owners, database experts, security experts. |
For small teams or low-severity incidents, one person can hold multiple roles. Each role's responsibilities should still be explicit.
The IC's authority:
Non-decisions to avoid:
When in doubt: act. A wrong action that can be rolled back beats inaction while users suffer.
Initial: > "We are investigating reports of issue]. Updates to follow."
Identified: > "We have identified the issue affecting scope]. Engineers are working on a fix. Next update by time]."
Monitoring: > "A fix has been applied. We are monitoring to confirm resolution. Next update by time]."
Resolved: > "This incident has been resolved. Service has been restored. A full incident report will be posted within timeframe]."
Patterns to avoid:
During an active incident: incident channel updates and status page updates as per the framework above.
After incident close: a brief incident summary feeding into the AAR.
markdown# Incident: [Brief title] **Date:** [YYYY-MM-DD] **Severity:** [SEV-1 / 2 / 3 / 4] **Duration:** [Detection to resolution] **Customer impact:** [Who, how many, how, or state the gap per the data-availability rule] ## Summary [1 to 2 paragraphs] ## Timeline [Timestamped events] ## Mitigation [What was done] ## Action items [Follow-ups, with owners] ## AAR scheduled for [Date]
This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
references/incident-playbook.md - Severity definitions, roles, status page templates, decision rubrics.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | pass→pass | 7,876 | 5,811 | -26% | 1 | 1 | 0% | 1,305 | 3,086 | +136% | 0 | 0 | — |
case-11 | pass→pass | 11,481 | 8,178 | -29% | 1 | 1 | 0% | 1,765 | 3,407 | +93% | 0 | 0 | — |
case-01 | fail→pass | 16,430 | 10,859 | -34% | 1 | 1 | 0% | 2,716 | 4,090 | +51% | 0 | 0 | — |
case-02 | fail→fail | 20,609 | 20,103 | -2% | 1 | 1 | 0% | 3,250 | 5,287 | +63% | 0 | 0 | — |
case-03 | pass→pass | 24,081 | 29,312 | +22% | 1 | 1 | 0% | 3,823 | 6,091 | +59% | 0 | 0 | — |
case-04 | fail→fail | 26,378 | 30,213 | +15% | 1 | 1 | 0% | 4,455 | 7,202 | +62% | 0 | 0 | — |
case-05 | pass→pass | 18,359 | 23,882 | +30% | 1 | 1 | 0% | 2,797 | 5,762 | +106% | 0 | 0 | — |
case-06 | pass→pass | 14,144 | 10,286 | -27% | 1 | 1 | 0% | 2,149 | 3,823 | +78% | 0 | 0 | — |
case-07 | pass→pass | 10,558 | 6,442 | -39% | 1 | 1 | 0% | 1,558 | 3,113 | +100% | 0 | 0 | — |
case-08 | fail→pass | 15,510 | 10,205 | -34% | 1 | 1 | 0% | 2,471 | 3,625 | +47% | 0 | 0 | — |
case-09 | fail→pass | 8,080 | 7,859 | -3% | 1 | 1 | 0% | 1,289 | 3,335 | +159% | 0 | 0 | — |
case-10 | pass→pass | 11,517 | 7,869 | -32% | 1 | 1 | 0% | 1,647 | 3,327 | +102% | 0 | 0 | — |
case-13 | pass→pass | 9,488 | 8,047 | -15% | 1 | 1 | 0% | 1,382 | 3,440 | +149% | 0 | 0 | — |
case-14 | fail→fail | 13,809 | 12,836 | -7% | 1 | 1 | 0% | 2,035 | 4,109 | +102% | 0 | 0 | — |
case-15 | pass→pass | 12,301 | 8,989 | -27% | 1 | 1 | 0% | 1,886 | 3,505 | +86% | 0 | 0 | — |
case-16 | fail→pass | 10,918 | 8,811 | -19% | 1 | 1 | 0% | 1,721 | 3,576 | +108% | 0 | 0 | — |
case-17 | pass→pass | 11,676 | 4,741 | -59% | 1 | 1 | 0% | 1,758 | 2,954 | +68% | 0 | 0 | — |
case-18 | pass→pass | 7,490 | 6,583 | -12% | 1 | 1 | 0% | 1,144 | 3,149 | +175% | 0 | 0 | — |
case-19 | pass→fail | 9,828 | 7,249 | -26% | 1 | 1 | 0% | 1,588 | 3,266 | +106% | 0 | 0 | — |
case-20 | pass→pass | 11,175 | 3,753 | -66% | 1 | 1 | 0% | 1,763 | 2,720 | +54% | 0 | 0 | — |
case-21 | fail→pass | 10,084 | 6,969 | -31% | 1 | 1 | 0% | 1,638 | 3,262 | +99% | 0 | 0 | — |
case-22 | pass→pass | 9,400 | 6,784 | -28% | 1 | 1 | 0% | 1,449 | 3,229 | +123% | 0 | 0 | — |
case-23 | fail→pass | 8,359 | 5,257 | -37% | 1 | 1 | 0% | 1,358 | 2,965 | +118% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +22 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.