Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Identify single points of failure, assess recovery capabilities, and produce a prioritized remediation plan by analyzing IaC, scaling configs, and resilience patterns in the codebase.
.claude/skills/bilal140202-reliability-improvement-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 126% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 172% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 99% | 0% |
| case-03 | ✓→✗ | ▼ Worse | 169% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 130% | 0% |
Ask the user:
> What workload would you like me to assess for reliability? Please share: > - Workload name and code packages/directories to analyze > - Availability target (99.9%, 99.95%, 99.99%, etc.) > - Recovery objectives (RTO and RPO if defined) > - Past incidents (optional — recent outages or near-misses)
If context is already provided or you are in a codebase with IaC, proceed directly.
Analyze infrastructure for single points of failure.
You MUST examine:
For each component, document:
You MUST flag as HIGH RISK:
Analyze backup and recovery configurations.
You MUST examine:
For each stateful resource, document:
You MUST flag as HIGH RISK:
Analyze scaling and capacity configurations.
You MUST examine:
You MUST flag as HIGH RISK:
Analyze application code for resilience patterns.
You MUST examine:
For each external integration, document:
You MUST flag as HIGH RISK:
Analyze deployment safety configurations.
You MUST examine:
You MUST flag as HIGH RISK:
---STOP--- Checkpoint: Discovery complete — present findings before evaluation.
> Here is what I discovered about your workload's reliability: > - Architecture: {summary of components and dependencies} > - Single points of failure: {count identified so far} > - Recovery capabilities: {summary of backup/DR status} > > Shall I proceed with the full reliability evaluation, or would you like to adjust scope?
Do NOT proceed past this point until the user explicitly confirms.
For each question, provide: Status, Evidence (file:line), Gaps, Risk.
For each finding, assess using Impact × Likelihood:
Impact: Minor (brief degradation, automatic recovery) | Moderate (extended outage for subset of users, manual intervention needed) | Severe (full outage, data loss, cannot recover within RTO)
Likelihood: Low (requires multiple simultaneous failures) | Medium (single component failure could trigger) | High (normal operational event could trigger, no redundancy)
| Impact | Likelihood | Risk Level | |----------|------------|------------| | Severe | High | Critical | | Severe | Medium | High | | Severe | Low | High | | Moderate | High | High | | Moderate | Medium | Medium | | Moderate | Low | Medium | | Minor | High | Medium | | Minor | Medium | Low | | Minor | Low | Low |
---STOP--- Checkpoint: Assessment complete — confirm findings before generating remediation plan.
> Assessment summary: > - Critical findings: {count} > - High findings: {count} > - Medium/Low findings: {count} > > Shall I produce the full remediation plan, or would you like to discuss specific findings first?
Do NOT proceed past this point until the user explicitly confirms.
markdown# Reliability Improvement Plan: {Workload Name} ## Executive Summary - **Date**: {date} - **Availability Target**: {target} - **Packages Analyzed**: {list} - **Findings**: {X} Critical, {Y} High, {Z} Medium, {W} Low - **Overall Reliability Maturity**: {1-5} — {one-line justification} ## Reliability Scorecard | Domain | Score (1-5) | Key Strength | Key Gap | |--------|-------------|--------------|---------| | Fault Tolerance | {score} | {strength} | {gap} | | Recovery & Backup | {score} | {strength} | {gap} | | Scaling & Capacity | {score} | {strength} | {gap} | | Resilience Patterns | {score} | {strength} | {gap} | | Change Management | {score} | {strength} | {gap} | | Testing & Validation | {score} | {strength} | {gap} | ## Single Points of Failure | Component | Evidence | Failure Impact | Current Mitigation | Risk Level | |-----------|----------|---------------|-------------------|------------| | {component} | {file:line} | {impact} | {mitigation or "None"} | {Critical/High/Medium/Low} | ## Critical and High Risk Findings {For each: ID, domain, title, description, evidence (file:line), impact assessment, recommendation, effort, AWS services} ## Medium and Low Risk Findings {Condensed format} ## Prioritized Remediation Plan ### Quick Wins (< 1 week) | Finding | Action | Impact | Effort | |---------|--------|--------|--------| {Enable Multi-AZ, add health checks, configure DLQs, add timeouts} ### Foundation (1-4 weeks) | Finding | Action | Impact | Effort | Dependencies | |---------|--------|--------|--------|--------------| {Auto-scaling, circuit breakers, backup configs, deployment safety} ### Strategic (1-3 months) | Finding | Action | Impact | Effort | Dependencies | |---------|--------|--------|--------|--------------| {Multi-region DR, chaos engineering, cell-based architecture} ## Testing Plan | Test | Validates | Frequency | AWS Service | Evidence Exists | |------|-----------|-----------|-------------|-----------------| | AZ failover | Compute survives AZ loss | Monthly | FIS | {Yes/No} | | Database failover | RDS failover < 60s | Quarterly | FIS | {Yes/No} | | Load test | Handles 2x peak | Before releases | Load Testing | {Yes/No} | | Backup restore | RPO met, data recoverable | Monthly | AWS Backup | {Yes/No} | | Deployment rollback | Bad deploy reverted < 5 min | Every deploy | CodeDeploy | {Yes/No} | ## Next Steps {Top 5 concrete reliability actions the team should take this week}
After delivering the plan, offer:
> Would you like me to: > - Design multi-AZ architecture for a specific component? > - Generate FIS experiment templates for chaos engineering? > - Implement circuit breaker patterns for service dependencies? > - Create backup and DR IaC for stateful resources? > - Design a deployment safety configuration with automated rollback?
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 33,731 | 1,754 | -95% | 1 | 1 | 0% | 6,254 | 3,201 | -49% | 0 | 0 | — |
case-02 | pass→pass | 14,293 | 21,307 | +49% | 1 | 1 | 0% | 3,192 | 7,338 | +130% | 0 | 0 | — |
case-03 | pass→fail | 9,553 | 11,551 | +21% | 1 | 1 | 0% | 1,804 | 4,854 | +169% | 0 | 0 | — |
case-04 | pass→pass | 16,898 | 14,201 | -16% | 1 | 1 | 0% | 4,519 | 6,468 | +43% | 0 | 0 | — |
case-05 | pass→pass | 9,942 | 2,177 | -78% | 1 | 1 | 0% | 1,699 | 3,181 | +87% | 0 | 0 | — |
case-06 | pass→pass | 13,366 | 9,906 | -26% | 1 | 1 | 0% | 2,733 | 4,807 | +76% | 0 | 0 | — |
case-07 | pass→pass | 9,085 | 11,298 | +24% | 1 | 1 | 0% | 1,619 | 4,832 | +198% | 0 | 0 | — |
case-08 | pass→pass | 12,034 | 11,315 | -6% | 1 | 1 | 0% | 2,062 | 4,851 | +135% | 0 | 0 | — |
case-09 | pass→pass | 8,703 | 10,994 | +26% | 1 | 1 | 0% | 1,585 | 4,684 | +196% | 0 | 0 | — |
case-10 | pass→pass | 14,932 | 10,932 | -27% | 1 | 1 | 0% | 2,594 | 4,791 | +85% | 0 | 0 | — |
case-11 | pass→pass | 9,239 | 8,376 | -9% | 1 | 1 | 0% | 1,636 | 4,268 | +161% | 0 | 0 | — |
case-12 | pass→pass | 8,044 | 8,361 | +4% | 1 | 1 | 0% | 1,630 | 4,358 | +167% | 0 | 0 | — |
case-13 | pass→pass | 12,081 | 15,203 | +26% | 1 | 1 | 0% | 2,308 | 5,751 | +149% | 0 | 0 | — |
case-14 | pass→pass | 8,635 | 8,308 | -4% | 1 | 1 | 0% | 1,517 | 4,399 | +190% | 0 | 0 | — |
case-15 | pass→pass | 9,033 | 7,581 | -16% | 1 | 1 | 0% | 1,588 | 4,203 | +165% | 0 | 0 | — |
case-16 | pass→pass | 8,706 | 4,251 | -51% | 1 | 1 | 0% | 1,505 | 3,612 | +140% | 0 | 0 | — |
case-17 | fail→pass | 8,591 | 3,663 | -57% | 1 | 1 | 0% | 1,509 | 3,407 | +126% | 0 | 0 | — |
case-18 | pass→pass | 16,726 | 7,755 | -54% | 1 | 1 | 0% | 2,947 | 4,168 | +41% | 0 | 0 | — |
case-19 | fail→pass | 7,442 | 3,707 | -50% | 1 | 1 | 0% | 1,323 | 3,599 | +172% | 0 | 0 | — |
case-20 | pass→pass | 16,541 | 8,531 | -48% | 1 | 1 | 0% | 2,670 | 4,422 | +66% | 0 | 0 | — |
case-21 | pass→pass | 10,884 | 7,434 | -32% | 1 | 1 | 0% | 1,717 | 4,118 | +140% | 0 | 0 | — |
case-22 | fail→pass | 10,847 | 6,133 | -43% | 1 | 1 | 0% | 1,999 | 3,973 | +99% | 0 | 0 | — |
case-23 | pass→pass | 12,903 | 11,041 | -14% | 1 | 1 | 0% | 2,428 | 4,932 | +103% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.