Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Defines service level objectives, creates error budget policies, designs incident response procedures, develops capacity models, and produces monitoring configurations and automation scripts for production systems. Use when defining SLIs/SLOs, managing error budgets, building reliable systems at scale, incident management, chaos engineering, toil reduction, or capacity planning.
.claude/skills/jeffallan-sre-engineer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 193% | 0% |
| case-19 | ✓→✗ | ▼ Worse | 72% | 0% |
Load detailed guidance based on context:
| Topic | Reference | Load When | |-------|-----------|-----------| | SLO/SLI | references/slo-sli-management.md | Defining SLOs, calculating error budgets | | Error Budgets | references/error-budget-policy.md | Managing budgets, burn rates, policies | | Monitoring | references/monitoring-alerting.md | Golden signals, alert design, dashboards | | Automation | references/automation-toil.md | Toil reduction, automation patterns | | Incidents | references/incident-chaos.md | Incident response, chaos engineering |
When implementing SRE practices, provide:
# 99.9% availability SLO over a 30-day window
# Allowed downtime: (1 - 0.999) * 30 * 24 * 60 = 43.2 minutes/month
# Error budget (request-based): 0.001 * total_requests
# Example: 10M requests/month → 10,000 error budget requests
# If 5,000 errors consumed in week 1 → 50% budget burned in 25% of window
# → Trigger error budget policy: freeze non-critical releasesyamlgroups: - name: slo_availability rules: # Fast burn: 2% budget in 1h (14.4x burn rate) - alert: HighErrorBudgetBurn expr: | ( sum(rate(http_requests_total{status=~"5.."}[1h])) / sum(rate(http_requests_total[1h])) ) > 0.014400 and ( sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) ) > 0.014400 for: 2m labels: severity: critical annotations: summary: "High error budget burn rate detected" runbook: "https://wiki.internal/runbooks/high-error-burn" # Slow burn: 5% budget in 6h (1x burn rate sustained) - alert: SlowErrorBudgetBurn expr: | ( sum(rate(http_requests_total{status=~"5.."}[6h])) / sum(rate(http_requests_total[6h])) ) > 0.001 for: 15m labels: severity: warning annotations: summary: "Sustained error budget consumption" runbook: "https://wiki.internal/runbooks/slow-error-burn"
promql# Latency — 99th percentile request duration histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, service)) # Traffic — requests per second by service sum(rate(http_requests_total[5m])) by (service) # Errors — error rate ratio sum(rate(http_requests_total{status=~"5.."}[5m])) by (service) / sum(rate(http_requests_total[5m])) by (service) # Saturation — CPU throttling ratio sum(rate(container_cpu_cfs_throttled_seconds_total[5m])) by (pod) / sum(rate(container_cpu_cfs_periods_total[5m])) by (pod)
python#!/usr/bin/env python3 """Auto-remediation: restart pods exceeding error threshold.""" import subprocess, sys, json ERROR_THRESHOLD = 0.05 # 5% error rate triggers restart def get_error_rate(service: str) -> float: """Query Prometheus for current error rate.""" import urllib.request query = f'sum(rate(http_requests_total{{status=~"5..",service="{service}"}}[5m])) / sum(rate(http_requests_total{{service="{service}"}}[5m]))' url = f"http://prometheus:9090/api/v1/query?query={urllib.request.quote(query)}" with urllib.request.urlopen(url) as resp: data = json.load(resp) results = data["data"]["result"] return float(results[0]["value"][1]) if results else 0.0 def restart_deployment(namespace: str, deployment: str) -> None: subprocess.run( ["kubectl", "rollout", "restart", f"deployment/{deployment}", "-n", namespace], check=True ) print(f"Restarted {namespace}/{deployment}") if __name__ == "__main__": service, namespace, deployment = sys.argv[1], sys.argv[2], sys.argv[3] rate = get_error_rate(service) print(f"Error rate for {service}: {rate:.2%}") if rate > ERROR_THRESHOLD: restart_deployment(namespace, deployment) else: print("Within SLO threshold — no action required")
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | pass→pass | 13,060 | 12,234 | -6% | 1 | 1 | 0% | 2,324 | 3,804 | +64% | 0 | 0 | — |
case-01 | fail→pass | 29,158 | 19,115 | -34% | 1 | 1 | 0% | 6,232 | 5,978 | -4% | 0 | 0 | — |
case-02 | fail→fail | 30,870 | 26,492 | -14% | 1 | 1 | 0% | 6,220 | 7,112 | +14% | 0 | 0 | — |
case-03 | fail→pass | 27,257 | 22,589 | -17% | 1 | 1 | 0% | 5,983 | 6,474 | +8% | 0 | 0 | — |
case-04 | fail→fail | 14,524 | 12,707 | -13% | 1 | 1 | 0% | 2,763 | 4,350 | +57% | 0 | 0 | — |
case-05 | pass→pass | 15,471 | 12,915 | -17% | 1 | 1 | 0% | 3,343 | 4,245 | +27% | 0 | 0 | — |
case-06 | pass→pass | 15,317 | 17,718 | +16% | 1 | 1 | 0% | 2,407 | 4,921 | +104% | 0 | 0 | — |
case-07 | pass→pass | 8,774 | 12,388 | +41% | 1 | 1 | 0% | 1,778 | 4,008 | +125% | 0 | 0 | — |
case-08 | pass→pass | 10,893 | 10,196 | -6% | 1 | 1 | 0% | 1,802 | 3,497 | +94% | 0 | 0 | — |
case-09 | pass→pass | 10,138 | 14,093 | +39% | 1 | 1 | 0% | 1,797 | 4,282 | +138% | 0 | 0 | — |
case-10 | pass→pass | 9,916 | 7,959 | -20% | 1 | 1 | 0% | 1,888 | 3,097 | +64% | 0 | 0 | — |
case-12 | fail→pass | 10,538 | 8,556 | -19% | 1 | 1 | 0% | 1,675 | 3,143 | +88% | 0 | 0 | — |
case-13 | pass→pass | 16,559 | 16,109 | -3% | 1 | 1 | 0% | 3,413 | 5,136 | +50% | 0 | 0 | — |
case-14 | pass→pass | 13,179 | 13,004 | -1% | 1 | 1 | 0% | 2,410 | 4,044 | +68% | 0 | 0 | — |
case-15 | pass→pass | 6,819 | 5,879 | -14% | 1 | 1 | 0% | 1,142 | 2,760 | +142% | 0 | 0 | — |
case-16 | pass→pass | 6,406 | 4,659 | -27% | 1 | 1 | 0% | 1,230 | 2,529 | +106% | 0 | 0 | — |
case-17 | pass→pass | 4,920 | 25,659 | +422% | 1 | 1 | 0% | 1,111 | 3,942 | +255% | 0 | 0 | — |
case-18 | pass→pass | 13,228 | 10,256 | -22% | 1 | 1 | 0% | 2,400 | 3,485 | +45% | 0 | 0 | — |
case-19 | pass→fail | 16,077 | 21,143 | +32% | 1 | 1 | 0% | 3,615 | 6,234 | +72% | 0 | 0 | — |
case-20 | pass→fail | 25,996 | 25,587 | -2% | 1 | 1 | 0% | 4,497 | 6,631 | +47% | 0 | 0 | — |
case-21 | pass→pass | 13,431 | 12,223 | -9% | 1 | 1 | 0% | 2,398 | 3,838 | +60% | 0 | 0 | — |
case-22 | fail→pass | 11,327 | 19,071 | +68% | 1 | 1 | 0% | 1,901 | 5,561 | +193% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.