Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design production-ready observability strategies combining metrics, logs, and traces. Includes SLI/SLO design, golden-signals monitoring, alert optimization. Use when adding observability to a new service, refactoring alerting that is too noisy, or designing an SLO program before scaling production load.
.claude/skills/alirezarezvani-observability-designer/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flashlowest | 96% | 54 |
| gemini-3.1-pro-preview | 100% | 1 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 138% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 142% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 71% | 0% |
Category: Engineering Tier: POWERFUL Description: Design comprehensive observability strategies for production systems including SLI/SLO frameworks, alerting optimization, and dashboard generation.
Observability Designer creates production-ready dashboards, alert configurations, and monitoring strategies across the three pillars (metrics, logs, traces).
When NOT to use → slo-architect. For SLO/SLI design with error-budget math, multi-window burn-rate alerting thresholds, and SLO review gates, route to slo-architect — it is the authoritative skill for that half. This skill's slo_designer.py produces a quick scaffold only. This skill's lane: dashboards (dashboard_generator.py) and alert-noise reduction (alert_optimizer.py).
bash# Dashboard spec (Grafana JSON + docs) for a service python3 scripts/dashboard_generator.py --service-type api --name payments --criticality critical --role sre --format grafana -o dashboard.json --doc-output dashboard.md # Analyze an existing alert config for noise, duplicates, and coverage gaps python3 scripts/alert_optimizer.py --input alerts.json --analyze-only --report alert_report.json # ...then emit the optimized config once the report is reviewed: python3 scripts/alert_optimizer.py --input alerts.json --output alerts_optimized.json # Quick SLO scaffold (hand off to slo-architect for the real error-budget work) python3 scripts/slo_designer.py --service-type api --criticality high --user-facing true --service-name payments -o slo_scaffold.json
Verification loop: after deploying optimized alerts, track the report's noise metrics for one on-call rotation — if the actionable-alert ratio didn't improve, re-run --analyze-only against the live config and iterate. Import the generated dashboard into Grafana and confirm every golden-signal panel renders with live data before closing the task.
This skill includes three powerful Python scripts for comprehensive observability design:
slo_designer.py)Generates complete SLI/SLO frameworks based on service characteristics:
alert_optimizer.py)Analyzes and optimizes existing alert configurations:
dashboard_generator.py)Creates comprehensive dashboard specifications:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 23,900 | 3,410 | -86% | 1 | 1 | 0% | 6,214 | 2,886 | -54% | 0 | 0 | — |
case-02 | fail→pass | 12,372 | 11,927 | -4% | 1 | 1 | 0% | 2,910 | 5,395 | +85% | 0 | 0 | — |
case-03 | fail→pass | 7,829 | 5,140 | -34% | 1 | 1 | 0% | 1,556 | 3,655 | +135% | 0 | 0 | — |
case-04 | fail→pass | 6,900 | 2,800 | -59% | 1 | 1 | 0% | 1,341 | 3,187 | +138% | 0 | 0 | — |
case-05 | fail→pass | 7,201 | 3,866 | -46% | 1 | 1 | 0% | 1,453 | 3,517 | +142% | 0 | 0 | — |
case-06 | pass→pass | 14,269 | 10,203 | -28% | 1 | 1 | 0% | 2,745 | 4,752 | +73% | 0 | 0 | — |
case-07 | fail→fail | 27,776 | 32,973 | +19% | 1 | 1 | 0% | 6,214 | 8,889 | +43% | 0 | 0 | — |
case-08 | fail→fail | 42,405 | 19,244 | -55% | 1 | 1 | 0% | 4,791 | 7,044 | +47% | 0 | 0 | — |
case-09 | fail→fail | 25,817 | 10,373 | -60% | 1 | 1 | 0% | 1,838 | 4,553 | +148% | 0 | 0 | — |
case-10 | fail→pass | 16,745 | 12,672 | -24% | 1 | 1 | 0% | 2,979 | 5,091 | +71% | 0 | 0 | — |
case-11 | pass→pass | 9,796 | 8,649 | -12% | 1 | 1 | 0% | 1,854 | 4,280 | +131% | 0 | 0 | — |
case-12 | fail→pass | 12,973 | 8,692 | -33% | 1 | 1 | 0% | 2,315 | 4,102 | +77% | 0 | 0 | — |
case-13 | pass→pass | 11,595 | 11,160 | -4% | 1 | 1 | 0% | 2,292 | 4,690 | +105% | 0 | 0 | — |
case-14 | fail→pass | 13,796 | 9,403 | -32% | 1 | 1 | 0% | 2,384 | 4,269 | +79% | 0 | 0 | — |
case-15 | pass→pass | 7,064 | 3,568 | -49% | 1 | 1 | 0% | 1,218 | 3,277 | +169% | 0 | 0 | — |
case-16 | pass→pass | 11,540 | 9,553 | -17% | 1 | 1 | 0% | 2,413 | 4,648 | +93% | 0 | 0 | — |
case-17 | pass→pass | 10,628 | 9,641 | -9% | 1 | 1 | 0% | 1,989 | 4,407 | +122% | 0 | 0 | — |
case-18 | pass→pass | 12,680 | 12,954 | +2% | 1 | 1 | 0% | 2,602 | 5,136 | +97% | 0 | 0 | — |
case-19 | fail→pass | 11,640 | 14,723 | +26% | 1 | 1 | 0% | 2,034 | 5,109 | +151% | 0 | 0 | — |
case-20 | fail→pass | 5,835 | 1,964 | -66% | 1 | 1 | 0% | 1,187 | 3,071 | +159% | 0 | 0 | — |
case-21 | pass→pass | 13,766 | 17,974 | +31% | 1 | 1 | 0% | 2,467 | 5,839 | +137% | 0 | 0 | — |
case-22 | pass→pass | 10,813 | 3,762 | -65% | 1 | 1 | 0% | 2,018 | 3,201 | +59% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.