▸case-01 We are launching a critical customer-facing API service and need to set up formal reliability targets. Can you guide us through defining measurable indicators, setting reasonable objectives, calculating our allowed failure allowance, and creating a step-by-step implementation plan with verification strategies? | fail→fail | 23,092 | 22,651 | -2% | 1 | 1 | 0% | 3,425 | 3,947 | +15% | 0 | 0 | — |
▸case-02 Our engineering team wants to transition from basic infrastructure alerts to service-level objective monitoring using Prometheus. Please provide actionable implementation steps, including how to structure recording metrics, build alerting rules based on target consumption rates, and establish review cadences for our team. | fail→fail | 30,608 | 23,327 | -24% | 1 | 1 | 0% | 4,038 | 4,281 | +6% | 0 | 0 | — |
▸case-03 We need to design a operational reliability framework for our e-commerce checkout service. Please provide a complete guide covering the structural hierarchy between customer agreements, internal targets, and actual metrics, along with operational policies for managing our error allowance and validating the outcomes. | fail→fail | 39,217 | 29,434 | -25% | 1 | 1 | 0% | 7,285 | 5,427 | -26% | 0 | 0 | — |
▸case-04 Our product team is confused about the difference between SLAs, SLOs, and SLIs for our online payment gateway. Can you explain how these three terms relate hierarchically from customer-facing commitments down to raw metrics? | fail→fail | 20,397 | 14,938 | -27% | 1 | 1 | 0% | 2,750 | 2,314 | -16% | 0 | 0 | — |
▸case-05 Our REST API processed 10,000,000 requests last month with an internal SLO target of 99.9% availability. How many failed requests can we tolerate before our error budget is completely exhausted? | fail→fail | 9,147 | 4,110 | -55% | 1 | 1 | 0% | 869 | 1,248 | +44% | 0 | 0 | — |
▸case-06 Our microservice team has burned through 100% of its quarterly error budget in the first month. What operational policy actions should be triggered when an error budget is fully consumed? | fail→fail | 19,375 | 18,707 | -3% | 1 | 1 | 0% | 2,296 | 2,601 | +13% | 0 | 0 | — |
▸case-07 We want to measure availability for our HTTP API microservice using server response codes. How should we define an availability SLI as a mathematical ratio of successful requests to total requests? | fail→fail | 22,867 | 14,043 | -39% | 1 | 1 | 0% | 2,982 | 2,080 | -30% | 0 | 0 | — |
▸case-08 Our frontend team wants an SLI to measure endpoint response times for our search service. How should a latency SLI be structured to capture user-perceived performance rather than average response times? | fail→fail | 14,830 | 16,690 | +13% | 1 | 1 | 0% | 2,486 | 2,490 | +0% | 0 | 0 | — |
▸case-09 We are writing Prometheus configuration for a gRPC user service. How do we construct a recording rule to compute a 5-minute rolling HTTP request error rate from raw metric counters? | fail→fail | 15,244 | 16,972 | +11% | 1 | 1 | 0% | 2,896 | 2,857 | -1% | 0 | 0 | — |
▸case-10 Instead of triggering alerts immediately on a single failed request, we want to build alert rules based on error budget burn rate over multiple evaluation windows. What strategy should we use for alert thresholds? | fail→fail | 18,711 | 18,609 | -1% | 1 | 1 | 0% | 2,465 | 3,077 | +25% | 0 | 0 | — |
▸case-11 We are establishing a weekly operational meeting for our SRE team to review reliability. What key metrics and burn rate signals should be examined during a weekly review? | fail→fail | 16,983 | 13,860 | -18% | 1 | 1 | 0% | 2,740 | 2,815 | +3% | 0 | 0 | — |
▸case-12 We need an operational template for our monthly service review. What long-term trends and budget consumption figures should be evaluated at the end of each month? | fail→fail | 23,096 | 19,952 | -14% | 1 | 1 | 0% | 2,933 | 2,959 | +1% | 0 | 0 | — |
▸case-13 At the end of Q2, our leadership wants to evaluate our existing reliability targets against business strategy. What actions should be taken during a quarterly SLO review? | fail→fail | 14,808 | 15,447 | +4% | 1 | 1 | 0% | 2,348 | 2,045 | -13% | 0 | 0 | — |
▸case-14 Our team is tempted to set a 99.999% ('five nines') availability target for our internal admin dashboard. How should we determine whether this target is appropriate or counterproductive? | fail→fail | 17,096 | 15,794 | -8% | 1 | 1 | 0% | 2,824 | 3,007 | +6% | 0 | 0 | — |
▸case-15 We have defined Prometheus alert rules for our service's error budget burn rates. How can we formally verify and test that our Prometheus alerting configuration is syntactically valid and firing correctly? | fail→fail | 16,992 | 19,429 | +14% | 1 | 1 | 0% | 3,128 | 3,182 | +2% | 0 | 0 | — |
▸case-16 We run a nightly data pipeline that processes batch files. How should we measure the reliability of this batch pipeline using job completion time and freshness SLIs? | fail→fail | 23,064 | 24,173 | +5% | 1 | 1 | 0% | 3,260 | 3,716 | +14% | 0 | 0 | — |
▸case-17 When setting up burn rate alerts for a 30-day SLO of 99.9%, how do short lookback windows (e.g., 5 minutes) and long lookback windows (e.g., 1 hour) combine to avoid false positives and delayed alerts? | fail→fail | 15,894 | 18,906 | +19% | 1 | 1 | 0% | 2,900 | 3,170 | +9% | 0 | 0 | — |
▸case-18 Our payment processing service has an SLO stating that 95% of requests must complete in under 200ms over a 30-day window. If we receive 2,000,000 requests in 30 days, what is our latency error budget? | fail→fail | 9,389 | 9,282 | -1% | 1 | 1 | 0% | 853 | 1,305 | +53% | 0 | 0 | — |
▸case-19 Our order entry service suffered a 10-minute outage, but because our monthly SLO is 99.0% and overall uptime remains at 99.4%, our error budget is not exhausted. Should feature deployments be frozen? | fail→fail | 16,147 | 16,097 | -0% | 1 | 1 | 0% | 1,785 | 2,265 | +27% | 0 | 0 | — |
▸case-20 We operate an Apache Kafka event streaming cluster. What is an appropriate user-centric SLI to measure consumer processing lag for our event processors? | fail→fail | 14,377 | 17,445 | +21% | 1 | 1 | 0% | 2,418 | 2,546 | +5% | 0 | 0 | — |
▸case-21 Our customer agreement promises 99.5% availability with financial penalties if breached. What internal SLO threshold should engineering target to ensure we do not breach our external SLA? | fail→fail | 18,379 | 11,014 | -40% | 1 | 1 | 0% | 2,332 | 2,359 | +1% | 0 | 0 | — |
▸case-22 We just deployed error budget tracking dashboards in Grafana connected to Prometheus. What steps should our SRE team perform to verify that the error budget calculation is working accurately in practice? | fail→fail | 22,706 | 21,014 | -7% | 1 | 1 | 0% | 3,089 | 2,980 | -4% | 0 | 0 | — |
▸case-23 We want to auto-scale our Kubernetes deployment based on high CPU and memory utilization. Can you write an HPA resource manifest targeting 70% CPU and 80% RAM utilization? | pass→pass | 7,353 | 11,917 | +62% | 1 | 1 | 0% | 1,500 | 1,816 | +21% | 0 | 0 | — |
▸case-24 Our DevOps team needs to configure a PagerDuty escalation policy so that unacknowledged server alerts escalate from primary on-call to secondary on-call after 15 minutes. Can you show us how to set up this escalation policy in PagerDuty? | pass→pass | 16,433 | 11,220 | -32% | 1 | 1 | 0% | 2,289 | 2,666 | +16% | 0 | 0 | — |
▸case-25 Our production database suffered a 30-minute outage yesterday due to deadlocks during a migration. Can you provide a post-mortem incident template using the 5 Whys analysis framework to document the root cause? | pass→pass | 21,850 | 19,849 | -9% | 1 | 1 | 0% | 2,844 | 2,943 | +3% | 0 | 0 | — |