▸case-01 Our team needs an incident response guide for handling Kubernetes cluster node memory exhaustion. Please construct a comprehensive runbook outlining required access prerequisites, quick reference commands, step-by-step mitigation actions, rollback mechanisms, diagnostic steps for troubleshooting, and escalation contacts. | fail→fail | 21,998 | 25,377 | +15% | 1 | 1 | 0% | 3,645 | 4,503 | +24% | 0 | 0 | — |
▸case-02 We need an operational runbook for executing a manual failover of a primary PostgreSQL 15 database instance on AWS EC2 to a standby replica. A developer suggested just giving a list of psql commands to switch roles. Generate a complete operational runbook for this procedure. | fail→fail | 23,040 | 27,565 | +20% | 1 | 1 | 0% | 3,980 | 5,167 | +30% | 0 | 0 | — |
▸case-03 Create a troubleshooting and mitigation runbook for high Elasticsearch node JVM heap pressure above 90%. Most engineers usually just restart the service without checking cluster health or logging outputs, leading to split-brain issues. Provide the operational procedure document. | fail→pass | 23,529 | 30,394 | +29% | 1 | 1 | 0% | 3,825 | 5,352 | +40% | 0 | 0 | — |
▸case-04 Write an operational procedure for clearing stuck RabbitMQ queues during high message backlog incidents. Engineers often forget what access they need and run commands blindly without checking queue state output. | fail→fail | 18,688 | 21,384 | +14% | 1 | 1 | 0% | 2,993 | 3,810 | +27% | 0 | 0 | — |
▸case-05 Draft an incident response runbook for an unannounced HashiCorp Vault cluster seal event. Engineers usually just search internal Slack for unseal keys without following a formal escalation or verification flow. | fail→pass | 29,420 | 24,007 | -18% | 1 | 1 | 0% | 3,526 | 3,940 | +12% | 0 | 0 | — |
▸case-06 Provide an operational runbook for emergency disk cleanup on Linux Docker container hosts when disk space hits 95%. Engineers tend to run destructive prune commands immediately without checking disk usage breakdown. | fail→fail | 16,920 | 19,673 | +16% | 1 | 1 | 0% | 2,881 | 3,518 | +22% | 0 | 0 | — |
▸case-07 Generate an operational runbook for removing a degraded member from a MongoDB replica set. Engineers usually try running admin commands from memory without checking sync status or specifying expected command output. | fail→pass | 20,272 | 18,201 | -10% | 1 | 1 | 0% | 3,399 | 3,373 | -1% | 0 | 0 | — |
▸case-08 Create a routine maintenance runbook for rotating TLS/SSL certificates on frontend NGINX web servers. Developers often treat this as a quick shell script request and skip documentation for rollback and escalation. | pass→pass | 18,523 | 23,858 | +29% | 1 | 1 | 0% | 2,994 | 4,481 | +50% | 0 | 0 | — |
▸case-09 Construct an incident runbook for handling an active BGP route leak on core datacenter routers. Network admins usually default to pasting raw router CLI snippets without structured prerequisites or expected output examples. | pass→pass | 25,003 | 26,981 | +8% | 1 | 1 | 0% | 4,115 | 4,662 | +13% | 0 | 0 | — |
▸case-10 Write a runbook for manual Sentinel-triggered failover of a Redis cluster. On-call engineers frequently perform failover without recording verification steps or defining escalation paths when Sentinel fails to elect a new master. | fail→fail | 24,181 | 23,933 | -1% | 1 | 1 | 0% | 4,213 | 4,510 | +7% | 0 | 0 | — |
▸case-11 Create an operational guide for fixing stuck cert-manager certificate renewals in a staging Kubernetes environment. Engineers often run resource deletion commands without logging output or specifying prerequisites. | fail→pass | 18,092 | 25,033 | +38% | 1 | 1 | 0% | 3,005 | 4,095 | +36% | 0 | 0 | — |
▸case-12 Draft an operational runbook for emergency rotation of compromised AWS KMS customer managed keys. Operators tend to just issue key deletion commands without documenting rollback or quick references. | fail→fail | 21,939 | 24,106 | +10% | 1 | 1 | 0% | 3,806 | 4,701 | +24% | 0 | 0 | — |
▸case-13 Generate a runbook for force-unlocking a corrupted or stuck Terraform state file in AWS S3 with DynamoDB locking. Engineers often execute unlock commands blindly without checking lock state or verifying integrity. | fail→pass | 19,375 | 23,041 | +19% | 1 | 1 | 0% | 3,224 | 4,237 | +31% | 0 | 0 | — |
▸case-14 Write an operational runbook for gracefully decommissioning a degraded Apache Kafka broker in a 5-node cluster. Operators often execute partition reassignment without checking copy-pasteable commands or verifying topic status outputs. | fail→pass | 20,734 | 21,472 | +4% | 1 | 1 | 0% | 3,833 | 4,317 | +13% | 0 | 0 | — |
▸case-15 Create an incident response runbook for adjusting Cloudflare Security Level and WAF rules during an ongoing Layer 7 HTTP flood attack. Support engineers often execute changes directly in the UI without documenting quick reference commands or rollback steps. | fail→pass | 23,022 | 22,439 | -3% | 1 | 1 | 0% | 3,986 | 4,356 | +9% | 0 | 0 | — |
▸case-16 Draft a runbook for resolving ArgoCD GitOps application sync status stuck in OutOfSync or Degraded state. Developers frequently attempt manual overrides without recording prerequisites or expected terminal output. | fail→fail | 20,904 | 27,018 | +29% | 1 | 1 | 0% | 3,481 | 4,556 | +31% | 0 | 0 | — |
▸case-17 Write an operational runbook for troubleshooting and recovering Ceph storage cluster degraded placement groups. Sysadmins often issue repair commands without verifying cluster health outputs or escalation contacts. | fail→fail | 24,322 | 26,913 | +11% | 1 | 1 | 0% | 3,879 | 4,662 | +20% | 0 | 0 | — |
▸case-18 Create an incident response runbook for recovering a MySQL Galera cluster from a split-brain condition where primary component status is lost. Engineers often attempt state file edits without documenting prerequisites or step verification. | fail→fail | 28,098 | 24,282 | -14% | 1 | 1 | 0% | 4,795 | 4,710 | -2% | 0 | 0 | — |
▸case-19 Provide a runbook for handling Datadog Agent high CPU consumption above 80% on Linux production servers. System administrators often issue service stop commands without documenting verification, troubleshooting, or escalation procedures. | fail→fail | 21,766 | 22,772 | +5% | 1 | 1 | 0% | 3,450 | 4,184 | +21% | 0 | 0 | — |
▸case-20 We are experiencing high CPU usage across our Kubernetes nodes. Please write a Prometheus PrometheusRule YAML manifest that triggers a HighNodeCPUAlert alert when node CPU utilization exceeds 85% for more than 5 minutes. | pass→pass | 8,450 | 8,549 | +1% | 1 | 1 | 0% | 1,464 | 1,888 | +29% | 0 | 0 | — |
▸case-21 Write a Terraform configuration file in HCL to provision an AWS S3 bucket with AES256 server-side encryption enabled and public access blocked. | pass→pass | 8,145 | 9,763 | +20% | 1 | 1 | 0% | 1,737 | 2,381 | +37% | 0 | 0 | — |
▸case-22 Provide a GitHub Actions workflow YAML file that triggers on push to the main branch, builds a Docker image from the repository Dockerfile, and pushes it to Amazon ECR. | pass→pass | 9,429 | 9,229 | -2% | 1 | 1 | 0% | 1,917 | 1,988 | +4% | 0 | 0 | — |