▸case-01 We are currently experiencing an issue where European merchants cannot process credit card payments at checkout; it started 20 minutes ago. We don't know the root cause yet, but users can switch to PayPal as a temporary workaround. Can you write a status page message, a notification email for affected merchants, and the follow-up posts for the remaining incident stages? | fail→fail | 15,379 | 13,986 | -9% | 1 | 1 | 0% | 1,823 | 2,285 | +25% | 0 | 0 | — |
▸case-02 Our database replication lag is causing dashboard analytics to be delayed by up to two hours for all accounts. We identified the memory leak, applied a patch, and are currently verifying recovery. There's no workaround right now. Please draft our status page announcement, an email template for our full customer base, and the sequence of future updates for subsequent phases. | fail→fail | 19,494 | 13,855 | -29% | 1 | 1 | 0% | 2,138 | 2,446 | +14% | 0 | 0 | — |
▸case-03 SSO logins are failing for enterprise accounts using Okta integration, blocking team members from accessing the workspace. We found an expired signing certificate on our auth server and are generating a replacement. No alternative login method is available right now. Please prepare the status page post, a direct broadcast email for enterprise administrators, and the follow-on status drafts for the upcoming lifecycle steps. | fail→fail | 18,756 | 10,540 | -44% | 1 | 1 | 0% | 2,109 | 2,435 | +15% | 0 | 0 | — |
▸case-04 Draft an internal blameless root cause analysis (RCA) postmortem document for engineering following yesterday's PostgreSQL database failure, including 5-why analysis and action items. | pass→fail | 29,916 | 26,456 | -12% | 1 | 1 | 0% | 3,924 | 4,260 | +9% | 0 | 0 | — |
▸case-05 Calculate the SLA uptime credit refund for enterprise customer Acme Corp following 47 minutes of service outage on their $10,000 monthly commitment tier. | pass→pass | 22,574 | 21,603 | -4% | 1 | 1 | 0% | 2,258 | 3,329 | +47% | 0 | 0 | — |
▸case-06 Configure PagerDuty escalation policies, alert thresholds, and on-call rotation schedules for P1 production incidents on the payment cluster. | pass→fail | 25,827 | 27,506 | +7% | 1 | 1 | 0% | 3,521 | 4,732 | +34% | 0 | 0 | — |
▸case-07 Our primary database cluster for Acme Cloud regional storage is failing health checks, resulting in 500 errors across all customer dashboards in EU-West. Engineers suspect either a network switch drop or a corrupted migration, but haven't confirmed either. Draft a status page announcement and an email update for affected accounts. A colleague suggests stating 'Database migration failure' in the status update to sound knowledgeable. | pass→pass | 17,614 | 14,451 | -18% | 1 | 1 | 0% | 1,796 | 2,556 | +42% | 0 | 0 | — |
▸case-08 We deployed a patch 15 minutes ago to fix latency spikes on the payments gateway for PayFast Inc., and metrics have returned to normal baseline. We want to monitor for another 45 minutes before declaring resolved. Prepare the status page update. The product manager wants to mark the post stage directly as Resolved right now to prevent customer alarm. | pass→pass | 14,956 | 15,649 | +5% | 1 | 1 | 0% | 1,432 | 2,370 | +66% | 0 | 0 | — |
▸case-09 The API routing service for Vortex Logistics is down, blocking all warehouse shipping requests. We're actively investigating. Provide the status page update and direct email notice. Support wants to put 'We will issue another update in 30 minutes' as the next update timing commitment. | fail→fail | 11,440 | 12,716 | +11% | 1 | 1 | 0% | 1,212 | 2,194 | +81% | 0 | 0 | — |
▸case-10 A major fiber line cut at our upstream data center provider has rendered the cloud console inaccessible for global users. The provider gave no estimate for physical repair completion. Write a status page notice and broadcast email. Marketing insists on promising customers 'Full service restoration within 2 hours'. | fail→pass | 16,984 | 16,695 | -2% | 1 | 1 | 0% | 1,877 | 2,719 | +45% | 0 | 0 | — |
▸case-11 Our background job processor crashed due to an out-of-memory exception in worker pool B, causing delayed invoice generation for enterprise accounts. Draft the customer status page update. The technical lead wants the post to start with 'Worker pool B ran out of memory, so invoices are delayed'. | fail→pass | 13,209 | 14,177 | +7% | 1 | 1 | 0% | 1,288 | 2,376 | +84% | 0 | 0 | — |
▸case-12 User authentication via OAuth is intermittent across all mobile apps. Draft an incident notification email for affected users. A customer support rep suggested adding apologies at the start of every paragraph ('We apologize for the inconvenience', 'We deeply regret this issue', 'Again, our sincere apologies') to show maximum empathy. | pass→pass | 15,342 | 16,599 | +8% | 1 | 1 | 0% | 1,690 | 2,830 | +67% | 0 | 0 | — |
▸case-13 Our main application server cluster is completely offline, preventing 100% of SaaS platform users from logging in or using the product. Write the status page notice. The VP of Comms suggested opening with 'Some users may be experiencing minor latency issues'. | pass→pass | 24,914 | 13,933 | -44% | 1 | 1 | 0% | 1,807 | 2,371 | +31% | 0 | 0 | — |
▸case-14 The web application file upload button is broken due to a front-end script error, preventing document submissions. Users can still upload files successfully by using our mobile app or sending them via email to uploads@example.com. Draft the status page message and customer email notice. | pass→pass | 12,636 | 13,500 | +7% | 1 | 1 | 0% | 1,368 | 2,324 | +70% | 0 | 0 | — |
▸case-15 The main relational database is corrupt and undergoing restoration from backups. All API endpoints return 503 errors and no alternative system or fallback exists. Prepare the status page update. Someone suggested adding 'Workaround: Please try refreshing your browser periodically'. | fail→pass | 16,446 | 15,649 | -5% | 1 | 1 | 0% | 1,725 | 2,670 | +55% | 0 | 0 | — |
▸case-16 Third-party SMS gateway failure is delaying two-factor authentication codes for log-ins on secure accounts. Draft an email notification to affected account holders. The security manager wants the subject line to read: 'CRITICAL SECURITY ALERT: 2FA SYSTEM BROKEN'. | fail→pass | 14,494 | 18,562 | +28% | 1 | 1 | 0% | 1,529 | 2,990 | +96% | 0 | 0 | — |
▸case-17 Our hosted search index service crashed, causing search results on all e-commerce stores to return empty lists. We identified a bad configuration deployment and are reverting it. A temporary workaround is using direct category navigation. Draft the broadcast email for store merchants. A team member suggested putting the apology first, followed by workaround, then status, then impact. | fail→pass | 7,473 | 16,232 | +117% | 1 | 1 | 0% | 1,304 | 2,731 | +109% | 0 | 0 | — |
▸case-18 We are currently in the Investigating stage for a DNS resolution failure affecting web app access. Draft the full update ladder pre-drafting follow-on posts for the subsequent incident stages so our team can publish them as work progresses. | pass→pass | 12,238 | 18,388 | +50% | 1 | 1 | 0% | 2,042 | 3,302 | +62% | 0 | 0 | — |
▸case-19 Single Sign-On (SSO) integration is failing for enterprise tier clients on Acme Platform. Draft a status page notice and direct notification email tailored specifically for enterprise account administrators. The writer wants to use casual slang like 'Oops! Our bad, our auth servers hit a snag!'. | fail→pass | 17,284 | 17,923 | +4% | 1 | 1 | 0% | 1,899 | 2,767 | +46% | 0 | 0 | — |
▸case-20 We found the memory leak in the microservice container cluster responsible for report generation, and developers are building a patch. Draft the status page update. A marketer wants to publish 'We have fixed the issue and reports are generating normally'. | pass→pass | 9,091 | 15,116 | +66% | 1 | 1 | 0% | 1,561 | 2,438 | +56% | 0 | 0 | — |
▸case-21 The real-time notification socket service is disconnected for all active sessions. We are investigating the network cluster split. Prepare the status page post. A team member drafted a 500-word essay with detailed architecture diagrams and multiple subheadings. | fail→pass | 15,953 | 14,456 | -9% | 1 | 1 | 0% | 1,579 | 2,489 | +58% | 0 | 0 | — |
▸case-22 After 1 hour of monitoring post-patch deployment, latency and error rates on the checkout API have remained at baseline zero for 60 consecutive minutes. Draft the final status page announcement ending the incident lifecycle. | pass→pass | 12,944 | 13,927 | +8% | 1 | 1 | 0% | 1,269 | 2,244 | +77% | 0 | 0 | — |
▸case-23 A CDN edge node disruption is causing slow image loading speeds across all customer blogs globally. We are investigating. Prepare a status page update for the general public customer base. A teammate suggested writing internal engineering abbreviations like 'BGP route flap on AS13335 transit link'. | pass→pass | 14,760 | 16,070 | +9% | 1 | 1 | 0% | 1,289 | 2,406 | +87% | 0 | 0 | — |
▸case-24 Database connection pool exhaustion is causing intermittent 504 gateway timeouts on the customer portal. We identified the runaway query and are killing long-running connections. Prepare status page comms and email comms for impacted portal users. | fail→pass | 16,298 | 14,423 | -12% | 1 | 1 | 0% | 1,820 | 2,299 | +26% | 0 | 0 | — |
▸case-25 The export to CSV feature on user analytics tables is failing with error 500. We identified a missing dependency in the build pipeline. Prepare a direct email notice for affected users. A reviewer suggested starting the email with two paragraphs of corporate fluff about our mission to provide world-class reliability and enterprise-grade infrastructure before mentioning the bug. | fail→pass | 13,602 | 14,589 | +7% | 1 | 1 | 0% | 1,406 | 2,398 | +71% | 0 | 0 | — |