▸case-01 We ran a performance investigation on our Postgres database and backend API. CPU usage spikes to 98% during peak load, latency jumps from 50ms to 4500ms. CPU scaling rules were proposed. However, pg_stat_statements shows a single query `SELECT * FROM orders WHERE customer_id = $1 AND status = 'pending'` takes 92% of DB execution time, doing a full sequential scan of 15M rows. Synthesize these findings into prioritized engineering recommendations. | pass→pass | 14,546 | 11,818 | -19% | 1 | 1 | 0% | 2,347 | 2,077 | -12% | 0 | 0 | — |
▸case-02 During an SRE performance investigation of a Node.js microservice, heap memory consumption increases monotonically by 12MB/hour regardless of traffic volume, eventually triggering OOMKills every 24 hours. The team suggests increasing the Kubernetes pod memory limit from 1GiB to 4GiB and adding scheduled nightly pod restarts. Synthesize these findings into an actionable recommendation report. | pass→pass | 32,368 | 14,949 | -54% | 1 | 1 | 0% | 3,221 | 2,691 | -16% | 0 | 0 | — |
▸case-03 Investigation of a Redis cluster latency spike shows CPU core 1 pinned at 100% while other cores sit at 5% load. Slowlog reveals frequent calls to `KEYS pattern:*` every minute by a background analytics worker. The developer proposes upgrading to a larger multi-core Redis instance. Synthesize these findings into clear technical recommendations. | pass→pass | 11,867 | 11,283 | -5% | 1 | 1 | 0% | 2,052 | 1,828 | -11% | 0 | 0 | — |
▸case-04 Flamegraphs and APM traces for `/api/v1/dashboard` show 250 individual SQL queries executed per HTTP request to fetch order line items inside a loop. Latency is 1.8s for 100 requests/sec. The team wants to place a Varnish cache in front of the entire API endpoint. Synthesize these profiling results into actionable engineering steps. | pass→pass | 16,911 | 13,877 | -18% | 1 | 1 | 0% | 2,763 | 2,298 | -17% | 0 | 0 | — |
▸case-05 Investigation into 504 Gateway Timeouts under moderate load shows HTTP server thread count at 100% utilization, waiting on DB connection checkout from HikariCP (pool size 10). DB server CPU is at 15%. A proposal was submitted to double the HTTP server thread pool. Synthesize these telemetry findings into recommendations. | pass→pass | 14,146 | 11,047 | -22% | 1 | 1 | 0% | 2,296 | 1,823 | -21% | 0 | 0 | — |
▸case-06 Java service p99 latency spikes to 800ms every 30 seconds. JVM GC logs show Stop-The-World (STW) Young Generation collection pauses lasting 750ms due to high allocation rates of temporary short-lived JSON objects. Engineers recommend adding a Redis caching layer to offload the service. Synthesize these findings into specific performance recommendations. | pass→pass | 14,406 | 14,398 | -0% | 1 | 1 | 0% | 2,233 | 2,273 | +2% | 0 | 0 | — |
▸case-07 Tracing data shows microservice-to-microservice REST calls spending 300ms in network connect time out of 320ms total latency. Packet capture reveals external DNS lookup performed on every HTTP request without local caching or connection reuse. The team suggests refactoring the backend services to Go for better performance. Synthesize these findings into action items. | pass→pass | 10,552 | 10,860 | +3% | 1 | 1 | 0% | 1,702 | 1,777 | +4% | 0 | 0 | — |
▸case-20 Write a k6 load test script in JavaScript that sends 100 HTTP GET requests per second to `https://api.example.com/health` with a 2-second timeout and checks for HTTP status 200. | pass→pass | 6,583 | 6,401 | -3% | 1 | 1 | 0% | 1,381 | 1,336 | -3% | 0 | 0 | — |
▸case-08 Performance profiling of a cloud-hosted video metadata service indicates high latency during payload generation. Network interface metrics show outbound throughput hitting 1 Gbps (the instance limit), while CPU is at 20% and memory at 30%. The proposed fix is to rewrite payload generation in Rust. Synthesize these results into a clear recommendation plan. | pass→pass | 13,795 | 13,278 | -4% | 1 | 1 | 0% | 2,235 | 2,253 | +1% | 0 | 0 | — |
▸case-09 Kafka broker disk write latency increases from 2ms to 450ms during peak batch processing. CloudWatch metrics show VolumeQueueLength spiking and disk IOPS hitting the baseline AWS EBS gp2 limit of 3,000 IOPS. The team proposes increasing the Kafka batch size. Synthesize this finding into clear recommendations. | pass→pass | 14,371 | 12,934 | -10% | 1 | 1 | 0% | 2,384 | 2,292 | -4% | 0 | 0 | — |
▸case-10 Pprof mutex profiles for a concurrent Go service reveal 85% of goroutine execution time spent waiting on a single sync.RWMutex protecting a global counter map. Latency degraded 10x as concurrent users increased from 100 to 1,000. Synthesize these profiling results into technical recommendations. | pass→pass | 17,243 | 16,867 | -2% | 1 | 1 | 0% | 2,824 | 2,831 | +0% | 0 | 0 | — |
▸case-11 Website static asset load times are averaging 1.2s globally. CloudFront metrics show a cache hit ratio of 12%. Header inspection reveals `Cache-Control: private, no-store` set on all immutable JS/CSS bundle assets by the Webpack build server. The ops team wants to double CDN edge capacity. Synthesize these findings into an action plan. | pass→pass | 11,887 | 10,527 | -11% | 1 | 1 | 0% | 2,060 | 1,940 | -6% | 0 | 0 | — |
▸case-12 Lighthouse performance score for an e-commerce landing page is 32/100, with First Contentful Paint at 4.2s. Chrome DevTools Network trace shows a single 12MB monolithic JavaScript bundle downloaded before page render. The frontend team suggests switching from React to Svelte to solve this. Synthesize these profiling findings into concrete recommendations. | pass→pass | 16,358 | 14,686 | -10% | 1 | 1 | 0% | 2,759 | 2,472 | -10% | 0 | 0 | — |
▸case-13 Profiling microservice RPC calls shows 60% of total response time spent establishing new TLS connections (3-way TCP handshake + TLS 1.3 handshake) for every individual request. The backend team recommends rewriting the HTTP endpoints as GraphQL. Synthesize these findings into actionable recommendations. | pass→pass | 14,436 | 21,917 | +52% | 1 | 1 | 0% | 2,328 | 2,365 | +2% | 0 | 0 | — |
▸case-14 Kafka consumer group processing latency spikes intermittently every 15 minutes. Logs show frequent consumer rebalances triggered because `max.poll.interval.ms` (300,000ms) is exceeded when processing large individual message batches. Developers suggest adding 20 more consumer instances. Synthesize these investigation findings. | pass→pass | 11,778 | 11,697 | -1% | 1 | 1 | 0% | 1,925 | 1,949 | +1% | 0 | 0 | — |
▸case-15 An eBPF CPU profile on a Linux microservice reveals 70% of CPU cycles consumed inside string regex validation functions in an inline HTTP request logger filter. The infrastructure team proposes adding auto-scaling nodes. Synthesize these profiling results into performance recommendations. | pass→pass | 16,046 | 14,707 | -8% | 1 | 1 | 0% | 2,557 | 2,301 | -10% | 0 | 0 | — |
▸case-16 APM transaction breakdown shows user checkout requests taking 3.5 seconds. 3.2 seconds of that time is spent synchronously waiting for a third-party fraud detection HTTP API endpoint. The team proposes setting up auto-scaling for the checkout deployment. Synthesize these results into engineering recommendations. | pass→pass | 15,386 | 12,698 | -17% | 1 | 1 | 0% | 2,388 | 2,228 | -7% | 0 | 0 | — |
▸case-17 Application monitoring shows DB connection count growing steadily until hitting Postgres `max_connections` (200), causing all subsequent API requests to fail with connection errors. Code review shows unclosed DB rows in exception handling blocks. The team suggests raising `max_connections` to 2000. Synthesize these results. | pass→pass | 9,792 | 9,857 | +1% | 1 | 1 | 0% | 1,543 | 1,582 | +3% | 0 | 0 | — |
▸case-18 Go service CPU profile shows 65% of CPU time spent in `encoding/json.Marshal` when converting large slice structures (100k items) for WebSocket broadcast. Developers propose scaling out to 10 extra Kubernetes pods. Synthesize these findings into targeted performance recommendations. | fail→pass | 14,164 | 12,687 | -10% | 1 | 1 | 0% | 2,389 | 2,121 | -11% | 0 | 0 | — |
▸case-19 Linux server monitoring shows system load average spiking to 45 on a 4-core VM. Disk I/O wait is at 80%, swap usage is at 95%, and memory utilization is 99%. A developer suggests upgrading to a faster NVMe SSD drive. Synthesize these findings into remediation recommendations. | pass→pass | 31,553 | 11,331 | -64% | 1 | 1 | 0% | 2,186 | 1,846 | -16% | 0 | 0 | — |
▸case-21 Provide a Helm values.yaml configuration snippet for installing the Datadog Agent on a Kubernetes cluster with APM and process monitoring enabled. | pass→pass | 8,083 | 11,548 | +43% | 1 | 1 | 0% | 1,508 | 1,286 | -15% | 0 | 0 | — |
▸case-22 Provide the exact Linux `perf record` command syntax to capture CPU profile data at 99Hz for PID 1234 for 30 seconds and output to `perf.data`. | pass→pass | 3,697 | 5,303 | +43% | 1 | 1 | 0% | 678 | 948 | +40% | 0 | 0 | — |