▸case-01 Can you help us diagnose and fix our web application's slow loading times and layout instability? Please evaluate http://localhost:3000 and provide a structured plan covering client-side rendering metric targets, bundle reduction strategies, and automated audit execution steps. | fail→pass | 20,914 | 16,539 | -21% | 1 | 1 | 0% | 3,894 | 3,409 | -12% | 0 | 0 | — |
▸case-02 Our backend services are experiencing high response latencies and database bottlenecks during traffic spikes. Could you create a performance improvement proposal that outlines database query fixes, multi-tier caching options, microservice distributed tracing, and container resource tuning? | fail→fail | 24,360 | 41,777 | +71% | 1 | 1 | 0% | 4,021 | 8,494 | +111% | 0 | 0 | — |
▸case-03 We are preparing our React web application for Google Search ranking audits. Our team is debating whether an Interaction to Next Paint (INP) score of 350ms is acceptable and how to handle heavy JavaScript tasks during click events. Provide recommended target thresholds and code execution strategies. | pass→pass | 18,353 | 16,687 | -9% | 1 | 1 | 0% | 3,347 | 3,849 | +15% | 0 | 0 | — |
▸case-04 Our media portal is experiencing visual page jumps while loading articles with embedded images and video embeds. Engineers suggest using CSS animations to smooth out layout shifts. Recommend target thresholds for visual stability and the correct technical remedy for media elements. | pass→pass | 15,546 | 10,928 | -30% | 1 | 1 | 0% | 2,264 | 2,778 | +23% | 0 | 0 | — |
▸case-05 Our single page web application's JavaScript bundle has grown to 4MB. Developers are considering converting CommonJS modules to ECMAScript modules and dynamically loading route components. Explain how ECMAScript imports and dynamic imports optimize bundle performance. | pass→pass | 14,496 | 16,930 | +17% | 1 | 1 | 0% | 2,408 | 3,509 | +46% | 0 | 0 | — |
▸case-06 We are configuring our Nginx front-end reverse proxy compression settings for static assets like JS and CSS files. The operations team plans to use standard gzip level 6 compression only. What modern compression format should be prioritized alongside gzip? | pass→pass | 9,225 | 8,102 | -12% | 1 | 1 | 0% | 1,544 | 2,394 | +55% | 0 | 0 | — |
▸case-07 Our PostgreSQL database logs show slow query warnings during high-concurrency read requests. The junior developer suggests adding indexes on every column in the table without running execution analysis. What exact database diagnostic tool and indexing strategy should be employed? | fail→fail | 16,300 | 19,732 | +21% | 1 | 1 | 0% | 2,871 | 4,019 | +40% | 0 | 0 | — |
▸case-08 When users submit a report generation form on our web application, the HTTP request hangs for 15 seconds while synchronously rendering PDF documents. What architectural pattern and queue options should be used to offload this heavy work? | pass→pass | 14,001 | 17,194 | +23% | 1 | 1 | 0% | 2,312 | 3,894 | +68% | 0 | 0 | — |
▸case-09 Our microservices on Kubernetes keep getting killed due to OOMKilled errors during unexpected traffic bursts, but during normal hours they consume less than 10% of their allocated CPU and memory. Engineers are debating manual pod scaling via HPA alone. How should resource limits be tuned? | pass→fail | 18,969 | 19,718 | +4% | 1 | 1 | 0% | 3,036 | 3,884 | +28% | 0 | 0 | — |
▸case-10 We are establishing standard Grafana dashboard metrics for microservices using OpenTelemetry. The team wants to track CPU utilization, disk IOPS, and memory percentage as our main dashboard metrics. What standard set of application metrics should be monitored instead? | pass→fail | 17,360 | 17,815 | +3% | 1 | 1 | 0% | 2,900 | 3,289 | +13% | 0 | 0 | — |
▸case-11 When analyzing failure logs in Grafana Loki during a distributed transaction across 5 microservices, developers struggle to group related log lines together. What exact technique should be implemented in OpenTelemetry logging? | pass→pass | 9,548 | 10,324 | +8% | 1 | 1 | 0% | 1,794 | 2,511 | +40% | 0 | 0 | — |
▸case-12 We need to verify whether our Node.js backend application accumulates uncollected garbage memory over a 48-hour period prior to a major promotional event. What specific category of load test should be conducted? | pass→pass | 6,560 | 5,028 | -23% | 1 | 1 | 0% | 1,072 | 1,517 | +42% | 0 | 0 | — |
▸case-13 Our QA team wants to find the exact point of failure where our API gateway crashes under excessive traffic beyond normal operational limits. Should we run a standard load test, a soak test, or a stress test? | pass→pass | 8,715 | 9,122 | +5% | 1 | 1 | 0% | 1,462 | 2,293 | +57% | 0 | 0 | — |
▸case-14 After running a k6 test on our staging environment, the test runner produces summary output metrics. How should engineering evaluate whether the test results pass performance regression checks? | pass→pass | 15,015 | 15,837 | +5% | 1 | 1 | 0% | 2,523 | 3,600 | +43% | 0 | 0 | — |
▸case-15 In our SRE handbook setup, the product manager asks for a definition of Service Level Indicators (SLI). How should SLI be defined in contrast to SLO and error budget? | pass→pass | 11,663 | 13,420 | +15% | 1 | 1 | 0% | 2,022 | 3,091 | +53% | 0 | 0 | — |
▸case-16 Our team has defined a target of 99.9% successful HTTP responses for our payment service. How does this target relate to an Error Budget, and what operational rule applies when the budget is exhausted? | pass→pass | 9,695 | 10,868 | +12% | 1 | 1 | 0% | 1,658 | 2,605 | +57% | 0 | 0 | — |
▸case-17 We are architecting a high-throughput content publishing platform. Design the complete caching layers from the client browser down to the database cache in the exact sequence recommended for peak efficiency. | pass→pass | 26,123 | 32,701 | +25% | 1 | 1 | 0% | 4,380 | 5,725 | +31% | 0 | 0 | — |
▸case-18 We are setting up an automated performance review protocol for a client application served at http://localhost:3000. Outline the execution workflow starting from initial scanning down through bundle inspection and Web Vitals verification. | fail→pass | 16,557 | 16,678 | +1% | 1 | 1 | 0% | 2,879 | 3,873 | +35% | 0 | 0 | — |
▸case-19 An e-commerce site has a Largest Contentful Paint (LCP) measurement of 4.2 seconds on mobile devices. The marketing team suggested changing font styles. What are the recommended optimization tactics to lower LCP below target thresholds? | pass→pass | 16,662 | 19,575 | +17% | 1 | 1 | 0% | 2,836 | 3,735 | +32% | 0 | 0 | — |
▸case-20 We are selecting developer-friendly load testing tooling for our DevOps pipeline that supports scriptable traffic scenarios and CLI integration without heavy Java GUIs. What open-source tools are recommended for this purpose? | pass→pass | 16,829 | 20,378 | +21% | 1 | 1 | 0% | 2,783 | 3,920 | +41% | 0 | 0 | — |
▸case-21 Can you generate a full Kubernetes Deployment YAML manifest for a Python FastAPI application including container ports, liveness probes, readiness probes, and volume mounts? | pass→pass | 12,205 | 14,779 | +21% | 1 | 1 | 0% | 2,635 | 3,734 | +42% | 0 | 0 | — |
▸case-22 Can you write a React functional component using TypeScript and Tailwind CSS for a credit card payment input form with field validation for card number, expiration date, and CVC? | pass→pass | 28,784 | 34,535 | +20% | 1 | 1 | 0% | 6,357 | 8,560 | +35% | 0 | 0 | — |
▸case-23 Can you write a PostgreSQL SQL migration script to create a range-partitioned table for system audit logs based on created_at timestamp columns, including monthly child partitions? | pass→pass | 16,924 | 13,850 | -18% | 1 | 1 | 0% | 3,449 | 3,410 | -1% | 0 | 0 | — |