▸case-01 Create a Prometheus Rule YAML manifest that fires a critical alert named 'HighNodeMemoryUsage' whenever node_memory_MemAvailable_bytes divided by node_memory_MemTotal_bytes falls below 0.10 for more than 5 minutes. | pass→pass | 4,767 | 6,203 | +30% | 1 | 1 | 0% | 1,018 | 1,864 | +83% | 0 | 0 | — |
▸case-02 Write a Terraform HCL configuration using the aws_autoscaling_group and aws_launch_template resources to maintain a minimum of 2 and maximum of 10 EC2 instances across US East availability zones. | pass→pass | 10,963 | 12,803 | +17% | 1 | 1 | 0% | 2,562 | 3,541 | +38% | 0 | 0 | — |
▸case-03 Our primary PostgreSQL cluster crashed 10 minutes ago with error 'FATAL: remaining connection slots are reserved for non-replication superuser connections'. Help us troubleshoot the active outage and identify which microservice exhausted the connection pool. | pass→pass | 11,918 | 10,496 | -12% | 1 | 1 | 0% | 2,329 | 2,489 | +7% | 0 | 0 | — |
▸case-04 Our PostgreSQL database cluster stores 850 GB on a 1 TB SSD volume. Storage grew from 500 GB to 850 GB over the past 7 months. Management wants to know when we will hit 95% disk usage and what disk scaling options exist. Identify the projected depletion month and compare vertical vs horizontal scaling options with rough monthly cost impacts (assuming $0.10/GB-month for EBS gp3). | pass→pass | 14,335 | 17,427 | +22% | 1 | 1 | 0% | 2,800 | 4,119 | +47% | 0 | 0 | — |
▸case-05 Our Kubernetes web service running 10 pods on c6i.xlarge instances (4 vCPU, 8 GiB RAM each at $0.17/hr) reaches 88% average CPU during 4-hour daily peak traffic, causing latency spikes. Off-peak CPU is 15%. Evaluate whether vertical scaling (moving to c6i.2xlarge) or horizontal scaling (HPA from 10 to 20 pods) is better, including cost and bottleneck implications. | pass→pass | 17,226 | 18,661 | +8% | 1 | 1 | 0% | 3,521 | 4,441 | +26% | 0 | 0 | — |
▸case-06 An Redis cluster with 3 nodes (r6g.xlarge, 26.32 GiB RAM each) currently consumes 22 GiB RAM per node. Dataset grows by 1.5 GiB per node monthly. Determine when memory utilization will reach 90% threshold, risk of OOM eviction during Redis snapshotting, and recommend node resizing vs shard addition. | pass→pass | 14,777 | 17,279 | +17% | 1 | 1 | 0% | 2,977 | 4,022 | +35% | 0 | 0 | — |
▸case-07 Our video streaming platform transmits 400 TB of egress traffic per month via AWS CloudFront ($0.085/GB). Monthly egress increases by 15% month-over-month. Forecast network bandwidth costs for the next 3 months and evaluate multi-CDN or committed use pricing discounts. | pass→pass | 17,990 | 26,801 | +49% | 1 | 1 | 0% | 3,869 | 6,539 | +69% | 0 | 0 | — |
▸case-08 An e-commerce checkout microservice receives 1,200 requests/sec during regular hours and handles up to 3,500 req/sec on promotional days. Each container handles up to 150 req/sec safely. Forecast pod capacity required for an upcoming flash sale expected to reach 6,000 req/sec, and estimate compute costs assuming $0.04 per pod-hour. | pass→pass | 9,096 | 13,158 | +45% | 1 | 1 | 0% | 2,033 | 3,364 | +65% | 0 | 0 | — |
▸case-09 Perform a capacity analysis across compute, memory, disk, and network for a payment processing cluster: CPU at 40%, RAM at 82% (growing 4% monthly), Disk at 60% (growing 20 GB/mo on 500 GB total), Network egress at 15% line rate. Identify the immediate bottleneck resource and recommend a target roadmap. | fail→fail | 13,925 | 15,077 | +8% | 1 | 1 | 0% | 2,794 | 3,532 | +26% | 0 | 0 | — |
▸case-10 A batch processing service runs on 4 m5.2xlarge instances (8 vCPU, 32 GB RAM, $0.384/hr each). CPU is at 90% during 8-hour batch runs. We can either scale horizontally to 8 m5.2xlarge instances or scale vertically to 4 m5.4xlarge instances ($0.768/hr each). Analyze the throughput and cost trade-offs between both strategies. | pass→pass | 18,554 | 19,482 | +5% | 1 | 1 | 0% | 3,743 | 4,307 | +15% | 0 | 0 | — |
▸case-11 An Elasticsearch cluster ingests 250 GB of raw log data daily. Uncompressed indices with replicas require 600 GB of storage per day. Current storage array capacity is 30 TB with 18 TB used. Calculate when current storage will fill up and estimate monthly cost savings if hot retention is reduced from 30 days to 14 days (EBS storage at $0.08/GB-month). | pass→pass | 11,163 | 15,492 | +39% | 1 | 1 | 0% | 2,464 | 3,915 | +59% | 0 | 0 | — |
▸case-12 An AWS Lambda microservice handles 10,000,000 invocations monthly with an average duration of 450ms. Concurrent executions peak at 800 during top hour. Invocations grow 20% every month. Determine in how many months peak concurrency will exceed the default account limit of 1,000, and calculate estimated monthly Lambda execution costs ($0.0000166667 per GB-second, configured with 1024 MB RAM). | fail→pass | 23,464 | 15,037 | -36% | 1 | 1 | 0% | 2,742 | 4,316 | +57% | 0 | 0 | — |
▸case-13 A 3-node Apache Kafka cluster processes 45 MB/s ingress with a replication factor of 3 (total broker write throughput 135 MB/s). Disk IOPS are saturated at 85% utilization on gp2 volumes. Ingress data volume is projected to reach 80 MB/s next quarter. Recommend cluster scaling options (adding broker nodes vs upgrading to gp3/io2 storage) and quantify cost implications. | pass→pass | 24,180 | 26,621 | +10% | 1 | 1 | 0% | 4,954 | 5,262 | +6% | 0 | 0 | — |
▸case-14 An Aurora MySQL primary DB instance handles 8,000 read queries/sec and 1,500 write queries/sec. Max throughput per db.r6g.xlarge replica is 3,500 read QPS. Read query volume is growing by 1,000 QPS each month. Determine how many read replicas are needed now and in 6 months, along with monthly instance costs ($0.52/hr per replica). | fail→fail | 10,052 | 12,552 | +25% | 1 | 1 | 0% | 2,226 | 3,461 | +55% | 0 | 0 | — |
▸case-15 An internal Docker image registry stores 4.2 TB of layer data. CI/CD pipelines build and upload 120 GB of new image layers weekly. No image cleanup policy is currently enabled. Compute disk capacity growth over 6 months, predict when 10 TB storage limit will be breached, and estimate cost differential between increasing storage vs implementing tag retention policies. | pass→pass | 19,333 | 16,564 | -14% | 1 | 1 | 0% | 3,861 | 4,013 | +4% | 0 | 0 | — |
▸case-16 An API Gateway route handles 5,000 requests per minute during peak business hours. User onboarding traffic is expected to double overall API requests within 90 days. Current API rate limit throttling threshold is set to 8,000 RPM. Analyze capacity saturation timeline and recommend throttling quota adjustments and infrastructure scale. | fail→pass | 15,604 | 20,179 | +29% | 1 | 1 | 0% | 3,095 | 4,675 | +51% | 0 | 0 | — |
▸case-17 An analytics platform stores 150 TB of unstructured JSON in AWS S3 Standard ($0.023/GB-month). Access logs show 80% of data is never accessed after 30 days. Data volume increases by 10 TB monthly. Project 12-month storage costs without tiering vs transitioning data older than 30 days to S3 Standard-IA ($0.0125/GB-month). | pass→pass | 25,988 | 27,969 | +8% | 1 | 1 | 0% | 6,164 | 6,760 | +10% | 0 | 0 | — |
▸case-18 A Kubernetes cluster runs 45 microservice workloads requiring a total of 120 vCPUs and 480 GB RAM. Node pool consists of m5.2xlarge nodes (8 vCPU, 32 GB RAM at $0.384/hr). System components reserve 1 vCPU and 4 GB RAM per node. Calculate minimum node count required to maintain a 20% system capacity buffer for pod scheduling spikes and estimate monthly cluster node cost. | pass→pass | 13,303 | 13,346 | +0% | 1 | 1 | 0% | 2,885 | 3,556 | +23% | 0 | 0 | — |
▸case-19 An AWS Application Load Balancer handles 25 Mbps bandwidth, 1,200 active connections/sec, and 15,000 new connections/min. Traffic doubles during annual sales events. Evaluate ALB capacity against AWS LCU (Load Balancer Capacity Units) limits and calculate monthly LCU costs before and during peak traffic. | pass→pass | 17,787 | 18,088 | +2% | 1 | 1 | 0% | 4,198 | 4,668 | +11% | 0 | 0 | — |
▸case-20 An on-premises VMware vSphere cluster has 8 ESXi hosts, each with 2x 16-core CPUs (32 physical cores, 64 hyperthreads per host, 512 total vCPUs). The cluster hosts 120 virtual machines using 768 allocated vCPUs (1.5:1 overcommit). VMs grow by 10 VMs per quarter (6 vCPUs each). Forecast when vCPU overcommit ratio will exceed safe 3:1 threshold and evaluate host hardware addition timing. | fail→pass | 17,401 | 28,220 | +62% | 1 | 1 | 0% | 3,931 | 6,824 | +74% | 0 | 0 | — |
▸case-21 An SQS queue receives 500 job messages per minute. Each background worker container processes 15 jobs per minute. Queue depth currently grows by 50 messages per minute during 6-hour daily peaks. Calculate worker pod count needed to eliminate backlog accumulation and calculate compute cost difference using Fargate spot vs on-demand. | pass→pass | 15,049 | 20,821 | +38% | 1 | 1 | 0% | 3,233 | 5,025 | +55% | 0 | 0 | — |
▸case-22 A Memcached cache tier running 2 nodes (16 GB RAM each) experiences a cache hit ratio drop from 94% to 78% as key eviction rate hits 2,500 keys/sec. Total key namespace increases by 2 GB per week. Calculate required cache cluster RAM expansion to restore >92% hit ratio and estimate cost to scale node count or instance sizes. | pass→pass | 23,694 | 21,274 | -10% | 1 | 1 | 0% | 4,782 | 4,811 | +1% | 0 | 0 | — |