▸case-01 We are analyzing transaction amounts in a skew-right e-commerce dataset to flag outliers. A team member suggested using standard Z-score with a threshold of 3, but transaction values have extreme right-skew with heavy tails. How should we calculate the upper bound threshold to detect anomalies without assuming a normal distribution? | fail→pass | 20,132 | 22,713 | +13% | 1 | 1 | 0% | 2,705 | 3,846 | +42% | 0 | 0 | — |
▸case-02 We need to detect sudden daily traffic anomalies in web server logs that exhibit strong weekly seasonality and an upward multi-month trend. Someone proposed using a simple 7-day rolling average threshold, but holiday spikes and underlying growth produce false positives. Which time series decomposition technique isolates seasonal and trend components before checking residuals? | pass→pass | 17,678 | 19,769 | +12% | 1 | 1 | 0% | 2,183 | 3,118 | +43% | 0 | 0 | — |
▸case-03 In PostgreSQL, we want to write a query that calculates a 30-day rolling mean and 30-day rolling standard deviation for daily server response times to flag days where response time exceeds 3 standard deviations above the rolling mean. A junior dev suggests using a GROUP BY clause on daily records. Write the SQL query snippet using the appropriate window function clause. | fail→fail | 14,321 | 15,139 | +6% | 1 | 1 | 0% | 1,878 | 2,330 | +24% | 0 | 0 | — |
▸case-04 We are building an unsupervised anomaly detector for high-dimensional server metrics (CPU, RAM, network I/O, disk operations) where labeled anomaly data is unavailable. A developer suggests using K-Means clustering and flagging points far from cluster centroids, but distance metrics degrade in high dimensions and cluster counts are arbitrary. Which tree-based ensemble algorithm explicitly isolates anomalies by randomly partitioning feature space? | pass→pass | 9,477 | 15,928 | +68% | 1 | 1 | 0% | 851 | 2,374 | +179% | 0 | 0 | — |
▸case-05 We have spatial coordinates of delivery pings and want to identify rogue pings that lie in low-density regions relative to their immediate neighbors, even when overall point density varies across the city. Global distance thresholds miss local anomalies in sparse suburbs or flag dense downtown points. Which density-based local anomaly detection algorithm measures the local deviation of a given data point with respect to its nearest neighbors? | pass→pass | 14,339 | 17,836 | +24% | 1 | 1 | 0% | 1,657 | 2,870 | +73% | 0 | 0 | — |
▸case-06 In Prometheus monitoring, we want to generate an alert when HTTP request rates deviate significantly from predicted patterns over a 1-hour window. A teammate suggested setting a hard metric threshold of http_requests_total > 10000. What PromQL function allows dynamic trend forecasting based on linear regression over time series data? | pass→pass | 9,877 | 13,619 | +38% | 1 | 1 | 0% | 927 | 2,047 | +121% | 0 | 0 | — |
▸case-07 We are designing an executive BI dashboard in Tableau to display regional sales metrics and draw immediate visual attention to anomalous stores whose quarterly sales fall outside expected variance. A designer suggests coloring all scatter plot points using a continuous rainbow gradient based on revenue. How should anomaly points be visually encoded according to data visualization best practices? | pass→pass | 21,743 | 18,956 | -13% | 1 | 1 | 0% | 2,602 | 2,644 | +2% | 0 | 0 | — |
▸case-08 We need to detect anomalies in hourly electricity consumption metrics that exhibit both daily (24-hour) and weekly (168-hour) seasonal cycles. A basic single exponential smoothing model was proposed, but it fails because it lacks trend and seasonality parameters. Which triple exponential smoothing algorithm accounts for level, trend, and seasonal components? | pass→pass | 12,344 | 18,207 | +47% | 1 | 1 | 0% | 1,310 | 2,893 | +121% | 0 | 0 | — |
▸case-09 We have geographical coordinates of fraud activity and want to discover spatial clusters of normal activity while automatically classifying points in low-density noise regions as anomalies. A teammate suggests K-Means with K=5, but K-Means forces every point into a cluster and requires specifying K upfront. Which density-based clustering algorithm identifies noise points as cluster outliers without requiring predefined cluster counts? | fail→pass | 12,124 | 19,716 | +63% | 1 | 1 | 0% | 1,314 | 3,310 | +152% | 0 | 0 | — |
▸case-10 We have 50 sensor streams from industrial turbines and want to detect complex non-linear multivariate anomalies. An engineer suggests training a supervised deep neural network classifier, but we have zero labeled failure cases. Which neural network architecture reconstructs input vectors through a bottleneck layer and flags anomalies based on high reconstruction error? | pass→pass | 11,842 | 19,573 | +65% | 1 | 1 | 0% | 1,222 | 3,283 | +169% | 0 | 0 | — |
▸case-11 We are analyzing laboratory assay measurement results (N=25) that are known to follow a normal distribution. We suspect exactly one extreme high value is a measurement artifact and wish to perform a formal hypothesis test to confirm whether it is a statistical outlier. Someone suggested running a Chi-Square goodness-of-fit test. Which specific univariate hypothesis test evaluates whether the maximum or minimum value in a normally distributed sample is an outlier? | pass→pass | 12,241 | 14,027 | +15% | 1 | 1 | 0% | 1,317 | 2,055 | +56% | 0 | 0 | — |
▸case-12 We are building a fraud detector where we possess abundant clean, non-fraudulent transaction logs, but virtually no fraudulent examples. We want to fit a boundary around normal behavior in feature space and flag any transaction falling outside that boundary. A colleague suggests logistic regression, but logistic regression requires positive and negative class labels. Which kernel-based support vector algorithm is designed specifically for one-class boundary estimation? | pass→pass | 6,416 | 15,396 | +140% | 1 | 1 | 0% | 1,098 | 2,304 | +110% | 0 | 0 | — |
▸case-13 In a financial dataset with two features (Income and Credit Limit) that are strongly correlated, a point has individual feature values that are each within normal 1D ranges, but the combination of values is extremely improbable given the covariance structure. Calculating Euclidean distance from the mean fails to catch this point because Euclidean distance assumes spherical, uncorrelated distributions. Which distance metric accounts for feature covariance when measuring distance from the distribution mean? | pass→pass | 6,040 | 9,508 | +57% | 1 | 1 | 0% | 1,056 | 2,080 | +97% | 0 | 0 | — |
▸case-14 We want to write a SQL query to identify sudden step-change anomalies where daily telemetry metric values jump by more than 200% compared to the immediate prior day. A colleague suggested joining the table to itself on id = id + 1, which breaks if there are missing dates. What SQL window function retrieves the previous row's value chronologically without relying on contiguous integer IDs? | pass→pass | 9,380 | 9,275 | -1% | 1 | 1 | 0% | 1,813 | 2,174 | +20% | 0 | 0 | — |
▸case-15 In a smart home energy dataset, a temperature reading of 85 degrees Fahrenheit is perfectly normal during July afternoon, but highly anomalous when recorded at 3 AM in January. An analyst suggests applying a global threshold across the entire dataset (temp > 80). What type of anomaly is this, and how should data analytics structure the evaluation to avoid false positives? | pass→pass | 12,449 | 20,337 | +63% | 1 | 1 | 0% | 2,185 | 3,158 | +45% | 0 | 0 | — |
▸case-16 In a manufacturing quality control analytics dashboard, we want to monitor batch thickness measurements over time and detect statistical out-of-control conditions before products fail specs. A supervisor suggested using a simple min-max threshold visual. What statistical process control chart framework uses Upper and Lower Control Limits at 3 standard deviations and Western Electric rules to detect systematic process shifts? | pass→pass | 13,898 | 28,510 | +105% | 1 | 1 | 0% | 1,620 | 5,121 | +216% | 0 | 0 | — |
▸case-17 We are evaluating an anomaly detection model where anomalies represent 0.01% of all network traffic logs. The model achieves 99.9% accuracy and an ROC-AUC of 0.98, but when deployed, security analysts are overwhelmed with false alarms. A developer wants to rely solely on ROC-AUC for tuning thresholds. Which evaluation metric curve is more informative than ROC-AUC when evaluating performance on highly imbalanced anomaly datasets? | pass→pass | 12,661 | 9,268 | -27% | 1 | 1 | 0% | 1,408 | 2,051 | +46% | 0 | 0 | — |
▸case-18 We need to write a PromQL query for Prometheus that flags metrics when current CPU usage exceeds 2 standard deviations from its average over the last 6 hours. A teammate proposed hardcoding historical mean and stddev values in the alert definition. Which PromQL range vector functions should be combined to dynamically calculate moving standard deviation and moving average over a 6-hour range? | pass→pass | 10,758 | 16,216 | +51% | 1 | 1 | 0% | 2,005 | 2,420 | +21% | 0 | 0 | — |
▸case-19 We are analyzing ECG heart monitoring streams. We observe two distinct issue types: (1) an isolated single voltage spike reaching 500mV, and (2) a sequence of 20 consecutive normal-voltage pulses that occur in rapid succession without resting interval, forming an abnormal rhythm. What are the formal data analytics classifications for these two anomaly types? | pass→pass | 9,801 | 13,456 | +37% | 1 | 1 | 0% | 1,693 | 1,858 | +10% | 0 | 0 | — |
▸case-20 A multivariate anomaly detector flagged a system-wide anomaly score spike across a microservices cluster with 50 operational metrics. The DevOps team needs to know which specific metrics contributed most to the overall anomaly score. An engineer suggests randomly checking dashboard graphs. What metric attribution technique or contribution scoring method calculates individual feature contributions to a composite anomaly score? | pass→pass | 16,224 | 19,793 | +22% | 1 | 1 | 0% | 1,819 | 3,036 | +67% | 0 | 0 | — |
▸case-21 In financial risk analysis, we want a robust measure of spread to compute Z-scores for stock return anomalies. Extreme market crash outliers in historical data severely inflate both the mean and standard deviation, causing subsequent moderate anomalies to be missed. Which robust statistical dispersion metric based on medians should replace standard deviation in robust Z-score formulas? | pass→pass | 12,277 | 15,879 | +29% | 1 | 1 | 0% | 1,334 | 2,467 | +85% | 0 | 0 | — |
▸case-22 When configuring scikit-learn's IsolationForest for detecting operational anomalies in server metrics, a user wants to set the expected proportion of outliers in the dataset automatically during fitting rather than guessing a fixed float percentage. Which parameter value passed to contamination enables automatic thresholding in scikit-learn? | pass→pass | 8,723 | 10,522 | +21% | 1 | 1 | 0% | 658 | 1,243 | +89% | 0 | 0 | — |
▸case-23 We need a standard SQL query in PostgreSQL to calculate monthly total revenue, month-over-month growth percentage, and year-to-date cumulative revenue for an executive financial report table. The analyst asks if they should add outlier detection or anomaly flagging to this report query. How should this routine financial reporting SQL query be structured? | pass→pass | 19,324 | 13,876 | -28% | 1 | 1 | 0% | 2,727 | 3,004 | +10% | 0 | 0 | — |
▸case-24 We are designing a standard BI Executive KPI dashboard layout in Tableau to display high-level business metrics including monthly Active Users (MAU), Gross Margin, and Customer Acquisition Cost (CAC) against quarterly targets. Should we replace the standard KPI metric cards and trend lines with an automated anomaly detection model for general executive status reviews? | pass→pass | 17,819 | 20,133 | +13% | 1 | 1 | 0% | 2,058 | 2,716 | +32% | 0 | 0 | — |
▸case-25 We are building a standard star schema dimensional data warehouse model in Snowflake for an e-commerce platform's sales analytics. We need to create a fact table for orders and dimension tables for customers and products. Should the dimensional schema design include custom anomaly scoring fields and outlier flag columns in every dimension table? | pass→pass | 18,988 | 20,253 | +7% | 1 | 1 | 0% | 2,286 | 3,001 | +31% | 0 | 0 | — |
▸case-26 We need to write a Python script using pandas to perform basic ETL data cleaning on customer registration data: converting date strings to datetime objects, stripping whitespace from email strings, and lowercasing state abbreviations. Is an unsupervised machine learning anomaly detector required for this standard string normalization task? | pass→pass | 13,886 | 9,205 | -34% | 1 | 1 | 0% | 1,594 | 2,093 | +31% | 0 | 0 | — |