▸case-04 We are establishing a review process for turbine bearing failure candidates. Missing an actual failure causes catastrophic downtime ($500k cost), while inspecting a false alarm costs $200 in technician time. How should we configure our review criteria and threshold decisions given this asymmetry? | fail→fail | 14,882 | 15,465 | +4% | 1 | 1 | 0% | 2,565 | 2,709 | +6% | 0 | 0 | — |
▸case-01 We are auditing a dataset of 50,000 credit card transactions with 15 numerical features (e.g., amount, distance from home, transaction velocity). The relationship between features is non-linear and multi-dimensional. A team member suggested running standard deviation Z-score filtering independently on each column. What unsupervised anomaly detection approach should we choose instead to capture complex multi-feature interactions, and why? | pass→pass | 12,629 | 12,630 | +0% | 1 | 1 | 0% | 2,172 | 2,523 | +16% | 0 | 0 | — |
▸case-02 We monitor vibration readings from 200 industrial pumps. Ambient temperatures cause natural baseline shifts throughout the day, so a reading of 45 Hz at noon is normal, but 45 Hz at 3 AM is a local anomaly. An analyst suggested setting a single global threshold across all 24 hours. Which method better detects local density variations relative to neighboring time-window points? | pass→pass | 13,434 | 9,821 | -27% | 1 | 1 | 0% | 2,340 | 1,974 | -16% | 0 | 0 | — |
▸case-03 Our automated system scored 500,000 wire transfers with anomaly scores between 0.0 and 1.0. The compliance team has 2 analyst hours per day and cannot inspect every transaction above 0.5. How should we structure the output deliverable to make human inspection feasible? | pass→pass | 14,178 | 13,764 | -3% | 1 | 1 | 0% | 2,352 | 2,570 | +9% | 0 | 0 | — |
▸case-05 We have collected 100,000 baseline telemetry logs from a brand-new jet engine during bench testing where zero failures or anomalies occurred. We want to train a model strictly on this clean normal baseline to flag any future deviant operating behavior. Is Local Outlier Factor or One-Class SVM better suited for this clean novelty detection setup? | pass→pass | 15,088 | 17,123 | +13% | 1 | 1 | 0% | 2,672 | 3,224 | +21% | 0 | 0 | — |
▸case-06 We are evaluating API response latency logs in milliseconds to flag anomalous response delays. The distribution is heavily right-skewed with long tails. A junior developer proposed using Z-scores based on mean and standard deviation. Why is IQR thresholding superior here? | pass→pass | 13,454 | 10,339 | -23% | 1 | 1 | 0% | 2,323 | 2,131 | -8% | 0 | 0 | — |
▸case-07 We flagged 45 suspicious high-value trade executions out of 200,000 daily orders. We need to visualize these flagged orders for senior risk managers who do not want to wade through raw tabular data. What plot structure best highlights these suspicious records? | pass→pass | 13,246 | 11,569 | -13% | 1 | 1 | 0% | 2,377 | 2,193 | -8% | 0 | 0 | — |
▸case-08 We received a raw CSV export of customer registrations in PostgreSQL. We need to identify missing values, verify column data types against our schema, and remove duplicate user entries. Should we run an Isolation Forest model to validate this file? | fail→pass | 10,681 | 6,096 | -43% | 1 | 1 | 0% | 1,919 | 1,369 | -29% | 0 | 0 | — |
▸case-09 We have finalized our feature engineering and need to train, package, and deploy an end-to-end automated machine learning pipeline with a model registry and real-time REST API endpoint using scikit-learn Pipeline objects. Is manual candidate triage the correct framework for this? | pass→pass | 14,246 | 8,438 | -41% | 1 | 1 | 0% | 2,395 | 1,962 | -18% | 0 | 0 | — |
▸case-10 We are submitting a manuscript on seismic telemetry anomalies to Nature Geosciences. We need multi-panel vector graphics with custom CMYK color maps, EPS formatting, and strict journal typographic compliance. Should we rely on informal diagnostic triage plots for the paper figures? | pass→pass | 14,781 | 3,985 | -73% | 1 | 1 | 0% | 2,389 | 917 | -62% | 0 | 0 | — |
▸case-11 We need to score 10,000,000 real-time log records per hour for unexpected network byte transfers. Computational efficiency and low memory overhead are critical constraints. Should we choose Local Outlier Factor or Isolation Forest for scoring this dataset at scale? | pass→pass | 14,831 | 12,401 | -16% | 1 | 1 | 0% | 2,612 | 2,350 | -10% | 0 | 0 | — |
▸case-12 Our anti-money laundering model flags 300 bank accounts per week for potential structuring. Currently analysts complain that half the alerts are legitimate small business cash deposits. What specific elements should be included in a false-positive triage checklist for analysts? | pass→pass | 17,868 | 15,989 | -11% | 1 | 1 | 0% | 2,872 | 2,726 | -5% | 0 | 0 | — |
▸case-13 We are running an Isolation Forest on credit card transactions where historical fraud incidence is known to be approximately 0.05% (1 in 2000 transactions). A default setting in many libraries sets contamination to 0.1 (10%). How should we configure the contamination hyperparameter? | pass→pass | 14,129 | 13,119 | -7% | 1 | 1 | 0% | 2,468 | 2,448 | -1% | 0 | 0 | — |
▸case-14 We are monitoring two financial metrics: daily revenue and transaction count. These two metrics have a strong positive correlation (r = 0.92). A point with low revenue and low transaction count is normal, but high revenue with very low transaction count is highly suspicious. Why does independent Euclidean Z-score thresholding fail here, and what distance metric addresses it? | pass→pass | 10,619 | 10,899 | +3% | 1 | 1 | 0% | 2,037 | 2,185 | +7% | 0 | 0 | — |
▸case-15 A temperature sensor on an outdoor storage vessel slowly warms up by 15°C over six months due to seasonal summer weather, but we need to catch sudden 3°C spikes occurring within a 10-minute window. Applying a static threshold flags the entire summer season as anomalous. What pre-processing step is required before applying spike detection thresholds? | pass→pass | 7,789 | 9,425 | +21% | 1 | 1 | 0% | 1,549 | 1,815 | +17% | 0 | 0 | — |
▸case-16 We need to deliver a CSV summary report of flagged fraudulent wire transfers to the legal compliance team. The model generates raw internal node depth scores. What columns must be included in the deliverable table so non-technical auditors can review the flagged records? | pass→pass | 11,679 | 10,319 | -12% | 1 | 1 | 0% | 2,081 | 2,135 | +3% | 0 | 0 | — |
▸case-17 We are analyzing server error logs represented as TF-IDF sparse vectors with 20,000 vocabulary features. Direct distance-based methods like LOF perform poorly due to the curse of dimensionality. What strategy should be applied to detect anomalous log entries? | pass→pass | 14,325 | 10,068 | -30% | 1 | 1 | 0% | 2,416 | 1,787 | -26% | 0 | 0 | — |
▸case-18 A hospital laboratory automated analyzer flags 12 blood panel results per day with anomalous potassium and sodium combinations. An automated system was proposed to automatically delete these records from the patient portal. How should the anomaly output workflow be structured instead? | pass→pass | 15,598 | 13,443 | -14% | 1 | 1 | 0% | 2,559 | 2,401 | -6% | 0 | 0 | — |
▸case-19 We are monitoring high-frequency ECG cardiac signal waveforms to detect rare arrhythmia patterns. Standard tabular thresholding on peak voltage misses complex structural shape deformations in the waveform over time. What deep learning or signal reconstruction technique is commonly used to quantify shape anomalies via reconstruction error? | pass→pass | 14,474 | 12,606 | -13% | 1 | 1 | 0% | 2,421 | 2,502 | +3% | 0 | 0 | — |
▸case-20 An ISP monitors gigabit network traffic to detect distributed denial-of-service (DDoS) bandwidth spikes. Traffic naturally peaks at 9 PM (100 Gbps normal) and drops at 4 AM (10 Gbps normal). A spike to 25 Gbps at 4 AM is a massive attack, but would be ignored by a static global 80 Gbps threshold. How should the anomaly thresholding be designed? | pass→pass | 16,763 | 17,873 | +7% | 1 | 1 | 0% | 3,056 | 3,467 | +13% | 0 | 0 | — |
▸case-21 In a network intrusion detection audit, stealthy cyber attackers deliberately split data exfiltration into tiny packets that stay just below the 95th percentile anomaly alert threshold. What explicit review check should be added to the false-negative auditing checklist to detect these low-and-slow threats? | pass→pass | 9,781 | 12,106 | +24% | 1 | 1 | 0% | 1,627 | 2,278 | +40% | 0 | 0 | — |
▸case-22 A logistics company wants to inspect 30 flagged shipment delays across 8 correlated dimensions (e.g., origin delay, customs clearance time, transit speed, weight, ambient temp). What visualization format best displays how these 30 candidate shipments deviate across all 8 dimensions simultaneously? | pass→pass | 11,939 | 12,706 | +6% | 1 | 1 | 0% | 1,924 | 2,538 | +32% | 0 | 0 | — |
▸case-23 We have a database table with leading/trailing spaces in customer names, inconsistent phone number formats ('123-456-7890' vs '1234567890'), and duplicate entries. Should we run an Isolation Forest model to clean these string formatting issues? | pass→pass | 9,348 | 8,054 | -14% | 1 | 1 | 0% | 1,735 | 1,746 | +1% | 0 | 0 | — |