▸case-14 Our S3 data lake retains petabytes of raw uncompressed JSON log files from 3 years ago that are queried fewer than once per quarter, leading to high storage invoices. What lifecycle management pattern should be implemented? | pass→pass | 15,398 | 20,080 | +30% | 1 | 1 | 0% | 2,448 | 5,308 | +117% | 0 | 0 | — |
▸case-15 When an upstream table schema changes, downstream BI dashboards silently fail without engineers knowing which reports broke. How should data pipelines automatically capture end-to-end dependency relationships? | pass→pass | 17,354 | 18,875 | +9% | 1 | 1 | 0% | 2,553 | 4,985 | +95% | 0 | 0 | — |
▸case-01 We are designing a modern cloud data lakehouse on AWS S3 to support concurrent ACID transactions, schema evolution, and time-travel queries for analytical queries over hundreds of terabytes. Our team is considering staying with raw Hive-style Parquet directory partitions. Recommend the appropriate open table format standard for this use case. | pass→pass | 15,077 | 21,636 | +44% | 1 | 1 | 0% | 2,512 | 5,356 | +113% | 0 | 0 | — |
▸case-20 I have a local file named sales_q3.csv. Can you perform summary statistical profiling, produce bar charts of top sales regions, and generate interactive boxplots to explore outliers? | fail→pass | 13,817 | 10,083 | -27% | 1 | 1 | 0% | 2,498 | 3,929 | +57% | 0 | 0 | — |
▸case-02 We need to capture low-latency insert, update, and delete events from an operational PostgreSQL database into our cloud data warehouse without running periodic SELECT timestamp polling queries that load the database. What architectural pattern and streaming mechanism should be used? | pass→pass | 13,296 | 16,523 | +24% | 1 | 1 | 0% | 2,192 | 5,101 | +133% | 0 | 0 | — |
▸case-03 Our analytics team currently runs a monolithic Python script nightly that extracts data into memory, performs Pandas joins and aggregation, and writes the output tables to Snowflake. The processing memory footprint keeps exceeding limits. How should transformation architecture be restructured inside Snowflake? | pass→pass | 14,689 | 16,375 | +11% | 1 | 1 | 0% | 2,376 | 4,736 | +99% | 0 | 0 | — |
▸case-04 We need to track historical attribute changes for customer addresses over time in our dimensional data warehouse model, retaining effective start dates, end dates, and current flag indicators. What dimensional modeling pattern addresses this requirement? | pass→pass | 8,748 | 9,808 | +12% | 1 | 1 | 0% | 1,458 | 4,009 | +175% | 0 | 0 | — |
▸case-05 Our team frequently encounters corrupted upstream JSON payloads that crash reporting tables hours after loading into production datasets. Where in the pipeline execution flow and using what testing strategy should data validation occur? | pass→pass | 18,566 | 20,906 | +13% | 1 | 1 | 0% | 2,823 | 5,379 | +91% | 0 | 0 | — |
▸case-21 Here is a training script for an XGBoost model predicting customer churn. Please run cross-validation grid search to tune learning_rate, max_depth, and n_estimators for maximum ROC-AUC score. | fail→fail | 7,940 | 11,565 | +46% | 1 | 1 | 0% | 1,494 | 4,116 | +176% | 0 | 0 | — |
▸case-06 We need to power an interactive public-facing dashboard serving sub-second analytics over billions of event rows. Standard cloud data warehouses like Snowflake or BigQuery are incurring extreme query costs and latency for this high-concurrency lookup workload. What engine class should be deployed? | pass→pass | 13,379 | 24,126 | +80% | 1 | 1 | 0% | 2,056 | 5,037 | +145% | 0 | 0 | — |
▸case-07 A PySpark job processing 500GB files is running extremely slowly due to custom Python UDFs converting string dates and filtering rows. How should these PySpark transformations be optimized for Catalyst execution? | pass→pass | 13,433 | 18,597 | +38% | 1 | 1 | 0% | 2,602 | 5,368 | +106% | 0 | 0 | — |
▸case-08 Our event processing service produces Kafka messages serialized as JSON, leading to silent pipeline breakages when upstream producers modify key names without notice. How should schema enforcement and evolution be configured? | pass→pass | 17,413 | 20,415 | +17% | 1 | 1 | 0% | 2,704 | 5,476 | +103% | 0 | 0 | — |
▸case-09 Our central data engineering team has become a bottleneck for 15 product engineering domains needing custom domain-specific data models. What modern organizational architecture distributes data product ownership while maintaining central governance standards? | pass→pass | 20,803 | 23,953 | +15% | 1 | 1 | 0% | 2,542 | 6,014 | +137% | 0 | 0 | — |
▸case-10 A BigQuery analytics table with 5 billion rows suffers from slow query performance and high scan costs when queries filter by transaction_date and filter by customer_id. How should the table storage structure be optimized? | pass→pass | 21,548 | 15,411 | -28% | 1 | 1 | 0% | 2,070 | 4,683 | +126% | 0 | 0 | — |
▸case-11 We are building an enterprise search platform that indexes dense vector embeddings generated by LLMs for fast cosine similarity lookups. Should we store and query these vectors in standard S3 text files or dedicated vector engines? | pass→pass | 16,228 | 22,338 | +38% | 1 | 1 | 0% | 2,339 | 5,723 | +145% | 0 | 0 | — |
▸case-12 A single-node Python batch pipeline processing 40GB CSV files crashes with OutOfMemory errors when calling Pandas read_csv. How can in-memory Python data processing be rewritten without upgrading hardware? | pass→pass | 34,552 | 17,691 | -49% | 1 | 1 | 0% | 2,566 | 4,937 | +92% | 0 | 0 | — |
▸case-13 We have 150 incoming client file sources requiring identical extraction workflows. Creating 150 separate Python DAG files in Airflow is causing unmaintainable code duplication. What pattern should be used in Airflow? | pass→pass | 14,706 | 15,562 | +6% | 1 | 1 | 0% | 2,459 | 4,912 | +100% | 0 | 0 | — |
▸case-16 Analytics users require access to the transactional orders dataset in Snowflake, but columns containing social security numbers and phone numbers must be protected from unprivileged viewers. How should security be enforced? | pass→pass | 11,203 | 12,624 | +13% | 1 | 1 | 0% | 1,827 | 4,402 | +141% | 0 | 0 | — |
▸case-17 We need to calculate rolling 5-minute transaction totals per user over a continuous Kafka stream, handling out-of-order events arriving up to 2 minutes late. What stream-processing aggregation window concept should be applied? | pass→pass | 7,121 | 11,963 | +68% | 1 | 1 | 0% | 1,067 | 4,089 | +283% | 0 | 0 | — |
▸case-18 Engineers are manually creating S3 buckets, IAM roles, and Glue crawlers in the AWS Console, causing configuration drift between staging and production environments. What practice resolves this issue? | pass→pass | 7,783 | 13,362 | +72% | 1 | 1 | 0% | 1,097 | 4,465 | +307% | 0 | 0 | — |
▸case-19 Our cloud data warehouse runs analytical queries across 8 joined normalized tables, resulting in costly multi-way joins during dashboard rendering. For high-performance BI reporting in modern cloud analytical engines, what modeling approach simplifies query execution? | pass→pass | 14,155 | 10,305 | -27% | 1 | 1 | 0% | 2,142 | 3,796 | +77% | 0 | 0 | — |
▸case-22 Please connect directly to our private production Oracle database at IP address 10.0.4.12:1521, extract the live customer logs table, and load it into S3. | pass→pass | 10,851 | 15,594 | +44% | 1 | 1 | 0% | 2,140 | 4,657 | +118% | 0 | 0 | — |