▸case-01 I am building an analytics pipeline for my e-commerce site. I need to know whether to use batch or stream processing. Please provide a detailed breakdown of the tradeoffs for my use case, including a clear final recommendation on which architecture to choose. | pass→pass | 15,600 | 17,117 | +10% | 1 | 1 | 0% | 2,730 | 2,810 | +3% | 0 | 0 | — |
▸case-02 Our infrastructure generates 10TB of daily application logs. The business team reviews the aggregated dashboard every morning at 9 AM. Propose the processing paradigm (batch or stream) and justify it based on compute utilization. | pass→pass | 12,913 | 14,306 | +11% | 1 | 1 | 0% | 2,168 | 2,075 | -4% | 0 | 0 | — |
▸case-03 We are implementing a credit card fraud detection system that must block transactions before authorization completes. Compare the two main data processing paradigms and select the appropriate one, specifying the latency requirement. | pass→pass | 10,638 | 7,686 | -28% | 1 | 1 | 0% | 1,773 | 1,305 | -26% | 0 | 0 | — |
▸case-04 We need to update our revenue calculation logic and apply it to 5 years of historical data. We currently use a real-time event pipeline. Explain the architectural challenge of doing this in a pure real-time paradigm and recommend the standard pattern to solve it. | pass→fail | 15,207 | 17,051 | +12% | 1 | 1 | 0% | 2,885 | 2,895 | +0% | 0 | 0 | — |
▸case-05 Our mobile game sends telemetry data, but devices often go offline and sync hours later. We need 100% accurate daily active user counts. Which processing model handles this natively without complex state retention, and what is the mechanism? | pass→pass | 24,318 | 9,423 | -61% | 1 | 1 | 0% | 3,694 | 1,455 | -61% | 0 | 0 | — |
▸case-06 We need to join ad impressions with ad clicks. Clicks can happen up to 30 days after an impression. If we build this using a continuous processing engine, what specific infrastructural burden is created by this 30-day window? | pass→pass | 12,902 | 13,746 | +7% | 1 | 1 | 0% | 2,097 | 2,301 | +10% | 0 | 0 | — |
▸case-07 We need to calculate a 5-minute rolling average of CPU temperature from IoT sensors, updating every 10 seconds. Recommend the processing paradigm and the specific windowing function required. | pass→pass | 7,278 | 6,877 | -6% | 1 | 1 | 0% | 1,465 | 1,052 | -28% | 0 | 0 | — |
▸case-08 We need to run an hourly ETL job transforming 50TB of raw JSON in S3 into Parquet. The team wants to use Apache Flink. Critique this choice and suggest the standard distributed processing engine designed specifically for this paradigm. | pass→pass | 16,302 | 11,252 | -31% | 1 | 1 | 0% | 2,346 | 1,879 | -20% | 0 | 0 | — |
▸case-09 We are building a sub-second anomaly detection system on Kafka topics. A junior engineer suggests using Apache Airflow to trigger Python scripts every minute. Explain why this fails the latency requirement and suggest an appropriate continuous processing engine. | pass→pass | 15,342 | 13,819 | -10% | 1 | 1 | 0% | 2,613 | 2,249 | -14% | 0 | 0 | — |
▸case-10 Our trading platform needs both real-time P&L estimates (within 1 second) and highly accurate end-of-day P&L reports for compliance. Name the architecture pattern that uses both paradigms to achieve this and describe its main operational drawback. | pass→pass | 4,877 | 3,796 | -22% | 1 | 1 | 0% | 834 | 679 | -19% | 0 | 0 | — |
▸case-11 To avoid maintaining separate codebases for our real-time and daily pipelines, we want to treat everything as an append-only log and use a single processing engine. Name this architecture pattern. | pass→pass | 3,859 | 3,153 | -18% | 1 | 1 | 0% | 678 | 499 | -26% | 0 | 0 | — |
▸case-12 We are using Apache Spark for real-time processing. We noticed our latency is around 500ms and cannot go lower, unlike our previous Apache Storm setup. Explain the underlying processing model Spark uses that causes this floor. | pass→pass | 13,638 | 12,527 | -8% | 1 | 1 | 0% | 2,365 | 2,098 | -11% | 0 | 0 | — |
▸case-13 We are processing financial transactions from a message broker. We cannot tolerate duplicates or dropped messages. Compare the difficulty of achieving exactly-once semantics in continuous processing versus nightly processing. | fail→pass | 16,569 | 16,937 | +2% | 1 | 1 | 0% | 2,726 | 2,805 | +3% | 0 | 0 | — |
▸case-14 We added a new field to our user profile schema and need to back-populate it for all historical events. Compare the operational effort of doing this in a pure event-streaming architecture versus a data lake architecture. | pass→pass | 19,104 | 15,804 | -17% | 1 | 1 | 0% | 3,131 | 2,764 | -12% | 0 | 0 | — |
▸case-15 In our continuous processing pipeline, events often arrive out of order due to network delays. Name the two concepts required to handle this: the timestamp used for logic, and the threshold used to determine when to emit results. | pass→pass | 3,751 | 2,673 | -29% | 1 | 1 | 0% | 614 | 440 | -28% | 0 | 0 | — |
▸case-16 We are designing a system to ingest 10 million events per second. Latency is not a concern, but we need to minimize compute costs. Should we optimize for throughput or latency, and which processing paradigm aligns with this goal? | pass→pass | 15,598 | 11,612 | -26% | 1 | 1 | 0% | 2,631 | 1,685 | -36% | 0 | 0 | — |
▸case-17 We have chosen a batch processing architecture using Snowflake. We are deciding between Star Schema and Data Vault. Which of these two modeling techniques explicitly separates business keys into 'Hub' tables? | pass→pass | 3,149 | 3,399 | +8% | 1 | 1 | 0% | 556 | 609 | +10% | 0 | 0 | — |
▸case-18 We are using Kafka for our stream processing pipeline. We have a topic with 100 partitions but only 10 consumer instances in our group. How will the partitions be assigned to the consumers, and what happens if we add 100 more consumers? | pass→pass | 10,181 | 10,135 | -0% | 1 | 1 | 0% | 1,500 | 1,683 | +12% | 0 | 0 | — |
▸case-19 In our batch processing system, we use Apache Airflow. We have a DAG with 50 parallel tasks. We are hitting database connection limits. What Airflow feature should we configure to restrict the number of concurrent tasks hitting the database? | pass→pass | 7,931 | 4,818 | -39% | 1 | 1 | 0% | 1,432 | 805 | -44% | 0 | 0 | — |
▸case-20 Our nightly ETL job occasionally fails halfway through. When it restarts, it appends duplicate rows to the destination table. What property must we implement in our batch job design to prevent this, and what is a common technique to achieve it? | pass→pass | 10,609 | 7,524 | -29% | 1 | 1 | 0% | 1,630 | 1,274 | -22% | 0 | 0 | — |
▸case-21 Our continuous processing application crashes occasionally. Upon restart, it loses its internal aggregation state. What mechanism must be enabled in engines like Flink or Spark to save this state durably? | pass→pass | 8,945 | 6,648 | -26% | 1 | 1 | 0% | 1,537 | 1,221 | -21% | 0 | 0 | — |
▸case-22 A startup wants to use a continuous processing engine for all their data pipelines to be 'real-time', even though most reports are viewed weekly. Explain the primary financial disadvantage of this approach compared to scheduled jobs. | pass→pass | 11,369 | 10,981 | -3% | 1 | 1 | 0% | 1,637 | 1,500 | -8% | 0 | 0 | — |