▸case-01 We need to set up real-time change data capture from our PostgreSQL database to Apache Kafka for our `orders` and `line_items` tables. We require tight sub-second latency (under 200ms) and guaranteed exactly-once processing. Please provide a complete implementation package including the connector settings, initial load snapshot strategy, schema evolution management rules, replication monitoring guidelines, and summary documentation. | pass→pass | 27,512 | 19,024 | -31% | 1 | 1 | 0% | 5,113 | 4,339 | -15% | 0 | 0 | — |
▸case-02 Our team is establishing a streaming pipeline to capture changes from a MySQL database and publish them to Amazon Kinesis. The tables involved are `user_profiles` and `billing_events`. Latency should be under 1000ms, and standard at-least-once or exactly-once streaming needs to be addressed. Generate the full CDC configuration output containing the source connector config, snapshot strategy details, schema change handling options, monitoring metrics configuration, and supporting documentation. | pass→pass | 31,359 | 14,459 | -54% | 1 | 1 | 0% | 5,876 | 3,298 | -44% | 0 | 0 | — |
▸case-03 I am building an ETL ingestion pipeline from a Microsoft SQL Server database to Google Cloud Pub/Sub targeting our `inventory_levels` table. We need minimal replication lag (target 500ms latency) and robust delivery settings. Please output the full solution with the connector parameters, snapshot execution strategy, schema management settings, connector monitoring setup, and technical documentation. | fail→pass | 30,477 | 14,048 | -54% | 1 | 1 | 0% | 6,214 | 3,167 | -49% | 0 | 0 | — |
▸case-04 We are integrating an Oracle database storing `customer_accounts` into Apache Kafka. Our DBA team suggested adding shadow trigger tables for every table to capture updates, but we want minimal overhead on the source OLTP transaction processing. Return the configuration package as a JSON object containing connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation. | pass→pass | 17,033 | 13,269 | -22% | 1 | 1 | 0% | 3,243 | 3,067 | -5% | 0 | 0 | — |
▸case-05 When records in our PostgreSQL `shopping_carts` table are deleted, downstream microservices need to purge their local state in Apache Kafka. We are tempted to suppress delete events entirely to reduce network bandwidth. Output a full CDC JSON response with keys connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation that correctly handles record deletions. | pass→pass | 14,786 | 14,155 | -4% | 1 | 1 | 0% | 2,767 | 3,228 | +17% | 0 | 0 | — |
▸case-06 We have a 2 Terabyte `audit_logs` table in MySQL that needs to be synchronized to Amazon Kinesis. A naive blocking table lock for the initial load will bring down production web servers. Provide the JSON implementation response containing connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation detailing how to safely take the initial snapshot. | pass→pass | 18,833 | 15,851 | -16% | 1 | 1 | 0% | 3,518 | 3,494 | -1% | 0 | 0 | — |
▸case-07 In our PostgreSQL to Kafka streaming pipeline for `payments`, developers frequently alter table column types without notifying the data team, causing connector failures. Output a JSON payload with connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation that configures error handling for schema changes. | pass→pass | 15,489 | 16,507 | +7% | 1 | 1 | 0% | 3,758 | 4,049 | +8% | 0 | 0 | — |
▸case-08 Our SQL Server ingestion pipeline for `order_fulfillment` into Cloud Pub/Sub experiences undetected delays during peak sales events. Provide a JSON response formatted with connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation establishing observability. | pass→pass | 18,385 | 13,165 | -28% | 1 | 1 | 0% | 3,961 | 3,148 | -21% | 0 | 0 | — |
▸case-09 We are configuring PostgreSQL stream ingestion for `user_subscriptions` to Kafka. The base configuration uses default PostgreSQL settings, but CDC fails to stream row changes. Provide the CDC JSON output structure (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation) addressing the database write-ahead log configuration. | pass→pass | 11,545 | 10,862 | -6% | 1 | 1 | 0% | 2,431 | 2,470 | +2% | 0 | 0 | — |
▸case-10 Our MySQL `inventory_movements` CDC connector to Amazon Kinesis is skipping update events when statement-based replication is active. Return the required JSON object with connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation specifying the required MySQL binary log format. | pass→pass | 10,547 | 12,967 | +23% | 1 | 1 | 0% | 1,683 | 2,953 | +75% | 0 | 0 | — |
▸case-11 For our `financial_transactions` table streamed from SQL Server to Kafka, business auditors require that no message is processed twice even if the connector restarts. Produce the JSON payload containing connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation addressing delivery guarantees. | pass→pass | 13,743 | 14,043 | +2% | 1 | 1 | 0% | 2,747 | 3,019 | +10% | 0 | 0 | — |
▸case-12 We are capturing changes from Oracle DB `ledger_entries` to Pub/Sub. The team considers using continuous polling `SELECT * WHERE updated_at > timestamp`, but frequent updates are missed due to transaction commit order. Produce the CDC JSON object (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation) selecting the appropriate capture method. | pass→pass | 20,731 | 14,882 | -28% | 1 | 1 | 0% | 2,852 | 3,266 | +15% | 0 | 0 | — |
▸case-13 We are building a complete end-to-end CDC pipeline capturing MySQL `catalog_products` changes into Apache Kafka and consuming them into target storage. Provide the full output JSON with connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, and documentation defining connector settings. | pass→pass | 17,570 | 16,356 | -7% | 1 | 1 | 0% | 3,585 | 3,739 | +4% | 0 | 0 | — |
▸case-14 Our Kafka CDC pipeline for PostgreSQL `shipping_addresses` needs strict schema version governance so consumers do not crash on field additions. Provide the CDC JSON structure (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation) configuring schema evolution. | pass→pass | 16,354 | 13,248 | -19% | 1 | 1 | 0% | 3,197 | 2,985 | -7% | 0 | 0 | — |
▸case-15 We are deploying a CDC pipeline from SQL Server `logs` table to Google Cloud Pub/Sub, but the table already exists in the target system with historical data loaded. We do not want to re-read historical rows upon startup. Return the standard JSON response (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation). | pass→pass | 11,045 | 9,684 | -12% | 1 | 1 | 0% | 2,177 | 2,203 | +1% | 0 | 0 | — |
▸case-16 When failing over our MySQL `user_credits` master database, the CDC connector loses its position in the binlog files. Output the standard JSON package (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation) configuring resilient transaction offset tracking. | pass→pass | 11,423 | 15,683 | +37% | 1 | 1 | 0% | 2,400 | 3,293 | +37% | 0 | 0 | — |
▸case-17 Our PostgreSQL database `low_activity_events` has hours of zero write activity, causing the CDC stream position to stale and WAL logs to bloat on the DB server. Provide the JSON output package (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation) addressing idle connection WAL retention. | pass→pass | 12,095 | 15,060 | +25% | 1 | 1 | 0% | 2,374 | 2,723 | +15% | 0 | 0 | — |
▸case-18 When writing CDC messages for `customer_orders` from MySQL to Kafka, downstream join operations require the message key to contain the primary key columns rather than plain text strings. Return the CDC JSON package (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation). | pass→pass | 11,928 | 13,440 | +13% | 1 | 1 | 0% | 2,424 | 2,401 | -1% | 0 | 0 | — |
▸case-19 We are setting up CDC for an Oracle database table `wire_transfers`. When updates occur, the CDC stream only receives modified columns, but downstream needs the full primary key even when un-updated. Provide the CDC JSON output (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation). | pass→pass | 15,344 | 12,204 | -20% | 1 | 1 | 0% | 2,856 | 2,679 | -6% | 0 | 0 | — |
▸case-20 When streaming MySQL `web_clicks` updates to Amazon Kinesis, events for the same user must arrive in strict sequential order within the same Kinesis shard. Output the full JSON implementation response (connectorConfig, snapshotStrategy, schemaConfig, monitoringConfig, documentation). | pass→pass | 13,648 | 16,364 | +20% | 1 | 1 | 0% | 2,771 | 3,709 | +34% | 0 | 0 | — |
▸case-21 Our PostgreSQL database query `SELECT * FROM orders WHERE status = 'PENDING' AND created_at > '2023-01-01'` is running slowly. How should we create B-tree indexes or rewrite this SQL query to improve query performance on PostgreSQL? | pass→pass | 11,831 | 9,424 | -20% | 1 | 1 | 0% | 2,327 | 2,219 | -5% | 0 | 0 | — |
▸case-22 We have an existing Apache Kafka cluster running on Kubernetes with 10 brokers, and topic `telemetry-data` has hot partition issues. How do we run `kafka-reassign-partitions.sh` to rebalance partitions across brokers? | pass→pass | 14,314 | 13,189 | -8% | 1 | 1 | 0% | 2,770 | 2,950 | +6% | 0 | 0 | — |
▸case-23 We are building a Kimball star schema in Snowflake using dbt. How do we write a dbt SQL model for a slowly changing dimension (SCD Type 2) dimension table tracking `dim_customers` history using `dbt snapshot`? | pass→pass | 15,735 | 12,718 | -19% | 1 | 1 | 0% | 3,228 | 3,017 | -7% | 0 | 0 | — |