▸case-01 We are building an Apache Airflow workflow for daily model retraining. Developers want to write all dataset prep, model training, and evaluation directly inside a single PythonOperator function to avoid intermediate storage overhead. How should this pipeline be structured across Airflow task boundaries? | pass→pass | 17,351 | 17,488 | +1% | 1 | 1 | 0% | 3,050 | 3,263 | +7% | 0 | 0 | — |
▸case-02 In Kubeflow Pipelines SDK v2, a pipeline task needs access to a large persistent dataset stored on a Kubernetes PersistentVolumeClaim named 'ml-data-pvc'. A developer proposes manually copying dataset files inside the component container entrypoint via curl at start time. What is the standard declarative Kubernetes volume attachment method in KFP? | pass→pass | 8,385 | 6,689 | -20% | 1 | 1 | 0% | 1,479 | 1,305 | -12% | 0 | 0 | — |
▸case-03 We are modeling user transaction features for real-time fraud detection. A engineer suggests saving pre-aggregated feature vectors as raw CSV files on NFS and querying them via SQL during model inference. What feature store abstraction should be defined to manage entity keys and feature retrieval for online serving? | pass→pass | 15,583 | 11,321 | -27% | 1 | 1 | 0% | 2,647 | 2,415 | -9% | 0 | 0 | — |
▸case-04 After automated pipeline evaluation passes validation checks, a model artifact is logged in MLflow. The engineer wants to automatically serve this model to production by overwriting the local filesystem model file on the inference server. How should the model registry handle promotion safely? | pass→pass | 14,813 | 10,167 | -31% | 1 | 1 | 0% | 2,431 | 1,922 | -21% | 0 | 0 | — |
▸case-05 Our production churn model's accuracy drops periodically due to distribution shifts in input covariates. The team currently triggers retraining on a fixed calendar schedule regardless of data statistics. What automated mechanism should monitor input drift and trigger downstream retraining pipelines? | pass→fail | 17,834 | 22,172 | +24% | 1 | 1 | 0% | 3,102 | 4,409 | +42% | 0 | 0 | — |
▸case-06 An automated ML retraining Airflow DAG starts immediately when new raw sensor files land in S3, occasionally crashing downstream training tasks due to missing null-check validation. Where and how should data validation be integrated in the workflow? | pass→pass | 14,797 | 17,108 | +16% | 1 | 1 | 0% | 2,492 | 3,305 | +33% | 0 | 0 | — |
▸case-07 When running hyperparameter optimization across 100 trials, an engineer suggests hardcoding parameter search loops within a single non-parallel Python script. How should distributed hyperparameter search be integrated into an ML orchestration pipeline? | pass→pass | 15,978 | 18,395 | +15% | 1 | 1 | 0% | 2,860 | 3,432 | +20% | 0 | 0 | — |
▸case-08 A multi-node PyTorch training job needs to run across 4 GPU nodes in Kubernetes. A developer suggests launching 4 independent Pods without communication env vars and manually managing process synchronization. Which Kubernetes custom resource operator should manage worker discovery and rendezvous? | pass→pass | 9,308 | 6,856 | -26% | 1 | 1 | 0% | 1,435 | 1,473 | +3% | 0 | 0 | — |
▸case-09 A freshly trained model artifact achieves an F1-score of 0.82. The currently deployed model has an F1-score of 0.87. A script automatically registers and deploys every newly trained model regardless of metrics. What logic must be implemented in the evaluation component? | pass→pass | 11,451 | 10,121 | -12% | 1 | 1 | 0% | 1,765 | 1,918 | +9% | 0 | 0 | — |
▸case-10 We need to deploy a newly registered XGBoost model alongside the existing model to test live production traffic, starting with 10% traffic. The DevOps team proposes deploying a separate endpoint and manually updating DNS records. How should canary traffic splitting be configured natively on Kubernetes for ML serving? | pass→pass | 14,801 | 11,347 | -23% | 1 | 1 | 0% | 2,838 | 2,339 | -18% | 0 | 0 | — |
▸case-11 To comply with auditing requirements, our team needs to trace which specific dataset partition and git commit produced a deployed model artifact. Currently, logs are scattered across individual task stdout logs. How should pipeline lineage metadata be captured automatically? | pass→pass | 15,174 | 20,152 | +33% | 1 | 1 | 0% | 2,753 | 3,488 | +27% | 0 | 0 | — |
▸case-12 A daily batch prediction pipeline needs to generate scores for 50 million customers. A developer suggests loading all records into memory in a single API container and making sequential REST calls. What is the scalable architecture for batch inference DAGs? | pass→pass | 15,503 | 22,280 | +44% | 1 | 1 | 0% | 2,700 | 3,647 | +35% | 0 | 0 | — |
▸case-13 Our online feature store serves real-time latency requests, but training jobs read historical features directly from transactional databases, causing training-serving skew. How should feature storage be unified across training and serving paths? | pass→pass | 16,140 | 14,138 | -12% | 1 | 1 | 0% | 2,829 | 2,830 | +0% | 0 | 0 | — |
▸case-14 A newly promoted model endpoint begins returning 500 error responses in production after a canary deployment. The team currently relies on manual manual SSH commands to restart old container images. How should automated rollback be configured? | pass→pass | 14,822 | 15,881 | +7% | 1 | 1 | 0% | 2,414 | 3,041 | +26% | 0 | 0 | — |
▸case-15 A deep learning container running inside a Kubeflow Pipelines step fails due to CUDA out-of-memory errors because no GPU hardware was provisioned. How are GPU resources requested in KFP task definitions? | fail→pass | 12,564 | 9,562 | -24% | 1 | 1 | 0% | 2,066 | 1,999 | -3% | 0 | 0 | — |
▸case-16 In Kubeflow Pipelines, expensive data preprocessing steps rerun from scratch on every pipeline trigger, even when input dataset URIs have not changed. How should task caching be configured? | pass→pass | 14,552 | 10,363 | -29% | 1 | 1 | 0% | 2,115 | 2,445 | +16% | 0 | 0 | — |
▸case-17 When a new dataset file is uploaded to AWS S3, we want an automated Airflow ML pipeline to start immediately without continuously running polling tasks. What AWS and Airflow pattern enables event-driven DAG execution? | pass→pass | 12,213 | 12,544 | +3% | 1 | 1 | 0% | 2,298 | 2,555 | +11% | 0 | 0 | — |
▸case-18 A team of data scientists runs training scripts locally, saving hyperparameters in local text files and loss curves as PNG images. How should experiment tracking be standardized across team runs? | pass→pass | 12,305 | 14,545 | +18% | 1 | 1 | 0% | 2,220 | 2,895 | +30% | 0 | 0 | — |
▸case-19 A Kubeflow pipeline task needs database credentials to pull training data. A developer embedded the plaintext Postgres password into the pipeline Python code repository. How should credentials be securely injected into pipeline components? | pass→pass | 11,884 | 14,516 | +22% | 1 | 1 | 0% | 2,130 | 2,993 | +41% | 0 | 0 | — |
▸case-20 We are selecting the context window size and attention head count for a custom transformer language model architecture to optimize key-value cache memory during generation. What design adjustments should be made to the attention heads? | pass→pass | 19,659 | 9,962 | -49% | 1 | 1 | 0% | 2,881 | 2,081 | -28% | 0 | 0 | — |
▸case-21 How do you write a custom CUDA C++ kernel using shared memory tile indexing to accelerate matrix multiplication on NVIDIA Ampere GPUs? | pass→pass | 18,954 | 21,134 | +12% | 1 | 1 | 0% | 3,793 | 4,424 | +17% | 0 | 0 | — |
▸case-22 How do you design a React functional component using Tailwind CSS to render an interactive confusion matrix heatmap with hover tooltips for end users? | pass→pass | 28,470 | 17,349 | -39% | 1 | 1 | 0% | 4,696 | 3,718 | -21% | 0 | 0 | — |