▸case-02 Our team is building an automated image classification dataset of roughly 60,000 medical dermatology scans gathered from several clinical partners. Before training, we want a complete curation and health assessment of the corpus. Please evaluate the class balance and source distributions, identify any harmful biases or representation gaps across subgroups, propose a stratified data partitioning strategy for split creation, specify recommendations for acquiring supplementary samples, and provide a filled-out ethical considerations checklist. | fail→pass | 38,082 | 31,070 | -18% | 1 | 1 | 0% | 6,229 | 6,385 | +3% | 0 | 0 | — |
▸case-03 We are reviewing a custom NLP dataset containing about 3,500 customer service dialogue transcripts labeled with intents and urgency levels across different regions. Could you perform a dataset curation report for us? Specifically, we want a break-down of label counts and tag co-occurrences, an evaluation of potential fairness issues or spurious correlations, a recommended validation split strategy suitable for our sample size, actionable steps for dataset augmentation or expansion, and an ethical safety checklist. | fail→fail | 26,937 | 25,066 | -7% | 1 | 1 | 0% | 4,343 | 5,152 | +19% | 0 | 0 | — |
▸case-04 We are training a ResNet-50 model on an imbalanced image dataset. Should we use Focal Loss or Cross-Entropy loss with class weights? Please evaluate the mathematical formulations and recommend hyperparameter settings for gamma and alpha. | pass→pass | 18,851 | 18,756 | -1% | 1 | 1 | 0% | 3,297 | 4,106 | +25% | 0 | 0 | — |
▸case-05 We are fine-tuning a Llama-3 8B model using QLoRA. What learning rate, LoRA rank (r), lora_alpha, and warmup ratio should we use for optimal convergence on instruction tuning? | pass→pass | 13,172 | 13,198 | +0% | 1 | 1 | 0% | 2,490 | 3,314 | +33% | 0 | 0 | — |
▸case-01 I have recently collected a multi-label dataset of around 12,000 video clips annotated with various genres, scene tags, and actor demographics. I am preparing to train our baseline models and need a comprehensive dataset audit. Please run a thorough analysis on this data. I need you to evaluate our class frequencies and label co-occurrences, analyze potential dataset biases and demographic fairness risks, recommend an optimal train, validation, and test splitting approach to prevent data leakage, outline a concrete data expansion strategy for weak spots, and complete an ethical review checklist before we proceed. | fail→pass | 32,018 | 28,991 | -9% | 1 | 1 | 0% | 6,085 | 6,300 | +4% | 0 | 0 | — |
▸case-14 We audited an annotated named-entity recognition dataset by running unit tests on label formatting. Is running format validation rules sufficient for a complete dataset quality assessment, or what other pillars must be included? | pass→pass | 14,406 | 15,140 | +5% | 1 | 1 | 0% | 2,336 | 3,267 | +40% | 0 | 0 | — |
▸case-06 Our object detection model output bounding boxes with confidence scores against ground truth annotations. How do we compute Mean Average Precision (mAP@0.5:0.95) step-by-step from raw prediction logs? | pass→pass | 15,834 | 23,573 | +49% | 1 | 1 | 0% | 3,130 | 5,621 | +80% | 0 | 0 | — |
▸case-07 In our 10,000-sample image classification dataset, class 'A' has 5,000 images, class 'B' has 1,000, and class 'C' has 150 images. A developer suggested setting the underrepresentation cutoff at 10% of the maximum class count. What threshold percentage identifies severely underrepresented classes in standard distribution analysis? | fail→pass | 21,096 | 5,804 | -72% | 1 | 1 | 0% | 2,052 | 1,992 | -3% | 0 | 0 | — |
▸case-08 We have a small dataset of 2,400 annotated clinical case notes. A colleague suggested using a standard 70/15/15 train/val/test split. What partitioning strategy is preferred for datasets under 5,000 samples? | pass→pass | 12,892 | 13,308 | +3% | 1 | 1 | 0% | 2,210 | 3,133 | +42% | 0 | 0 | — |
▸case-09 We have collected 120,000 audio clips for speech recognition. Should we use 70/15/15 train/val/test splitting or something else suitable for corpora over 50,000 samples? | fail→pass | 22,003 | 17,015 | -23% | 1 | 1 | 0% | 2,294 | 3,799 | +66% | 0 | 0 | — |
▸case-10 We stratified our tabular dataset splits by primary class labels. What statistical hypothesis test should be performed to validate that the label distributions across train, validation, and test splits are statistically similar? | pass→pass | 11,071 | 11,240 | +2% | 1 | 1 | 0% | 1,959 | 2,994 | +53% | 0 | 0 | — |
▸case-11 During analysis of an automated resume screening dataset, we noticed male applicants outnumber female applicants 4:1 in software engineering roles. Before taking remediation steps like resampling or dropping records, what three structured evaluation questions must be asked about this imbalance? | pass→pass | 11,096 | 8,626 | -22% | 1 | 1 | 0% | 1,635 | 2,184 | +34% | 0 | 0 | — |
▸case-12 We are curating a multi-hospital chest X-ray dataset. Standard label stratification was performed, but validation loss is artificially low while test performance drops drastically on new hospitals. How should secondary stratification be applied to prevent data leakage across splits? | pass→pass | 17,698 | 14,643 | -17% | 1 | 1 | 0% | 2,860 | 3,073 | +7% | 0 | 0 | — |
▸case-13 We hired 5 annotators to label sentiment on 10,000 tweets. We want to measure annotation consistency across all annotators for nominal labels. Which inter-annotator agreement metrics are appropriate for multi-annotator datasets? | pass→pass | 13,296 | 16,001 | +20% | 1 | 1 | 0% | 2,344 | 3,583 | +53% | 0 | 0 | — |
▸case-15 Our dataset shows severe gaps in low-resource language translations. We created a list of priority language pairs and suggested scraping target websites. What critical operational factor must be included in the expansion plan alongside priority classes, sources, and collection strategy? | fail→pass | 8,872 | 8,173 | -8% | 1 | 1 | 0% | 1,333 | 2,122 | +59% | 0 | 0 | — |
▸case-16 We completed privacy scrubbing and licensing checks for our open-source speech corpus. What specific documentation artifact is required in the ethical review checklist before dataset release? | pass→pass | 10,538 | 7,970 | -24% | 1 | 1 | 0% | 1,520 | 2,094 | +38% | 0 | 0 | — |
▸case-17 In an automated loan approval dataset, we want to measure whether different demographic groups receive positive outcomes at equal rates regardless of actual ground truth labels. Which specific fairness metric measures this? | pass→pass | 5,671 | 6,517 | +15% | 1 | 1 | 0% | 1,034 | 2,048 | +98% | 0 | 0 | — |
▸case-18 When auditing a recidivism prediction dataset, we need to ensure that true positive rates and false positive rates are equal across protected racial groups. What fairness metric targets both TPR and FPR parity? | pass→pass | 5,677 | 7,743 | +36% | 1 | 1 | 0% | 1,003 | 2,329 | +132% | 0 | 0 | — |
▸case-19 We are partitioning a medium-sized dataset of 25,000 news articles for topic classification. What train/val/test split ratios are standard for medium dataset sizes between 5,000 and 50,000 samples? | fail→pass | 10,975 | 6,806 | -38% | 1 | 1 | 0% | 1,892 | 1,950 | +3% | 0 | 0 | — |
▸case-20 In an object detection dataset for autonomous driving, we built a co-occurrence matrix and discovered that 'person' tags co-occur 99% of the time with 'sidewalk' tags. What issue does this co-occurrence analysis reveal? | pass→pass | 8,622 | 12,156 | +41% | 1 | 1 | 0% | 1,380 | 2,711 | +96% | 0 | 0 | — |
▸case-21 We are preparing a dataset audit report. What six required review items must be evaluated in the ethical checklist before publishing or training? | pass→pass | 10,653 | 3,096 | -71% | 1 | 1 | 0% | 1,666 | 1,361 | -18% | 0 | 0 | — |
▸case-22 We are delivering a dataset curation report to an executive team. What five specific structured sections must be included in the final output format? | fail→pass | 8,833 | 3,169 | -64% | 1 | 1 | 0% | 1,384 | 1,414 | +2% | 0 | 0 | — |