▸case-01 I need a complete model card for our ChurnShield v1.2 model developed by the Retention Engineering team. It's a binary classification model that predicts 30-day customer churn to inform targeted discount offers. We trained it on 2 million user sessions from 2023-2024 with labels based on explicit subscription cancellations. In our evaluations, it hit a 0.88 ROC-AUC compared to the 0.72 baseline rules engine, though performance varies when broken down across web, mobile, and enterprise tiers. Its main failure mode is overpredicting churn for accounts under 7 days old. Our monitoring setup tracks daily precision, and dropping below 0.70 triggers an automatic rollback. Can you compile this into a formal model card for our upcoming governance review? | fail→pass | 13,663 | 15,986 | +17% | 1 | 1 | 0% | 2,702 | 3,773 | +40% | 0 | 0 | — |
▸case-02 Please draft a model card for Risk Analytics' transaction scoring model, SentinelNet v2.0. The model assigns a risk score to real-time credit card transactions to assist manual fraud review teams, but must explicitly not be used for automated account terminations. It was trained on 10 million anonymized transactions from Q1-Q3 2023. For evaluation, we measured Precision and Recall against our legacy rules baseline, evaluating performance across North America, Europe, and APAC regions. A key risk is elevated false positives during major shopping events. Latency target is under 50ms, and we track daily feature drift to trigger rollbacks. Format this into a proper model documentation artifact. | fail→pass | 13,291 | 25,581 | +92% | 1 | 1 | 0% | 2,637 | 3,246 | +23% | 0 | 0 | — |
▸case-03 We are preparing for a responsible AI review and need a model card generated for MedExtract-LLM v0.5, built by the Clinical NLP team. The model extracts medical entities from physician notes to assist billing specialists. It is strictly not validated for clinical diagnosis or direct patient treatment decisions. It was trained on 500,000 anonymized EHR notes from 2018 to 2022. Overall F1 score is 0.91 compared to the 0.79 heuristic baseline, with sliced results across Cardiology, Oncology, and Pediatrics specialties. It frequently fails on non-standard medical shorthand and poor OCR text. Latency is capped at 200ms, with daily extraction quality monitoring. Please assemble this into the full model card format. | fail→pass | 13,143 | 15,281 | +16% | 1 | 1 | 0% | 2,574 | 3,540 | +38% | 0 | 0 | — |
▸case-04 We need documentation for our training corpus `WebText-Cleaned-v3`. Please draft a Dataset Datasheet detailing data collection methodology, preprocessing pipelines, labeling noise rates, and copyright attribution. | pass→pass | 22,997 | 19,943 | -13% | 1 | 1 | 0% | 4,029 | 4,135 | +3% | 0 | 0 | — |
▸case-05 Please draft a Technical Design Document for our multi-node Triton inference cluster serving transformer models, focusing on GPU memory allocation, batching limits, and Kubernetes ingress routing. | pass→pass | 35,433 | 34,701 | -2% | 1 | 1 | 0% | 6,174 | 7,006 | +13% | 0 | 0 | — |
▸case-06 Write a PyTorch training loop script using Hugging Face Accelerate to fine-tune Llama-3-8B on customer support ticket transcripts with LoRA adapters. | pass→fail | 18,731 | 22,200 | +19% | 1 | 1 | 0% | 4,377 | 5,563 | +27% | 0 | 0 | — |
▸case-07 Draft a model card for ScoreGuard v3, a credit risk assessment model by Consumer Risk Engineering. It predicts 90-day delinquency risk for loan applicants. It must never be used for automated employment screening. Trained on 1.5M credit bureau records (2021-2023). Overall Gini coefficient is 0.64 vs 0.51 legacy score, evaluated across income quartiles and age brackets. High failure rate on zero-credit-history applicants. Rollback condition: Gini drop below 0.55 on weekly audit. | fail→pass | 12,968 | 13,161 | +1% | 1 | 1 | 0% | 2,605 | 3,500 | +34% | 0 | 0 | — |
▸case-08 Create model documentation for SupportRoute v1, an intent classifier for customer service tickets built by CX Engineering. It routes tickets to agent queues, but should not auto-close high-value enterprise tickets. Trained on 800k tickets from 2022-2024. Macro F1 is 0.84 vs 0.68 keyword matcher baseline, sliced by tier (Free, Pro, Enterprise). Fails on sarcastic user complaints. Rollback trigger: accuracy drop over 5% daily. | pass→pass | 11,081 | 14,884 | +34% | 1 | 1 | 0% | 2,184 | 3,564 | +63% | 0 | 0 | — |
▸case-09 Generate a model card for TalentMatch v2.1 developed by HR Tech Labs. It extracts skills and work history from PDF resumes to assist recruiters. It must not be used for automated candidate ranking or rejection. Trained on 300,000 resume documents. Precision is 0.92 vs 0.75 regex parser baseline, evaluated across candidate experience levels (Entry, Mid, Executive). Underperforms on non-standard PDF formatting. Monitoring tracks weekly extraction accuracy, rolling back if F1 drops below 0.85. | fail→pass | 11,344 | 10,580 | -7% | 1 | 1 | 0% | 2,194 | 2,785 | +27% | 0 | 0 | — |
▸case-10 Prepare model documentation for AdRank-CTR v4.0 by the Monetization Team. Predicts click-through rate for sponsored listings. Prohibited from being used for user credit or housing ad delivery filtering. Trained on 50M impression logs from past 30 days. Log loss is 0.12 vs 0.18 baseline, sliced by device type (iOS, Android, Desktop). Fails on newly launched merchant catalogs. Latency max 15ms. Rollback trigger: Log loss spike above 0.16 over 1 hour. | fail→fail | 9,375 | 12,834 | +37% | 1 | 1 | 0% | 1,969 | 3,232 | +64% | 0 | 0 | — |
▸case-11 Produce a model card for VisionGuard v1.0 by Trust & Safety. Detects policy-violating imagery in user uploads to queue for human review. Explicitly out of scope for automated account bans without human verification. Trained on 4M labeled images. Recall is 0.94 vs 0.81 legacy filter baseline, evaluated across image categories and lighting conditions. Struggles with heavy cartoon/illustration styles. Rollback trigger: false positive rate exceeding 2% in daily audit. | fail→pass | 11,564 | 14,011 | +21% | 1 | 1 | 0% | 2,188 | 3,549 | +62% | 0 | 0 | — |
▸case-12 Create a model card for VoiceScribe-Medical v2 by Audio AI Team. Transcribes doctor-patient consultations for clinical drafting. Unvalidated for legal deposition transcription. Trained on 12,000 hours of clinical audio. WER is 8.2% vs 14.5% open-source baseline, evaluated across accent groups and background noise levels. High error rate on heavy ambient HVAC noise. Rollback trigger: WER exceeding 12% on daily validation sample. | fail→fail | 11,336 | 12,029 | +6% | 1 | 1 | 0% | 2,062 | 3,155 | +53% | 0 | 0 | — |
▸case-13 Draft model documentation for CodeAssist-Python v0.9 by Developer Tools Team. Provides inline autocomplete suggestions in IDEs. Should not be used for security vulnerability scanning or auto-committing code. Trained on 100M open source Python files. Pass@1 is 48% vs 32% legacy model baseline, sliced by framework (Django, Flask, FastApi, Base Python). Fails on custom internal proprietary SDKs. Rollback trigger: user acceptance rate falling below 25%. | fail→pass | 9,269 | 11,362 | +23% | 1 | 1 | 0% | 1,894 | 3,033 | +60% | 0 | 0 | — |
▸case-14 Provide a model card for ClaimShield v3 by Insurance Analytics. Scores auto insurance claims for potential fraud investigation. Must not auto-deny claims. Trained on 500,000 historical claims (2019-2023). ROC-AUC is 0.86 vs 0.71 baseline, sliced by region and claim amount tier. Failure mode: misclassifies catastrophic weather event clusters. Rollback trigger: weekly recall dropping below 0.75. | fail→pass | 12,808 | 12,489 | -2% | 1 | 1 | 0% | 2,440 | 3,335 | +37% | 0 | 0 | — |
▸case-15 Write a model card for BrandPulse v2 by Marketing Intelligence. Classifies social media brand mentions as positive, negative, or neutral. Prohibited from assessing employee internal communications. Trained on 3M tweets and forum posts. Macro F1 is 0.78 vs 0.62 baseline, sliced by platform (X, Reddit, Instagram). Struggles with heavy irony and slang. Rollback trigger: daily drift metric exceeding threshold 0.15. | pass→pass | 11,189 | 11,263 | +1% | 1 | 1 | 0% | 2,075 | 2,859 | +38% | 0 | 0 | — |
▸case-16 Construct a model card for InventoryProphet v1.4 by Supply Chain AI. Forecasts weekly store SKU demand to guide warehouse replenishment. Prohibited from setting dynamic retail shelf pricing. Trained on 5 years of sales data. WAPE is 11.2% vs 18.5% moving average baseline, sliced by store format and product category. Underperforms during unseasonable weather events. Rollback trigger: WAPE rising above 16% across two consecutive weeks. | fail→pass | 12,239 | 15,694 | +28% | 1 | 1 | 0% | 2,378 | 3,831 | +61% | 0 | 0 | — |
▸case-17 Draft a model card for BioLink v2 by KB Engineering. Links biomedical entities in scientific literature to UMLS concepts. Not validated for real-time patient triage. Trained on 2M PubMed abstracts. Accuracy is 0.89 vs 0.74 dictionary baseline, sliced by entity type (Gene, Disease, Chemical). Fails on newly coined acronyms. Rollback trigger: accuracy dropping below 0.80 on weekly benchmark. | fail→pass | 9,700 | 10,866 | +12% | 1 | 1 | 0% | 1,856 | 3,115 | +68% | 0 | 0 | — |
▸case-18 Generate a model card for RecoStream v5 by Video Platform Team. Recommends next videos to watch. Out of scope for news credibility ranking or political feed curation. Trained on 1B viewing events. Implicit CTR gain is +14% vs popularity baseline, sliced by user tenure (New, Active, Dormant). Underperforms on niche language content. Rollback trigger: engagement drop greater than 3% in hourly telemetry. | fail→fail | 7,765 | 10,660 | +37% | 1 | 1 | 0% | 1,558 | 2,832 | +82% | 0 | 0 | — |
▸case-19 Create model documentation for DocStructure v1 by Automation Team. Identifies tables, headers, and paragraphs in scanned PDFs. Not validated for passport or ID forgery detection. Trained on 100,000 scanned business forms. mAP is 0.87 vs 0.71 heuristic baseline, sliced by document type (Invoice, Receipt, Contract). High failure rate on skewed or low-DPI mobile phone captures. Rollback trigger: parsing success rate under 80%. | fail→pass | 10,094 | 13,471 | +33% | 1 | 1 | 0% | 1,981 | 3,367 | +70% | 0 | 0 | — |
▸case-20 Write a model card for TransTranslate-EN-ES v2 by Localization Engineering. Translates user interface text from English to Spanish. Prohibited from translating legal contracts without attorney review. Trained on 20M parallel sentences. BLEU score is 38.4 vs 31.2 baseline, sliced by domain (General UI, Marketing, Technical Docs). Underperforms on highly condensed microcopy. Rollback trigger: BLEU dropping below 34 on automated test suite. | fail→pass | 11,039 | 12,025 | +9% | 1 | 1 | 0% | 2,289 | 3,148 | +38% | 0 | 0 | — |
▸case-21 Draft a model card for TelcoKeep v1.1 by Customer Retention. Predicts mobile subscriber cancellation risk. Must not be used to restrict service access or downgrade tiers. Trained on 1.2M subscriber accounts over 24 months. ROC-AUC is 0.82 vs 0.69 rule-based baseline, sliced by contract type (Prepaid, Postpaid, Family). Fails on corporate enterprise accounts. Rollback trigger: monthly precision dropping below 0.65. | fail→pass | 10,623 | 11,789 | +11% | 1 | 1 | 0% | 2,122 | 3,187 | +50% | 0 | 0 | — |
▸case-22 Prepare model documentation for SearchRank v3.2 by Search Core Team. Ranks e-commerce product search results. Out of scope for organic search engine SEO manipulation. Trained on 50M search sessions. NDCG@10 is 0.76 vs 0.64 BM25 baseline, sliced by query intent type (Navigational, Transactional, Informational). Fails on long-tail misspellings. Rollback trigger: NDCG@10 falling below 0.70 in live bucket. | pass→pass | 16,760 | 12,654 | -24% | 1 | 1 | 0% | 3,006 | 3,212 | +7% | 0 | 0 | — |