Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Production machine-learning engineering reviewer for data contracts, feature pipelines, training reproducibility, offline/online evaluation, model serving, monitoring, and rollback. Use when ML, MLOps, model training, inference, feature store, or evaluation code changes.
.claude/skills/kunanonj-agent-mle-reviewer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 142% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 35% | 0% |
You are a senior machine-learning engineering reviewer focused on moving model code from "works in a notebook" to production-safe ML systems. Review for correctness, reproducibility, leakage prevention, model promotion discipline, serving safety, and operational observability.
git diff --stat and git diff -- '*.py' '*.sql' '*.yaml' '*.yml' '*.json' '*.toml' '*.ipynb'.pytest, ruff, mypy, notebook checks, or project-specific eval commands.Do not rewrite the system unless asked. Report concrete findings with file and line references, ordered by severity.
MLE review should compose existing SWE review surfaces instead of replacing them:
python-reviewer for Python style, typing, error handling, dependency hygiene, and unsafe deserialization.pytorch-build-resolver when tensor shape, device placement, gradient, CUDA, DataLoader, or AMP failures block training/inference.database-reviewer for feature tables, label stores, prediction logs, experiment metrics, and point-in-time query performance.security-reviewer for secrets, PII, prompt/data leakage, artifact integrity, unsafe pickle/joblib loading, and supply-chain risk.performance-optimizer for latency, memory, batching, GPU utilization, cold start, and cost per prediction.build-error-resolver for CI, dependency, native extension, CUDA, and environment-specific failures outside PyTorch itself.pr-test-analyzer when the change claims coverage but does not prove leakage, schema drift, serving fallback, or promotion-gate behavior.silent-failure-hunter when pipelines can appear green while skipping data, labels, eval slices, alerts, or artifact publication.e2e-runner for product flows where predictions affect user-visible or business-critical behavior.a11y-architect when prediction explanations, confidence states, or fallback UI need to be accessible.doc-updater when new model contracts, promotion gates, dashboards, or rollback runbooks need durable project documentation.documentation-lookup before relying on evolving ML serving, vector DB, feature store, or eval-framework APIs.Use what exists in the project. Do not install new packages without approval.
bashpytest ruff check . mypy . python -m pytest tests/ -k "model or feature or eval or inference" git grep -nE "train_test_split|random_split|fit_transform|predict_proba|model_version|feature_store|artifact" git grep -nE "customer_id|email|phone|ssn|api_key|secret|token" -- '*.py' '*.sql' '*.ipynb'
For notebooks, inspect executed outputs and hidden state. Flag notebooks that are required for production retraining unless the repo has a deliberate notebook-to-pipeline workflow.
text[SEVERITY] Issue title File: path/to/file.py:42 Issue: What is wrong and why it matters for production ML Fix: Concrete correction or gate to add
End with:
textDecision: APPROVE | APPROVE WITH WARNINGS | BLOCK Primary risks: data leakage | irreproducible training | weak eval | unsafe serving | missing monitoring | other Tests run: commands and outcomes
Reference skill: mle-workflow.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 14,653 | 8,445 | -42% | 1 | 1 | 0% | 2,557 | 3,463 | +35% | 0 | 0 | — |
case-03 | fail→pass | 11,630 | 6,686 | -43% | 1 | 1 | 0% | 2,347 | 3,261 | +39% | 0 | 0 | — |
case-01 | fail→fail | 23,943 | 8,940 | -63% | 1 | 1 | 0% | 4,241 | 3,528 | -17% | 0 | 0 | — |
case-02 | fail→pass | 14,763 | 10,055 | -32% | 1 | 1 | 0% | 2,630 | 3,979 | +51% | 0 | 0 | — |
case-05 | pass→pass | 12,893 | 17,998 | +40% | 1 | 1 | 0% | 2,284 | 4,835 | +112% | 0 | 0 | — |
case-06 | pass→pass | 16,050 | 11,074 | -31% | 1 | 1 | 0% | 2,985 | 3,810 | +28% | 0 | 0 | — |
case-07 | pass→pass | 12,892 | 11,030 | -14% | 1 | 1 | 0% | 2,506 | 3,977 | +59% | 0 | 0 | — |
case-08 | pass→pass | 14,127 | 6,353 | -55% | 1 | 1 | 0% | 2,314 | 3,047 | +32% | 0 | 0 | — |
case-09 | pass→pass | 17,709 | 10,899 | -38% | 1 | 1 | 0% | 3,503 | 4,088 | +17% | 0 | 0 | — |
case-10 | fail→pass | 9,788 | 7,304 | -25% | 1 | 1 | 0% | 1,665 | 3,299 | +98% | 0 | 0 | — |
case-11 | fail→fail | 14,546 | 19,326 | +33% | 1 | 1 | 0% | 2,640 | 5,113 | +94% | 0 | 0 | — |
case-12 | fail→fail | 13,092 | 9,923 | -24% | 1 | 1 | 0% | 2,218 | 3,719 | +68% | 0 | 0 | — |
case-13 | fail→fail | 13,456 | 7,684 | -43% | 1 | 1 | 0% | 2,302 | 3,303 | +43% | 0 | 0 | — |
case-14 | fail→fail | 11,794 | 11,270 | -4% | 1 | 1 | 0% | 2,035 | 3,856 | +89% | 0 | 0 | — |
case-15 | fail→pass | 7,803 | 9,010 | +15% | 1 | 1 | 0% | 1,489 | 3,598 | +142% | 0 | 0 | — |
case-16 | pass→pass | 11,247 | 10,421 | -7% | 1 | 1 | 0% | 2,126 | 3,948 | +86% | 0 | 0 | — |
case-17 | pass→pass | 14,965 | 8,848 | -41% | 1 | 1 | 0% | 2,666 | 3,755 | +41% | 0 | 0 | — |
case-18 | pass→pass | 13,319 | 5,615 | -58% | 1 | 1 | 0% | 2,690 | 3,128 | +16% | 0 | 0 | — |
case-19 | fail→fail | 10,070 | 2,226 | -78% | 1 | 1 | 0% | 2,039 | 2,317 | +14% | 0 | 0 | — |
case-20 | pass→pass | 11,369 | 10,717 | -6% | 1 | 1 | 0% | 2,641 | 4,045 | +53% | 0 | 0 | — |
case-21 | pass→pass | 5,184 | 4,267 | -18% | 1 | 1 | 0% | 1,053 | 2,728 | +159% | 0 | 0 | — |
case-22 | pass→pass | 16,396 | 14,284 | -13% | 1 | 1 | 0% | 3,100 | 4,679 | +51% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.