Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Intelligent companion for ML engineering with arXiv integration
.claude/skills/brycewang-stanford-mle-agent-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-02 | ✓→✗ | ▼ Worse | 54% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 58% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 52% | 0% |
| case-22 | ✓→✓ | = Same ✓ | 54% | 0% |
A skill for using an intelligent ML engineering companion that integrates arXiv paper discovery with experiment implementation, tracking, and iteration. Based on MLE-agent (2K stars), this skill helps researchers bridge the gap between reading about new ML techniques and implementing them in their own projects.
Machine learning research moves at an extraordinary pace, with hundreds of new papers appearing on arXiv daily. Researchers struggle not just to keep up with the literature but to translate promising ideas into working implementations. MLE-agent addresses this by combining paper discovery, technique extraction, implementation assistance, and experiment management into a unified workflow.
This skill is designed for ML researchers and engineers who want to quickly prototype ideas from papers, systematically compare approaches, and maintain organized experiment records throughout the research process.
The skill provides sophisticated arXiv paper discovery and analysis:
Paper Discovery
Paper Analysis
Technique Extraction
The core experiment management workflow:
Project Setup
Implementation Assistance
Experiment Execution
Result Analysis
The skill enforces ML engineering standards throughout the workflow:
Reproducibility
Code Quality
Resource Management
The skill recognizes and supports common research patterns:
Baseline Comparison - Implement and evaluate standard baselines before proposing improvements Ablation Study - Systematically remove or vary components to understand contributions Scaling Analysis - Test how performance changes with model size, data size, or compute Transfer Learning - Adapt pretrained models to new tasks with appropriate fine-tuning strategies Ensemble Methods - Combine multiple models for improved and more robust performance
This skill connects with the Research-Claw ecosystem:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 15,037 | 14,901 | -1% | 1 | 1 | 0% | 2,046 | 3,223 | +58% | 0 | 0 | — |
case-02 | pass→fail | 15,585 | 20,653 | +33% | 1 | 1 | 0% | 2,774 | 4,264 | +54% | 0 | 0 | — |
case-03 | pass→pass | 12,121 | 14,154 | +17% | 1 | 1 | 0% | 2,038 | 3,096 | +52% | 0 | 0 | — |
case-22 | pass→pass | 5,353 | 3,472 | -35% | 1 | 1 | 0% | 1,048 | 1,619 | +54% | 0 | 0 | — |
case-04 | fail→fail | 9,703 | 9,141 | -6% | 1 | 1 | 0% | 1,923 | 2,770 | +44% | 0 | 0 | — |
case-05 | pass→pass | 18,868 | 19,453 | +3% | 1 | 1 | 0% | 2,814 | 4,098 | +46% | 0 | 0 | — |
case-06 | pass→pass | 15,463 | 12,817 | -17% | 1 | 1 | 0% | 2,347 | 3,317 | +41% | 0 | 0 | — |
case-07 | pass→pass | 20,735 | 23,444 | +13% | 1 | 1 | 0% | 3,273 | 4,849 | +48% | 0 | 0 | — |
case-08 | pass→pass | 14,385 | 17,868 | +24% | 1 | 1 | 0% | 2,481 | 3,832 | +54% | 0 | 0 | — |
case-09 | pass→pass | 11,114 | 15,452 | +39% | 1 | 1 | 0% | 1,848 | 3,224 | +74% | 0 | 0 | — |
case-10 | fail→pass | 17,734 | 10,441 | -41% | 1 | 1 | 0% | 2,520 | 2,693 | +7% | 0 | 0 | — |
case-11 | pass→pass | 10,816 | 8,704 | -20% | 1 | 1 | 0% | 1,781 | 2,782 | +56% | 0 | 0 | — |
case-12 | pass→pass | 9,145 | 7,589 | -17% | 1 | 1 | 0% | 1,514 | 2,446 | +62% | 0 | 0 | — |
case-13 | pass→pass | 16,208 | 15,769 | -3% | 1 | 1 | 0% | 2,438 | 3,666 | +50% | 0 | 0 | — |
case-14 | pass→pass | 14,355 | 14,161 | -1% | 1 | 1 | 0% | 1,968 | 3,352 | +70% | 0 | 0 | — |
case-15 | pass→pass | 13,765 | 5,640 | -59% | 1 | 1 | 0% | 1,934 | 2,007 | +4% | 0 | 0 | — |
case-16 | pass→pass | 13,171 | 7,675 | -42% | 1 | 1 | 0% | 1,857 | 2,307 | +24% | 0 | 0 | — |
case-17 | pass→pass | 6,818 | 11,490 | +69% | 1 | 1 | 0% | 1,249 | 2,796 | +124% | 0 | 0 | — |
case-18 | pass→pass | 8,059 | 10,595 | +31% | 1 | 1 | 0% | 1,639 | 2,882 | +76% | 0 | 0 | — |
case-19 | pass→pass | 15,738 | 18,296 | +16% | 1 | 1 | 0% | 2,481 | 4,291 | +73% | 0 | 0 | — |
case-20 | pass→pass | 13,072 | 12,973 | -1% | 1 | 1 | 0% | 1,912 | 2,903 | +52% | 0 | 0 | — |
case-21 | pass→pass | 2,578 | 1,642 | -36% | 1 | 1 | 0% | 369 | 1,349 | +266% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of 0 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.