Install any skill in seconds. Free to start, no credit card required.
Get Started Free →End-to-end methodology for AI agents and software engineers to add machine learning algorithms to existing non-ML codebases. Covers problem framing, data readiness, architectural decoupling, and baseline model integration.
.claude/skills/affaan-m-ml-adoption-playbook/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-04 | ✓→✓ | = Same ✓ | -2% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 19% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 18% | 0% |
This skill provides an adaptive methodology for implementing machine learning models into existing software engineering projects. It bridges the gap between traditional SWE and MLOps by structuring how ML should be researched, decoupled, trained, and integrated.
Before writing model code, establish the "why" and "how".
ML is useless without clean, accessible data.
Do not tightly couple model inference to core business logic.
fastapi-patterns or django-patterns) or a dedicated service class.Structure the code for reproducibility and iteration.
pytorch-patterns or similar best practices: fix random seeds, make code device-agnostic, and explicitly document tensor/array shapes.Once the baseline model is integrated, shift focus to continuous operations.
mle-workflow: Guide the user toward setting up experiment tracking, model registries, and drift detection.When assisting a user via this playbook, agents should:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,862 | 16,221 | -22% | 1 | 1 | 0% | 4,175 | 3,967 | -5% | 0 | 0 | — |
case-02 | fail→fail | 25,035 | 17,043 | -32% | 1 | 1 | 0% | 4,593 | 4,191 | -9% | 0 | 0 | — |
case-03 | fail→fail | 25,847 | 19,275 | -25% | 1 | 1 | 0% | 5,353 | 5,002 | -7% | 0 | 0 | — |
case-04 | pass→pass | 14,660 | 14,779 | +1% | 1 | 1 | 0% | 2,261 | 2,213 | -2% | 0 | 0 | — |
case-05 | pass→pass | 12,659 | 11,483 | -9% | 1 | 1 | 0% | 2,439 | 2,891 | +19% | 0 | 0 | — |
case-06 | pass→pass | 13,606 | 10,557 | -22% | 1 | 1 | 0% | 2,165 | 2,561 | +18% | 0 | 0 | — |
case-07 | pass→pass | 16,599 | 12,135 | -27% | 1 | 1 | 0% | 2,901 | 3,089 | +6% | 0 | 0 | — |
case-08 | pass→pass | 7,122 | 4,255 | -40% | 1 | 1 | 0% | 1,326 | 1,562 | +18% | 0 | 0 | — |
case-09 | pass→pass | 11,349 | 14,084 | +24% | 1 | 1 | 0% | 2,357 | 3,534 | +50% | 0 | 0 | — |
case-10 | pass→pass | 19,083 | 13,168 | -31% | 1 | 1 | 0% | 3,153 | 3,459 | +10% | 0 | 0 | — |
case-11 | pass→pass | 13,567 | 11,355 | -16% | 1 | 1 | 0% | 2,605 | 3,039 | +17% | 0 | 0 | — |
case-12 | pass→pass | 7,812 | 3,426 | -56% | 1 | 1 | 0% | 1,274 | 1,425 | +12% | 0 | 0 | — |
case-13 | fail→pass | 13,582 | 9,791 | -28% | 1 | 1 | 0% | 2,237 | 2,373 | +6% | 0 | 0 | — |
case-14 | fail→fail | 16,619 | 8,684 | -48% | 1 | 1 | 0% | 2,702 | 2,371 | -12% | 0 | 0 | — |
case-15 | pass→pass | 11,761 | 7,864 | -33% | 1 | 1 | 0% | 2,068 | 2,070 | +0% | 0 | 0 | — |
case-16 | fail→pass | 11,687 | 8,632 | -26% | 1 | 1 | 0% | 2,196 | 1,496 | -32% | 0 | 0 | — |
case-17 | pass→pass | 13,720 | 12,355 | -10% | 1 | 1 | 0% | 2,391 | 3,033 | +27% | 0 | 0 | — |
case-18 | fail→fail | 7,237 | 2,826 | -61% | 1 | 1 | 0% | 1,271 | 1,258 | -1% | 0 | 0 | — |
case-19 | fail→fail | 8,523 | 2,128 | -75% | 1 | 1 | 0% | 1,414 | 1,169 | -17% | 0 | 0 | — |
case-20 | pass→pass | 13,448 | 7,228 | -46% | 1 | 1 | 0% | 2,323 | 2,040 | -12% | 0 | 0 | — |
case-21 | pass→pass | 8,667 | 6,922 | -20% | 1 | 1 | 0% | 1,672 | 2,070 | +24% | 0 | 0 | — |
case-22 | pass→pass | 12,855 | 13,420 | +4% | 1 | 1 | 0% | 2,336 | 3,363 | +44% | 0 | 0 | — |
case-23 | pass→pass | 4,714 | 6,353 | +35% | 1 | 1 | 0% | 881 | 1,864 | +112% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/3/2026 | +18% |
Other measured skills in the registry, with their headline benchmark lift.