Loading skill
Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 198% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-20 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 12% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 11% | 0% |
Use this skill as the direct owner for ML input-preparation pipelines.
It covers preprocessing-heavy tasks where the requested deliverable is a repeatable pipeline for cleaning, encoding, transforming, and validating input data.
Use this skill when:
scikit-learn or ml-pipeline-workflowml-data-leakage-guardscientific-data-preprocessingml-data-leakage-guard before trusting fitted preprocessing stepssplitting-datasets when the next narrow problem is partition strategyOther measured skills in the registry, with their headline benchmark lift.