Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Builds a leakage-safe tabular classification pipeline end to end: a clean train/test split, preprocessing inside a Pipeline/ColumnTransformer, baseline vs. tuned models with cross-validated hyperparameter search, and honest metrics (accuracy, precision/recall/F1, ROC-AUC, confusion matrix) on a held-out set. Use whenever the user wants to "train a model", "predict a class/label", "build a classifier", or compare models on a labeled tabular dataset — even if they just say "can we predict X from t
.claude/skills/thomson-li-ml-classification-pipeline/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 45% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 35% | 0% |
| case-07 | ✓→✓ | = Same ✓ | 50% | 0% |
A disciplined, reproducible tabular classification workflow that avoids the usual foot-guns (leakage, tuning on the test set, accuracy-only on imbalanced data).
Labeled tabular data + a categorical target the user wants to predict: binary or multiclass. Not for regression (continuous target), unstructured data (text/images), or pure data exploration (use eda-report first).
the class balance. Decide the primary metric up front based on the cost of errors (ROC-AUC / F1 / recall for imbalanced or asymmetric-cost problems; accuracy only when classes are balanced and errors are symmetric).
train_test_split with stratify=y and a fixed random_statebefore any fitting. The test set is touched exactly once, at the end.
ColumnTransformer:numeric → impute + scale; categorical → impute + one-hot (or ordinal where ordered). Everything fit on train folds only — this is what prevents leakage.
DummyClassifier for the floor). Record cross-validated scores. Every later model must beat this to justify its complexity.
booster (HistGradientBoosting / XGBoost / LightGBM if available). Compare via stratified k-fold CV on the training set using the primary metric.
GridSearchCV / RandomizedSearchCV (or Optuna for big spaces) withstratified CV, optimizing the primary metric. Handle imbalance via class_weight='balanced' or resampling inside CV folds — never resample the whole dataset before splitting.
test set, and report the full picture: accuracy, precision/recall/F1 (per class + macro), ROC-AUC (or PR-AUC for heavy imbalance), and the confusion matrix. State the operating threshold if it was tuned.
SHAP if the user wants per-prediction explanations). Note caveats and any drift between CV and test performance.
A runnable, reproducible script/notebook with fixed seeds, plus a short results summary: chosen metric and why, baseline-vs-tuned table, held-out test metrics, confusion matrix, top features, and honest limitations. Persist the fitted pipeline (joblib.dump) if the user wants to reuse it.
stratified, seed=42"), not as bare numbers.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 21,733 | 18,851 | -13% | 1 | 1 | 0% | 4,942 | 4,888 | -1% | 0 | 0 | — |
case-02 | fail→fail | 23,607 | 19,616 | -17% | 1 | 1 | 0% | 4,286 | 5,068 | +18% | 0 | 0 | — |
case-03 | fail→fail | 21,970 | 24,537 | +12% | 1 | 1 | 0% | 5,007 | 6,236 | +25% | 0 | 0 | — |
case-04 | pass→pass | 13,465 | 34,367 | +155% | 1 | 1 | 0% | 2,650 | 3,838 | +45% | 0 | 0 | — |
case-05 | fail→fail | 17,013 | 19,658 | +16% | 1 | 1 | 0% | 3,620 | 4,862 | +34% | 0 | 0 | — |
case-06 | pass→pass | 14,427 | 15,220 | +5% | 1 | 1 | 0% | 2,975 | 4,024 | +35% | 0 | 0 | — |
case-07 | pass→pass | 11,091 | 12,275 | +11% | 1 | 1 | 0% | 1,938 | 2,908 | +50% | 0 | 0 | — |
case-08 | pass→pass | 11,678 | 9,816 | -16% | 1 | 1 | 0% | 1,979 | 2,535 | +28% | 0 | 0 | — |
case-09 | pass→pass | 12,015 | 8,045 | -33% | 1 | 1 | 0% | 2,135 | 2,130 | -0% | 0 | 0 | — |
case-10 | pass→pass | 11,966 | 12,043 | +1% | 1 | 1 | 0% | 2,079 | 2,923 | +41% | 0 | 0 | — |
case-11 | pass→pass | 11,636 | 7,001 | -40% | 1 | 1 | 0% | 1,873 | 1,953 | +4% | 0 | 0 | — |
case-12 | pass→pass | 3,738 | 3,942 | +5% | 1 | 1 | 0% | 617 | 1,489 | +141% | 0 | 0 | — |
case-13 | pass→pass | 10,123 | 10,760 | +6% | 1 | 1 | 0% | 1,842 | 2,573 | +40% | 0 | 0 | — |
case-14 | pass→pass | 6,824 | 5,032 | -26% | 1 | 1 | 0% | 1,281 | 1,597 | +25% | 0 | 0 | — |
case-15 | pass→pass | 5,199 | 5,124 | -1% | 1 | 1 | 0% | 880 | 1,651 | +88% | 0 | 0 | — |
case-16 | fail→pass | 13,914 | 12,274 | -12% | 1 | 1 | 0% | 2,231 | 2,896 | +30% | 0 | 0 | — |
case-17 | fail→pass | 6,999 | 5,555 | -21% | 1 | 1 | 0% | 1,154 | 1,575 | +36% | 0 | 0 | — |
case-18 | pass→pass | 12,729 | 9,722 | -24% | 1 | 1 | 0% | 2,150 | 2,361 | +10% | 0 | 0 | — |
case-19 | pass→pass | 13,223 | 12,200 | -8% | 1 | 1 | 0% | 2,256 | 2,810 | +25% | 0 | 0 | — |
case-20 | pass→pass | 12,513 | 11,154 | -11% | 1 | 1 | 0% | 2,457 | 2,566 | +4% | 0 | 0 | — |
case-21 | fail→fail | 12,301 | 10,504 | -15% | 1 | 1 | 0% | 2,120 | 2,523 | +19% | 0 | 0 | — |
case-22 | fail→fail | 11,125 | 11,987 | +8% | 1 | 1 | 0% | 1,940 | 2,688 | +39% | 0 | 0 | — |
case-23 | pass→pass | 13,190 | 4,826 | -63% | 1 | 1 | 0% | 1,640 | 1,633 | -0% | 0 | 0 | — |
case-24 | pass→pass | 14,639 | 11,751 | -20% | 1 | 1 | 0% | 2,374 | 2,657 | +12% | 0 | 0 | — |
case-25 | pass→pass | 5,184 | 5,736 | +11% | 1 | 1 | 0% | 962 | 1,816 | +89% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +8 percentage points is the difference between those two pass rates over the 25 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.