Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Runs a structured exploratory data analysis on a tabular dataset (CSV / Excel / Parquet / DataFrame) and produces a readable EDA report with shape, dtypes, missingness, distributions, correlations, outliers, and target relationships. Use whenever the user hands over a dataset and wants to "explore", "profile", "understand", "sanity-check", or "look at" it before modeling — even if they don't say the words "EDA" or "exploratory".
.claude/skills/thomson-li-eda-report/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-16 | ✓→✗ | ▼ Worse | 27% | 0% |
| case-18 | ✓→✗ | ▼ Worse | 29% | 0% |
| case-07 | ✓→✓ | = Same ✓ | -7% | 0% |
Profile an unfamiliar tabular dataset and surface what matters before any modeling.
The user has a dataset and wants to understand its shape and quality: new data file, "what's in here?", "any data issues?", "should I clean anything?", pre-modeling sanity checks. Not for production dashboards (use a dashboard skill) or for the modeling itself (use ml-classification-pipeline).
pandas.read_csv/read_excel/read_parquet).Print shape and the first few rows. Confirm the intended target column with the user if one exists; if unclear, ask before going deep.
Flag columns that are all-null, constant, or near-constant (≥ 99% one value), and likely-ID columns (unique ≈ row count).
Distinguish "missing at random" vs structural missingness if patterns are obvious.
skew, and outliers via IQR (values beyond Q1−1.5·IQR or Q3+1.5·IQR). Call out suspicious values (negatives where impossible, placeholder codes like -999, 0 as missing).
high-cardinality, report cardinality and top-k. Flag rare categories.
redundancy/multicollinearity risk). If a target is set: feature-vs-target relationships (group means for classification, correlations for regression) and class balance for classification targets.
rows, inconsistent units, or mixed types within a column.
A concise EDA report (markdown by default; HTML if the user wants something shareable) with these sections: Overview → Data quality → Distributions → Relationships → Flags & recommendations. End with a short, prioritized "what to fix / watch before modeling" list. Keep prose tight — tables and bullets over paragraphs. Save plots only if the user wants visuals; otherwise describe findings numerically.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 21,722 | 5,859 | -73% | 1 | 1 | 0% | 4,144 | 942 | -77% | 0 | 0 | — |
case-02 | fail→fail | 15,687 | 34,095 | +117% | 1 | 1 | 0% | 2,127 | 1,114 | -48% | 0 | 0 | — |
case-03 | fail→fail | 19,540 | 7,632 | -61% | 1 | 1 | 0% | 3,809 | 1,001 | -74% | 0 | 0 | — |
case-04 | fail→fail | 19,036 | 19,027 | -0% | 1 | 1 | 0% | 4,603 | 4,813 | +5% | 0 | 0 | — |
case-05 | fail→fail | 10,733 | 15,923 | +48% | 1 | 1 | 0% | 2,320 | 3,865 | +67% | 0 | 0 | — |
case-06 | fail→fail | 14,371 | 15,538 | +8% | 1 | 1 | 0% | 3,333 | 3,759 | +13% | 0 | 0 | — |
case-07 | pass→pass | 15,960 | 18,415 | +15% | 1 | 1 | 0% | 2,738 | 2,548 | -7% | 0 | 0 | — |
case-08 | pass→pass | 9,446 | 4,179 | -56% | 1 | 1 | 0% | 1,463 | 1,296 | -11% | 0 | 0 | — |
case-09 | pass→pass | 8,299 | 6,857 | -17% | 1 | 1 | 0% | 1,486 | 1,783 | +20% | 0 | 0 | — |
case-10 | pass→pass | 6,934 | 5,022 | -28% | 1 | 1 | 0% | 1,212 | 1,507 | +24% | 0 | 0 | — |
case-11 | pass→pass | 12,038 | 6,593 | -45% | 1 | 1 | 0% | 2,169 | 1,652 | -24% | 0 | 0 | — |
case-12 | pass→pass | 11,206 | 10,007 | -11% | 1 | 1 | 0% | 2,103 | 2,342 | +11% | 0 | 0 | — |
case-13 | pass→pass | 12,036 | 6,438 | -47% | 1 | 1 | 0% | 2,057 | 1,817 | -12% | 0 | 0 | — |
case-14 | pass→pass | 12,735 | 7,847 | -38% | 1 | 1 | 0% | 1,828 | 1,876 | +3% | 0 | 0 | — |
case-15 | fail→pass | 11,729 | 8,066 | -31% | 1 | 1 | 0% | 2,070 | 1,909 | -8% | 0 | 0 | — |
case-16 | pass→fail | 12,580 | 12,305 | -2% | 1 | 1 | 0% | 2,210 | 2,817 | +27% | 0 | 0 | — |
case-17 | fail→pass | 9,813 | 6,732 | -31% | 1 | 1 | 0% | 1,747 | 1,842 | +5% | 0 | 0 | — |
case-18 | pass→fail | 5,495 | 3,731 | -32% | 1 | 1 | 0% | 894 | 1,154 | +29% | 0 | 0 | — |
case-19 | pass→pass | 9,783 | 5,465 | -44% | 1 | 1 | 0% | 1,636 | 1,446 | -12% | 0 | 0 | — |
case-20 | pass→pass | 11,039 | 8,239 | -25% | 1 | 1 | 0% | 1,820 | 1,943 | +7% | 0 | 0 | — |
case-21 | pass→pass | 14,225 | 10,865 | -24% | 1 | 1 | 0% | 2,329 | 2,371 | +2% | 0 | 0 | — |
case-22 | pass→pass | 10,782 | 7,537 | -30% | 1 | 1 | 0% | 1,920 | 1,868 | -3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 19 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/3/2026 | +23% |
Other measured skills in the registry, with their headline benchmark lift.