Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Imported from borghei/claude-skills
.claude/skills/borghei-senior-data-scientist/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 16% | 0% |
Expert data science for statistical modeling, experimentation, ML deployment, and data-driven decision making — A/B test design and analysis, feature engineering, model training/evaluation, production deployment, and causal inference.
data-science, machine-learning, statistics, a-b-testing, causal-inference, feature-engineering, mlops, experiment-design, model-deployment, python, scikit-learn, pytorch, tensorflow, spark, airflow
Before running an analysis or pipeline, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
| Script | Purpose | |--------|---------| | scripts/experiment_designer.py | A/B test design, power analysis, sample size calculation | | scripts/feature_engineering_pipeline.py | Automated feature generation, correlation analysis, feature selection | | scripts/statistical_analyzer.py | Hypothesis testing, causal inference, regression analysis | | scripts/model_evaluation_suite.py | Model comparison, cross-validation, deployment readiness checks |
> statistical_analyzer.py is referenced but not yet present in the repo — see the note in references/ds-operations.md. Use inline scipy/statsmodels in the meantime.
Load the reference that matches the task — keep this file lean and pull detail on demand:
This skill covers:
This skill does NOT cover:
senior-data-engineersenior-ml-engineersenior-prompt-engineersenior-computer-vision| Skill | Integration | Data Flow | |-------|-------------|-----------| | senior-data-engineer | Feature pipeline ingests data from ETL outputs; shares data quality validation patterns | Raw data stores --> feature engineering pipeline --> feature store | | senior-ml-engineer | Trained models handed off for MLOps deployment; shares model registry and serving configs | Evaluated model artifacts --> deployment pipeline --> production serving | | senior-prompt-engineer | Embedding features from LLMs feed into ML pipelines; experiment frameworks apply to prompt A/B tests | LLM embeddings --> feature vectors; experiment designs --> prompt evaluation | | senior-architect | Model serving architecture reviewed for scalability; data platform design aligned with training infrastructure | Architecture specs --> deployment topology --> monitoring dashboards | | senior-backend | Model inference endpoints integrated into backend services; API contracts defined for prediction requests | REST/gRPC model API --> backend service layer --> client applications | | senior-devops | CI/CD pipelines extended for model retraining triggers; containerized model images deployed via infrastructure-as-code | Docker images --> Kubernetes manifests --> production clusters |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→fail | 28,470 | 19,541 | -31% | 1 | 1 | 0% | 4,434 | 4,578 | +3% | 0 | 0 | — |
case-01 | fail→pass | 20,771 | 21,290 | +2% | 1 | 1 | 0% | 3,379 | 5,107 | +51% | 0 | 0 | — |
case-02 | fail→fail | 25,165 | 27,111 | +8% | 1 | 1 | 0% | 4,723 | 6,234 | +32% | 0 | 0 | — |
case-03 | fail→fail | 25,358 | 29,988 | +18% | 1 | 1 | 0% | 4,025 | 6,705 | +67% | 0 | 0 | — |
case-04 | fail→pass | 24,541 | 18,147 | -26% | 1 | 1 | 0% | 3,956 | 4,115 | +4% | 0 | 0 | — |
case-06 | fail→pass | 17,331 | 15,371 | -11% | 1 | 1 | 0% | 2,754 | 3,910 | +42% | 0 | 0 | — |
case-07 | pass→pass | 19,400 | 17,264 | -11% | 1 | 1 | 0% | 2,976 | 3,974 | +34% | 0 | 0 | — |
case-08 | pass→pass | 18,180 | 24,544 | +35% | 1 | 1 | 0% | 2,984 | 5,489 | +84% | 0 | 0 | — |
case-09 | fail→pass | 12,396 | 8,066 | -35% | 1 | 1 | 0% | 2,345 | 2,831 | +21% | 0 | 0 | — |
case-10 | pass→pass | 13,205 | 17,125 | +30% | 1 | 1 | 0% | 2,103 | 4,214 | +100% | 0 | 0 | — |
case-11 | fail→pass | 10,642 | 5,428 | -49% | 1 | 1 | 0% | 1,906 | 2,218 | +16% | 0 | 0 | — |
case-12 | pass→pass | 19,028 | 22,700 | +19% | 1 | 1 | 0% | 2,756 | 4,966 | +80% | 0 | 0 | — |
case-13 | fail→fail | 19,151 | 21,804 | +14% | 1 | 1 | 0% | 2,805 | 5,014 | +79% | 0 | 0 | — |
case-14 | pass→pass | 10,535 | 12,994 | +23% | 1 | 1 | 0% | 1,888 | 3,407 | +80% | 0 | 0 | — |
case-15 | pass→pass | 8,156 | 11,338 | +39% | 1 | 1 | 0% | 1,245 | 3,163 | +154% | 0 | 0 | — |
case-16 | fail→pass | 15,046 | 4,412 | -71% | 1 | 1 | 0% | 2,682 | 2,037 | -24% | 0 | 0 | — |
case-17 | pass→pass | 14,391 | 14,590 | +1% | 1 | 1 | 0% | 2,231 | 3,650 | +64% | 0 | 0 | — |
case-18 | pass→pass | 10,957 | 19,819 | +81% | 1 | 1 | 0% | 1,680 | 4,541 | +170% | 0 | 0 | — |
case-19 | pass→pass | 7,417 | 13,193 | +78% | 1 | 1 | 0% | 1,107 | 3,309 | +199% | 0 | 0 | — |
case-20 | pass→pass | 18,722 | 21,778 | +16% | 1 | 1 | 0% | 2,733 | 4,932 | +80% | 0 | 0 | — |
case-21 | pass→pass | 14,460 | 17,402 | +20% | 1 | 1 | 0% | 2,185 | 4,288 | +96% | 0 | 0 | — |
case-22 | fail→pass | 13,607 | 4,108 | -70% | 1 | 1 | 0% | 2,065 | 1,993 | -3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.