Install any skill in seconds. Free to start, no credit card required.
Get Started Free →10 statistical analysis skills. Trigger: statistical tests, Bayesian analysis, hypothesis testing, sampling. Design: method guides covering assumptions, code, and result interpretation.
.claude/skills/brycewang-stanford-statistics-skills/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | -37% | 0% |
| case-21 | ✓→✗ | ▼ Worse | -49% | 0% |
| case-22 | ✓→✗ | ▼ Worse | -61% | 0% |
| case-23 | ✓→✗ | ▼ Worse | -77% | 0% |
| case-01 | ✗→✗ | = Same ✗ | -61% | 0% |
Select the skill matching the user's need, then read its SKILL.md.
| Skill | Description | |-------|-------------| | bayesian-statistics-guide | Bayesian inference methods including prior selection, MCMC, and model comparison | | data-anomaly-detection | Detect anomalies and outliers in research data using statistical methods | | hypothesis-testing-guide | Statistical hypothesis testing, power analysis, and significance reporting | | meta-analysis-guide | Conduct systematic meta-analyses with effect size pooling and heterogeneity | | ml-experiment-tracker | Plan reproducible ML experiment runs with parameters and metrics tracking | | modeling-strategy-guide | Strategic statistical modeling, experimentation, and causal inference | | nonparametric-tests-guide | Apply Mann-Whitney, Kruskal-Wallis, and other nonparametric methods | | power-analysis-guide | Sample size calculation and statistical power analysis guide | | sem-guide | Structural equation modeling with latent variables guide | | survival-analysis-guide | Conduct Kaplan-Meier, Cox regression, and time-to-event analyses |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 12,235 | 14,790 | +21% | 1 | 1 | 0% | 1,888 | 741 | -61% | 0 | 0 | — |
case-02 | fail→fail | 21,259 | 15,855 | -25% | 1 | 1 | 0% | 2,599 | 737 | -72% | 0 | 0 | — |
case-03 | fail→fail | 14,475 | 15,004 | +4% | 1 | 1 | 0% | 2,592 | 785 | -70% | 0 | 0 | — |
case-04 | fail→fail | 14,374 | 10,443 | -27% | 1 | 1 | 0% | 2,644 | 743 | -72% | 0 | 0 | — |
case-05 | fail→fail | 15,715 | 13,228 | -16% | 1 | 1 | 0% | 277 | 746 | +169% | 0 | 0 | — |
case-06 | fail→fail | 15,631 | 13,838 | -11% | 1 | 1 | 0% | 3,226 | 759 | -76% | 0 | 0 | — |
case-07 | fail→fail | 27,791 | 15,654 | -44% | 1 | 1 | 0% | 2,726 | 786 | -71% | 0 | 0 | — |
case-08 | fail→fail | 23,603 | 11,987 | -49% | 1 | 1 | 0% | 3,517 | 716 | -80% | 0 | 0 | — |
case-09 | fail→fail | 12,012 | 12,239 | +2% | 1 | 1 | 0% | 1,363 | 757 | -44% | 0 | 0 | — |
case-10 | fail→fail | 25,894 | 14,127 | -45% | 1 | 1 | 0% | 3,071 | 542 | -82% | 0 | 0 | — |
case-11 | fail→fail | 20,094 | 10,068 | -50% | 1 | 1 | 0% | 2,277 | 620 | -73% | 0 | 0 | — |
case-12 | fail→fail | 4,656 | 16,861 | +262% | 1 | 1 | 0% | 154 | 709 | +360% | 0 | 0 | — |
case-13 | fail→fail | 17,382 | 9,953 | -43% | 1 | 1 | 0% | 2,262 | 530 | -77% | 0 | 0 | — |
case-14 | fail→fail | 37,914 | 4,608 | -88% | 1 | 1 | 0% | 5,853 | 604 | -90% | 0 | 0 | — |
case-15 | fail→fail | 16,283 | 6,574 | -60% | 1 | 1 | 0% | 327 | 728 | +123% | 0 | 0 | — |
case-16 | fail→fail | 23,856 | 10,040 | -58% | 1 | 1 | 0% | 4,155 | 605 | -85% | 0 | 0 | — |
case-17 | fail→fail | 17,230 | 3,408 | -80% | 1 | 1 | 0% | 1,647 | 533 | -68% | 0 | 0 | — |
case-18 | fail→fail | 38,007 | 7,114 | -81% | 1 | 1 | 0% | 6,516 | 762 | -88% | 0 | 0 | — |
case-19 | fail→pass | 8,094 | 3,968 | -51% | 1 | 1 | 0% | 1,395 | 877 | -37% | 0 | 0 | — |
case-20 | fail→fail | 25,312 | 6,675 | -74% | 1 | 1 | 0% | 4,775 | 805 | -83% | 0 | 0 | — |
case-21 | pass→fail | 11,543 | 8,697 | -25% | 1 | 1 | 0% | 1,962 | 991 | -49% | 0 | 0 | — |
case-22 | pass→fail | 12,767 | 9,154 | -28% | 1 | 1 | 0% | 2,686 | 1,048 | -61% | 0 | 0 | — |
case-23 | pass→fail | 22,512 | 7,865 | -65% | 1 | 1 | 0% | 4,037 | 914 | -77% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 1 counted toward the lift figure. The other 22 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. A headline lift is not published for this run.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.