Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant.
.claude/skills/phuryn-ab-test-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 96% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 80% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 104% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 47% | 0% |
Evaluate A/B test results with statistical rigor and translate findings into clear product decisions.
You are analyzing A/B test results for $ARGUMENTS.
If the user provides data files (CSV, Excel, or analytics exports), read and analyze them directly. Generate Python scripts for statistical calculations when needed.
If the user provides raw data, generate and run a Python script to calculate these.
| Outcome | Recommendation | |---|---| | Significant positive lift, no guardrail issues | Ship it — roll out to 100% | | Significant positive lift, guardrail concerns | Investigate — understand trade-offs before shipping | | Not significant, positive trend | Extend the test — need more data or larger effect | | Not significant, flat | Stop the test — no meaningful difference detected | | Significant negative lift | Don't ship — revert to control, analyze why |
## A/B Test Results: Test Name]
Hypothesis: What we expected] Duration: X days] | Sample: N control / M variant]
| Metric | Control | Variant | Lift | p-value | Significant? | |---|---|---|---|---|---| | Primary] | X% | Y% | +Z% | 0.0X | Yes/No | | Guardrail] | ... | ... | ... | ... | ... |
Recommendation: Ship / Extend / Stop / Investigate] Reasoning: Why] Next steps: What to do]
Think step by step. Save as markdown. Generate Python scripts for calculations if raw data is provided.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,086 | 17,259 | +23% | 1 | 1 | 0% | 2,820 | 3,886 | +38% | 0 | 0 | — |
case-02 | fail→fail | 21,592 | 23,355 | +8% | 1 | 1 | 0% | 3,942 | 6,354 | +61% | 0 | 0 | — |
case-03 | fail→fail | 13,963 | 16,764 | +20% | 1 | 1 | 0% | 2,886 | 3,722 | +29% | 0 | 0 | — |
case-04 | fail→pass | 13,211 | 26,486 | +100% | 1 | 1 | 0% | 2,398 | 4,703 | +96% | 0 | 0 | — |
case-05 | pass→pass | 11,529 | 15,157 | +31% | 1 | 1 | 0% | 2,033 | 2,992 | +47% | 0 | 0 | — |
case-06 | pass→pass | 8,740 | 11,036 | +26% | 1 | 1 | 0% | 1,631 | 3,090 | +89% | 0 | 0 | — |
case-07 | pass→pass | 16,141 | 12,593 | -22% | 1 | 1 | 0% | 2,208 | 3,648 | +65% | 0 | 0 | — |
case-08 | pass→pass | 13,476 | 15,496 | +15% | 1 | 1 | 0% | 2,490 | 3,936 | +58% | 0 | 0 | — |
case-09 | pass→pass | 15,859 | 16,333 | +3% | 1 | 1 | 0% | 2,859 | 4,579 | +60% | 0 | 0 | — |
case-10 | pass→pass | 11,866 | 37,428 | +215% | 1 | 1 | 0% | 2,024 | 2,869 | +42% | 0 | 0 | — |
case-11 | pass→pass | 13,915 | 14,795 | +6% | 1 | 1 | 0% | 2,889 | 4,401 | +52% | 0 | 0 | — |
case-12 | fail→fail | 11,295 | 11,035 | -2% | 1 | 1 | 0% | 2,175 | 3,204 | +47% | 0 | 0 | — |
case-13 | pass→pass | 11,732 | 12,544 | +7% | 1 | 1 | 0% | 2,159 | 3,756 | +74% | 0 | 0 | — |
case-14 | pass→pass | 11,428 | 12,278 | +7% | 1 | 1 | 0% | 2,730 | 3,978 | +46% | 0 | 0 | — |
case-15 | pass→pass | 12,456 | 15,063 | +21% | 1 | 1 | 0% | 2,402 | 4,365 | +82% | 0 | 0 | — |
case-16 | fail→pass | 11,880 | 11,041 | -7% | 1 | 1 | 0% | 1,931 | 3,477 | +80% | 0 | 0 | — |
case-17 | fail→pass | 10,786 | 11,956 | +11% | 1 | 1 | 0% | 2,292 | 3,469 | +51% | 0 | 0 | — |
case-18 | fail→pass | 11,434 | 16,110 | +41% | 1 | 1 | 0% | 1,791 | 3,655 | +104% | 0 | 0 | — |
case-19 | pass→pass | 16,053 | 19,966 | +24% | 1 | 1 | 0% | 2,155 | 3,845 | +78% | 0 | 0 | — |
case-20 | pass→pass | 12,711 | 14,319 | +13% | 1 | 1 | 0% | 2,572 | 2,956 | +15% | 0 | 0 | — |
case-21 | pass→pass | 9,130 | 9,780 | +7% | 1 | 1 | 0% | 1,624 | 2,487 | +53% | 0 | 0 | — |
case-22 | pass→pass | 11,821 | 16,932 | +43% | 1 | 1 | 0% | 2,349 | 4,435 | +89% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.