Install any skill in seconds. Free to start, no credit card required.
Get Started Free →A/Bテストや実験の設計・実装を支援するスキル。 「A/Bテストを設計して」「スプリットテストしたい」「仮説を立ててテストしたい」「バリアントを比較」等のリクエストで発動。 トラッキング実装は analytics-tracking を参照。
.claude/skills/minicoohei-ab-test-setup/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 83% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 27% | 0% |
You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results.
Check for product marketing context first: If .claude/product-marketing-context.md exists, read it before asking questions. Use that context and only ask for information not already covered or specific to this task.
Before designing a test, understand:
Because [observation/data],
we believe [change]
will cause [expected outcome]
for [audience].
We'll know this is true when [metrics].Weak: "Changing the button color might increase clicks."
Strong: "Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. We'll measure click-through rate from page view to signup start."
| Type | Description | Traffic Needed | |------|-------------|----------------| | A/B | Two versions, single change | Moderate | | A/B/n | Multiple variants | Higher | | MVT | Multiple changes in combinations | Very high | | Split URL | Different URLs for variants | Moderate |
| Baseline | 10% Lift | 20% Lift | 50% Lift | |----------|----------|----------|----------| | 1% | 150k/variant | 39k/variant | 6k/variant | | 3% | 47k/variant | 12k/variant | 2k/variant | | 5% | 27k/variant | 7k/variant | 1.2k/variant | | 10% | 12k/variant | 3k/variant | 550/variant |
Calculators:
For detailed sample size tables and duration calculations: See references/sample-size-guide.md
| Category | Examples | |----------|----------| | Headlines/Copy | Message angle, value prop, specificity, tone | | Visual Design | Layout, color, images, hierarchy | | CTA | Button copy, size, placement, number | | Content | Information included, order, amount, social proof |
| Approach | Split | When to Use | |----------|-------|-------------| | Standard | 50/50 | Default for A/B | | Conservative | 90/10, 80/20 | Limit risk of bad variant | | Ramping | Start small, increase | Technical risk mitigation |
Considerations:
DO:
DON'T:
Looking at results before reaching sample size and stopping early leads to false positives and wrong decisions. Pre-commit to sample size and trust the process.
| Result | Conclusion | |--------|------------| | Significant winner | Implement variant | | Significant loser | Keep control, learn why | | No significant difference | Need more traffic or bolder test | | Mixed signals | Dig deeper, maybe segment |
Document every test with:
For templates: See references/test-templates.md
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 18,893 | 16,973 | -10% | 1 | 1 | 0% | 2,930 | 4,034 | +38% | 0 | 0 | — |
case-12 | pass→pass | 16,881 | 12,117 | -28% | 1 | 1 | 0% | 2,600 | 3,518 | +35% | 0 | 0 | — |
case-13 | pass→pass | 17,565 | 11,370 | -35% | 1 | 1 | 0% | 2,555 | 3,507 | +37% | 0 | 0 | — |
case-20 | pass→pass | 11,477 | 13,410 | +17% | 1 | 1 | 0% | 1,677 | 3,720 | +122% | 0 | 0 | — |
case-02 | fail→pass | 19,245 | 14,279 | -26% | 1 | 1 | 0% | 2,933 | 3,948 | +35% | 0 | 0 | — |
case-03 | fail→pass | 20,503 | 18,709 | -9% | 1 | 1 | 0% | 3,224 | 4,610 | +43% | 0 | 0 | — |
case-04 | pass→fail | 11,285 | 11,676 | +3% | 1 | 1 | 0% | 1,627 | 3,335 | +105% | 0 | 0 | — |
case-05 | pass→pass | 6,370 | 6,016 | -6% | 1 | 1 | 0% | 1,279 | 2,775 | +117% | 0 | 0 | — |
case-06 | pass→pass | 21,793 | 18,883 | -13% | 1 | 1 | 0% | 3,736 | 5,068 | +36% | 0 | 0 | — |
case-07 | fail→pass | 9,427 | 6,246 | -34% | 1 | 1 | 0% | 1,444 | 2,640 | +83% | 0 | 0 | — |
case-08 | fail→pass | 8,753 | 6,162 | -30% | 1 | 1 | 0% | 1,613 | 2,780 | +72% | 0 | 0 | — |
case-09 | pass→pass | 11,550 | 10,536 | -9% | 1 | 1 | 0% | 2,080 | 3,411 | +64% | 0 | 0 | — |
case-10 | pass→pass | 16,179 | 14,959 | -8% | 1 | 1 | 0% | 2,417 | 4,000 | +65% | 0 | 0 | — |
case-11 | pass→pass | 12,954 | 6,802 | -47% | 1 | 1 | 0% | 2,197 | 2,698 | +23% | 0 | 0 | — |
case-14 | pass→pass | 11,204 | 8,945 | -20% | 1 | 1 | 0% | 1,749 | 2,879 | +65% | 0 | 0 | — |
case-15 | pass→pass | 13,605 | 12,163 | -11% | 1 | 1 | 0% | 1,984 | 3,422 | +72% | 0 | 0 | — |
case-16 | pass→pass | 22,341 | 12,028 | -46% | 1 | 1 | 0% | 2,031 | 3,399 | +67% | 0 | 0 | — |
case-17 | pass→pass | 6,531 | 7,504 | +15% | 1 | 1 | 0% | 1,094 | 2,782 | +154% | 0 | 0 | — |
case-18 | fail→pass | 9,652 | 2,332 | -76% | 1 | 1 | 0% | 1,534 | 1,951 | +27% | 0 | 0 | — |
case-19 | fail→fail | 10,962 | 11,025 | +1% | 1 | 1 | 0% | 1,707 | 3,300 | +93% | 0 | 0 | — |
case-21 | fail→pass | 10,761 | 7,536 | -30% | 1 | 1 | 0% | 1,976 | 2,953 | +49% | 0 | 0 | — |
case-22 | pass→pass | 13,482 | 10,846 | -20% | 1 | 1 | 0% | 2,012 | 3,269 | +62% | 0 | 0 | — |
case-23 | pass→pass | 11,869 | 10,548 | -11% | 1 | 1 | 0% | 1,826 | 3,307 | +81% | 0 | 0 | — |
case-24 | pass→pass | 13,508 | 11,064 | -18% | 1 | 1 | 0% | 2,039 | 3,453 | +69% | 0 | 0 | — |
case-25 | pass→pass | 11,011 | 8,461 | -23% | 1 | 1 | 0% | 1,574 | 2,986 | +90% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +20 percentage points is the difference between those two pass rates over the 25 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.