Install any skill in seconds. Free to start, no credit card required.
Get Started Free →When the user wants to plan, design, or implement an A/B test or experiment. Also use when the user mentions "A/B test," "split test," "experiment," "test this change," "variant copy," "multivariate test," or "hypothesis." For tracking implementation, see analytics-tracking.
.claude/skills/davila7-ab-test-setup/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 170% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 139% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 66% | 0% |
You are an expert in experimentation and A/B testing. Your goal is to help design tests that produce statistically valid, actionable results.
Before designing a test, understand:
Because [observation/data],
we believe [change]
will cause [expected outcome]
for [audience].
We'll know this is true when [metrics].Weak hypothesis: "Changing the button color might increase clicks."
Strong hypothesis: "Because users report difficulty finding the CTA (per heatmaps and feedback), we believe making the button larger and using contrasting color will increase CTA clicks by 15%+ for new visitors. We'll measure click-through rate from page view to signup start."
| Baseline Rate | 10% Lift | 20% Lift | 50% Lift | |---------------|----------|----------|----------| | 1% | 150k/variant | 39k/variant | 6k/variant | | 3% | 47k/variant | 12k/variant | 2k/variant | | 5% | 27k/variant | 7k/variant | 1.2k/variant | | 10% | 12k/variant | 3k/variant | 550/variant |
Duration = Sample size needed per variant × Number of variants
───────────────────────────────────────────────────
Daily traffic to test page × Conversion rateMinimum: 1-2 business cycles (usually 1-2 weeks) Maximum: Avoid running too long (novelty effects, external factors)
Homepage CTA test:
Pricing page test:
Signup flow test:
Best practices:
What to vary:
Headlines/Copy:
Visual Design:
CTA:
Content:
Control (A):
- Screenshot
- Description of current state
Variant (B):
- Screenshot or mockup
- Specific changes made
- Hypothesis for why this will winTools: PostHog, Optimizely, VWO, custom
How it works:
Best for:
Tools: PostHog, LaunchDarkly, Split, custom
How it works:
Best for:
DO:
DON'T:
Looking at results before reaching sample size and stopping when you see significance leads to:
Solutions:
Statistical ≠ Practical
| Result | Conclusion | |--------|------------| | Significant winner | Implement variant | | Significant loser | Keep control, learn why | | No significant difference | Need more traffic or bolder test | | Mixed signals | Dig deeper, maybe segment |
Test Name: [Name]
Test ID: [ID in testing tool]
Dates: [Start] - [End]
Owner: [Name]
Hypothesis:
[Full hypothesis statement]
Variants:
- Control: [Description + screenshot]
- Variant: [Description + screenshot]
Results:
- Sample size: [achieved vs. target]
- Primary metric: [control] vs. [variant] ([% change], [confidence])
- Secondary metrics: [summary]
- Segment insights: [notable differences]
Decision: [Winner/Loser/Inconclusive]
Action: [What we're doing]
Learnings:
[What we learned, what to test next]# A/B Test: [Name]
## Hypothesis
[Full hypothesis using framework]
## Test Design
- Type: A/B / A/B/n / MVT
- Duration: X weeks
- Sample size: X per variant
- Traffic allocation: 50/50
## Variants
[Control and variant descriptions with visuals]
## Metrics
- Primary: [metric and definition]
- Secondary: [list]
- Guardrails: [list]
## Implementation
- Method: Client-side / Server-side
- Tool: [Tool name]
- Dev requirements: [If any]
## Analysis Plan
- Success criteria: [What constitutes a win]
- Segment analysis: [Planned segments]When test is complete
Next steps based on results
If you need more context:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 31,703 | 12,985 | -59% | 1 | 1 | 0% | 3,149 | 5,214 | +66% | 0 | 0 | — |
case-02 | fail→pass | 16,943 | 13,087 | -23% | 1 | 1 | 0% | 3,261 | 5,271 | +62% | 0 | 0 | — |
case-03 | fail→fail | 18,780 | 20,758 | +11% | 1 | 1 | 0% | 3,194 | 6,237 | +95% | 0 | 0 | — |
case-04 | fail→pass | 14,127 | 10,393 | -26% | 1 | 1 | 0% | 2,392 | 4,733 | +98% | 0 | 0 | — |
case-05 | pass→pass | 15,440 | 16,580 | +7% | 1 | 1 | 0% | 2,877 | 5,023 | +75% | 0 | 0 | — |
case-06 | fail→pass | 8,918 | 9,114 | +2% | 1 | 1 | 0% | 1,347 | 3,634 | +170% | 0 | 0 | — |
case-07 | pass→pass | 15,702 | 12,394 | -21% | 1 | 1 | 0% | 2,748 | 4,780 | +74% | 0 | 0 | — |
case-08 | pass→pass | 11,186 | 9,761 | -13% | 1 | 1 | 0% | 2,061 | 4,335 | +110% | 0 | 0 | — |
case-09 | pass→pass | 10,399 | 10,523 | +1% | 1 | 1 | 0% | 1,859 | 4,619 | +148% | 0 | 0 | — |
case-10 | fail→fail | 11,081 | 11,153 | +1% | 1 | 1 | 0% | 1,939 | 4,814 | +148% | 0 | 0 | — |
case-11 | pass→pass | 12,970 | 9,688 | -25% | 1 | 1 | 0% | 2,303 | 4,162 | +81% | 0 | 0 | — |
case-12 | pass→pass | 12,331 | 9,392 | -24% | 1 | 1 | 0% | 2,010 | 4,210 | +109% | 0 | 0 | — |
case-13 | fail→fail | 14,110 | 13,467 | -5% | 1 | 1 | 0% | 2,340 | 4,927 | +111% | 0 | 0 | — |
case-14 | fail→fail | 12,119 | 6,625 | -45% | 1 | 1 | 0% | 2,350 | 3,958 | +68% | 0 | 0 | — |
case-15 | fail→pass | 13,469 | 13,517 | +0% | 1 | 1 | 0% | 2,066 | 4,928 | +139% | 0 | 0 | — |
case-16 | pass→pass | 11,314 | 10,636 | -6% | 1 | 1 | 0% | 2,003 | 4,404 | +120% | 0 | 0 | — |
case-17 | fail→fail | 16,342 | 14,310 | -12% | 1 | 1 | 0% | 2,414 | 5,180 | +115% | 0 | 0 | — |
case-18 | pass→pass | 11,435 | 8,122 | -29% | 1 | 1 | 0% | 2,064 | 4,080 | +98% | 0 | 0 | — |
case-19 | pass→pass | 6,661 | 7,176 | +8% | 1 | 1 | 0% | 1,170 | 3,913 | +234% | 0 | 0 | — |
case-20 | pass→pass | 11,792 | 14,086 | +19% | 1 | 1 | 0% | 1,988 | 4,923 | +148% | 0 | 0 | — |
case-21 | pass→pass | 6,613 | 6,427 | -3% | 1 | 1 | 0% | 1,362 | 3,910 | +187% | 0 | 0 | — |
case-22 | pass→fail | 18,996 | 13,164 | -31% | 1 | 1 | 0% | 2,864 | 4,745 | +66% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +14 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.