Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build product A/B test briefs with hypotheses, success metrics, guardrails, baselines, proxy metrics, eligibility, variants, randomization, confidence, and launch criteria. Use when planning an A/B test from a product idea, writing an experiment spec, defining test/control variants, choosing metrics, or checking whether an experiment is ready to run.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -5% | 0% |
Use this skill to turn a product change into a decision-ready A/B test brief. It focuses on experiment anatomy: hypothesis, metrics, baselines, variants, eligibility, randomization, confidence, and launch criteria.
Primary source: Practical A/B Testing by Leemay Nassery. Guidance is transformed and paraphrased from chapter 2, especially "Creating a Clear Hypothesis" through "Summarizing the For You A/B Test" in the working text analysis at lines 1203-1942. Related motivation and variant examples come from chapter 1 lines 394-718.
experiment-sensitivity-optimization: use when the brief is blocked by MDE,sample size, noisy metrics, CUPED, capping, or too many variants.
experiment-verification-monitoring: use when the brief needs prelaunch QA,canaries, exposure validation, or active experiment health checks.
long-term-impact-evaluation: use when the brief needs delayed or sustainedimpact measurement beyond the initial test window.
| Need | Read | |------|------| | Concepts and terminology | references/core/knowledge.md | | Design rules and readiness checks | references/core/rules.md | | Brief examples and anti-examples | references/core/examples.md | | Step-by-step brief creation | workflows/create-ab-test-brief.md |
markdown# A/B Test Brief ## Decision [What decision this test will support.] ## Hypothesis Because [observation], we believe [change] will cause [outcome] for [audience]. We will know this is true when [primary metric] changes without harming [guardrails]. ## Metrics | Metric | Role | Baseline | Target or Concern | Data Source | |--------|------|----------|-------------------|-------------| ## Variants and Eligibility - Population: - Eligibility criteria: - Exposure event: - Control: - Test: - Randomization unit: ## Confidence Plan - Minimum detectable effect: - Sample size or duration: - Risks to validity: ## Launch Criteria - Ship if: - Do not ship if: - Investigate if:
Other measured skills in the registry, with their headline benchmark lift.