Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Choose the right product experiment type: superiority, non-inferiority, equivalence, A/B/n, or holdback-backed validation. Use when deciding what kind of A/B test to run, when the question is not simply "is variant better," when validating no degradation, proving similarity, comparing multiple variants, or selecting an experiment design for a mature product.
.claude/skills/hashgraph-online-experiment-type-selection/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-23 | ✗→✓ | ▲ Improved | 125% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -6% | 0% |
Use this skill when the experiment question determines the test type. Not every experiment should be a simple superiority test; some decisions need evidence that a change is not worse, roughly equivalent, or durable over time.
Primary source: Practical A/B Testing by Leemay Nassery. Guidance is transformed and paraphrased from chapter 3, especially lines 2013-2870. Related variant design context comes from chapter 1 lines 539-571 and chapter 2 lines 1564-1735.
experimentation-throughput-strategy: use when the choice is isolated versusoverlapping testing or when testing availability constrains the design.
adaptive-experimentation-strategy: use when fixed-horizon A/B testing may bereplaced by sequential testing, bandits, or contextual bandits.
ml-experiment-evaluation: use when the experiment is evaluating ML models,rankers, offline metrics, interleaving, or model filtering.
long-term-impact-evaluation: use when the test type question is reallyabout delayed or sustained impact measurement.
| Need | Read | |------|------| | Test type concepts | references/core/knowledge.md | | Selection rules | references/core/rules.md | | Scenario examples | references/core/examples.md | | Step-by-step selection | workflows/choose-experiment-type.md |
show practical similarity.
markdown# Experiment Type Recommendation ## Decision Question [What the team needs to learn.] ## Recommended Type [Superiority | Non-inferiority | Equivalence | A/B/n | Holdback] ## Why This Type Fits - Goal: - Metric behavior needed: - Risk tolerance: - Time horizon: ## Design Notes - Primary metric: - Guardrails: - Variants: - Population: - Follow-up analysis: ## Do Not Use [Types that would answer the wrong question and why.]
band.
preserve interpretable learning.
holdback-experiment-design for detailed long-term holdback planning.Other measured skills in the registry, with their headline benchmark lift.