Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Plan A/B testing platform strategy, architecture, and build-vs-buy decisions for product engineering teams. Use when deciding whether to build or buy an experimentation platform, scoping feature flagging, targeting, assignment, exposure logging, metrics pipelines, dashboards, governance, or evolving a simple testing setup into a durable platform.
.claude/skills/hashgraph-online-ab-testing-platform-strategy/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-01 | ✓→✓ | = Same ✓ | -23% | 0% |
| case-02 | ✓→✓ | = Same ✓ | -35% | 0% |
| case-03 | ✓→✓ | = Same ✓ | -28% | 0% |
Use this skill to decide how an organization should support A/B testing through platform choices, architecture, ownership, and incremental scope.
Primary source: Practical A/B Testing by Leemay Nassery. Guidance is transformed and paraphrased from chapter 5 lines 3804-4443. Related startup and "start simple" context comes from preface lines 286-332 and chapter 2 lines 1622-1628.
experimentation-strategy-roadmap: use when deciding which platformcapability to prioritize across rate, quality, cost, usability, and company strategy.
experimentation-throughput-strategy: use when the platform needs capacityvisibility, isolated versus overlapping test policies, or coordination tools.
experiment-verification-monitoring: use when the platform needs QA tooling,canaries, A/A tests, active monitoring, leakage checks, or quality metrics.
adaptive-experimentation-strategy: use when considering sequential testing,bandits, Thompson sampling, contextual bandits, or dynamic allocation support.
| Need | Read | |------|------| | Platform concepts and components | references/core/knowledge.md | | Build-vs-buy and scoping rules | references/core/rules.md | | Scenario examples | references/core/examples.md | | Decision workflow | workflows/decide-platform-strategy.md |
markdown# A/B Testing Platform Strategy ## Recommendation [Build | Buy | Hybrid | Start manually] because [reason]. ## Current Context - Team: - Product surface: - Experiment volume: - Data maturity: - Engineering capacity: ## Required Capabilities | Capability | Need Now? | Build/Buy/Manual | Owner | |------------|-----------|------------------|-------| ## Tradeoffs - Build advantages: - Build risks: - Buy advantages: - Buy risks: - Hybrid notes: ## Incremental Roadmap 1. Minimum viable experimentation: 2. Reliability and governance: 3. Scale and self-service:
fit.
measurement creates false confidence.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 26,969 | 15,438 | -43% | 1 | 1 | 0% | 4,469 | 3,450 | -23% | 0 | 0 | — |
case-02 | pass→pass | 29,753 | 15,146 | -49% | 1 | 1 | 0% | 4,741 | 3,085 | -35% | 0 | 0 | — |
case-03 | pass→pass | 31,800 | 17,830 | -44% | 1 | 1 | 0% | 4,870 | 3,501 | -28% | 0 | 0 | — |
case-04 | pass→pass | 21,862 | 23,984 | +10% | 1 | 1 | 0% | 4,627 | 5,431 | +17% | 0 | 0 | — |
case-05 | pass→pass | 19,482 | 23,423 | +20% | 1 | 1 | 0% | 3,377 | 5,114 | +51% | 0 | 0 | — |
case-06 | pass→pass | 16,409 | 22,734 | +39% | 1 | 1 | 0% | 3,734 | 5,883 | +58% | 0 | 0 | — |
case-07 | pass→pass | 25,642 | 12,329 | -52% | 1 | 1 | 0% | 3,774 | 2,781 | -26% | 0 | 0 | — |
case-08 | pass→pass | 30,612 | 12,305 | -60% | 1 | 1 | 0% | 5,541 | 2,776 | -50% | 0 | 0 | — |
case-09 | pass→pass | 30,362 | 18,034 | -41% | 1 | 1 | 0% | 5,100 | 3,782 | -26% | 0 | 0 | — |
case-10 | pass→pass | 23,538 | 14,075 | -40% | 1 | 1 | 0% | 4,063 | 3,139 | -23% | 0 | 0 | — |
case-11 | pass→pass | 26,236 | 15,171 | -42% | 1 | 1 | 0% | 4,226 | 3,218 | -24% | 0 | 0 | — |
case-12 | pass→pass | 26,742 | 13,131 | -51% | 1 | 1 | 0% | 4,159 | 2,723 | -35% | 0 | 0 | — |
case-13 | pass→pass | 34,432 | 18,653 | -46% | 1 | 1 | 0% | 5,660 | 3,489 | -38% | 0 | 0 | — |
case-14 | pass→pass | 38,973 | 15,826 | -59% | 1 | 1 | 0% | 6,178 | 3,167 | -49% | 0 | 0 | — |
case-15 | pass→pass | 22,368 | 14,070 | -37% | 1 | 1 | 0% | 3,271 | 2,746 | -16% | 0 | 0 | — |
case-16 | fail→pass | 22,669 | 14,583 | -36% | 1 | 1 | 0% | 3,215 | 2,637 | -18% | 0 | 0 | — |
case-17 | pass→pass | 19,595 | 13,263 | -32% | 1 | 1 | 0% | 2,881 | 2,624 | -9% | 0 | 0 | — |
case-18 | pass→pass | 26,896 | 14,843 | -45% | 1 | 1 | 0% | 4,141 | 2,901 | -30% | 0 | 0 | — |
case-19 | pass→pass | 19,002 | 12,143 | -36% | 1 | 1 | 0% | 3,171 | 2,734 | -14% | 0 | 0 | — |
case-20 | pass→pass | 24,253 | 15,144 | -38% | 1 | 1 | 0% | 3,497 | 2,939 | -16% | 0 | 0 | — |
case-21 | pass→pass | 22,630 | 19,767 | -13% | 1 | 1 | 0% | 3,419 | 3,680 | +8% | 0 | 0 | — |
case-22 | fail→pass | 25,418 | 13,958 | -45% | 1 | 1 | 0% | 3,688 | 2,713 | -26% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.