Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Experiment design expert using pretotyping and lean validation for both new product concepts and existing product features.
.claude/skills/borghei-brainstorm-experiments/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 74% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 34% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 49% | 0% |
Design fast, low-cost experiments to validate product hypotheses before committing to full development. This skill applies Alberto Savoia's pretotyping philosophy ("Make sure you are building The Right It before you build It right") alongside lean experimentation methods for both new and existing products.
experiment_designer.py suggests 2-3 experiments per hypothesis with metric, threshold, effort, and duration.Before designing the experiment, confirm these inputs. If any is unknown or vague, ASK — do not assume:
hypothesis_text and the pass/fail threshold)Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
bashpython3 scripts/experiment_designer.py --demo # built-in sample (3 hypotheses) python3 scripts/experiment_designer.py input.json # design experiments for your hypotheses python3 scripts/experiment_designer.py input.json --format json
Each hypothesis needs hypothesis_text, target_segment, and product_type (new/existing). Document each experiment with assets/experiment_plan_template.md.
Load the reference that matches the task — keep this file lean and pull detail on demand:
experiment_designer.py usage and flags, output template, troubleshooting, success criteria, and bibliography. Read when designing or scripting an experiment.In Scope: XYZ hypothesis formulation and validation; experiment method selection for new products (landing page, pre-order, concierge, explainer video) and existing products (fake door, feature stub, A/B test, Wizard of Oz, in-app survey); automated experiment design from hypothesis keyword analysis; metric selection, success threshold definition, and effort/duration estimation.
Out of Scope: statistical power analysis or sample size calculation (use dedicated A/B test platforms); experiment infrastructure setup (feature flags, analytics instrumentation); running the actual experiment (this skill designs, not executes); long-term product strategy or roadmap decisions (execution/outcome-roadmap/).
Important Caveats: pretotyping validates demand and value, not usability or performance; in-app surveys are the weakest SITG signal — use only when behavioral experiments are impractical; the tool's keyword-to-signal matching is heuristic — override when domain knowledge dictates a better method.
| Integration | Direction | Description | |------------|-----------|-------------| | brainstorm-ideas/ | Receives from | Ideas generated become hypotheses for experiment design | | identify-assumptions/ | Receives from | "Test Now" assumptions become hypotheses for this skill | | pre-mortem/ | Feeds into | Experiment results inform pre-mortem risk assessment before full build | | execution/create-prd/ | Feeds into | Validated hypotheses become PRD assumptions with evidence | | execution/brainstorm-okrs/ | Feeds into | Experiment metrics may become OKR key results | | execution/outcome-roadmap/ | Feeds into | Experiment outcomes inform Now/Next/Later roadmap placement |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,405 | 6,594 | -57% | 1 | 1 | 0% | 2,342 | 1,760 | -25% | 0 | 0 | — |
case-02 | fail→fail | 15,288 | 15,540 | +2% | 1 | 1 | 0% | 2,389 | 3,792 | +59% | 0 | 0 | — |
case-03 | pass→pass | 17,553 | 20,721 | +18% | 1 | 1 | 0% | 2,805 | 4,871 | +74% | 0 | 0 | — |
case-04 | fail→pass | 14,316 | 13,996 | -2% | 1 | 1 | 0% | 2,164 | 3,382 | +56% | 0 | 0 | — |
case-05 | pass→pass | 12,241 | 9,007 | -26% | 1 | 1 | 0% | 1,956 | 2,630 | +34% | 0 | 0 | — |
case-06 | pass→pass | 13,111 | 10,527 | -20% | 1 | 1 | 0% | 1,935 | 2,883 | +49% | 0 | 0 | — |
case-07 | pass→pass | 13,308 | 12,892 | -3% | 1 | 1 | 0% | 2,051 | 3,288 | +60% | 0 | 0 | — |
case-08 | fail→fail | 8,372 | 4,160 | -50% | 1 | 1 | 0% | 1,319 | 1,835 | +39% | 0 | 0 | — |
case-09 | pass→pass | 5,261 | 2,041 | -61% | 1 | 1 | 0% | 789 | 1,569 | +99% | 0 | 0 | — |
case-10 | fail→fail | 11,896 | 9,291 | -22% | 1 | 1 | 0% | 1,859 | 2,693 | +45% | 0 | 0 | — |
case-11 | fail→fail | 11,521 | 11,524 | +0% | 1 | 1 | 0% | 2,418 | 3,546 | +47% | 0 | 0 | — |
case-12 | fail→fail | 27,777 | 14,266 | -49% | 1 | 1 | 0% | 1,457 | 4,053 | +178% | 0 | 0 | — |
case-13 | fail→fail | 11,878 | 10,721 | -10% | 1 | 1 | 0% | 1,771 | 2,820 | +59% | 0 | 0 | — |
case-14 | fail→pass | 10,376 | 5,255 | -49% | 1 | 1 | 0% | 1,602 | 2,147 | +34% | 0 | 0 | — |
case-15 | pass→pass | 15,204 | 13,223 | -13% | 1 | 1 | 0% | 2,128 | 3,272 | +54% | 0 | 0 | — |
case-16 | pass→pass | 15,474 | 12,152 | -21% | 1 | 1 | 0% | 2,264 | 2,808 | +24% | 0 | 0 | — |
case-17 | pass→pass | 14,925 | 10,923 | -27% | 1 | 1 | 0% | 2,179 | 2,865 | +31% | 0 | 0 | — |
case-18 | pass→pass | 11,403 | 11,337 | -1% | 1 | 1 | 0% | 1,845 | 2,901 | +57% | 0 | 0 | — |
case-19 | pass→pass | 12,808 | 10,750 | -16% | 1 | 1 | 0% | 1,999 | 2,824 | +41% | 0 | 0 | — |
case-20 | fail→fail | 5,931 | 14,752 | +149% | 1 | 1 | 0% | 806 | 3,780 | +369% | 0 | 0 | — |
case-21 | fail→fail | 13,734 | 13,247 | -4% | 1 | 1 | 0% | 2,126 | 3,393 | +60% | 0 | 0 | — |
case-22 | pass→pass | 14,209 | 8,788 | -38% | 1 | 1 | 0% | 2,400 | 2,603 | +8% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.