Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transform validated hypotheses into rigorous, executable experiment designs
.claude/skills/yogsoth-ai-experiment-design/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 36% | 0% |
Positioning: What experiment to run — transform a validated hypothesis into a rigorous experiment design that maximizes information yield per compute dollar.
Before entering this campaign, the following must be satisfied:
| Gate | Requirement | |------|-------------| | Hypothesis | A falsifiable hypothesis with clearly stated IV/DV exists | | Scope | Research question is bounded (not open-ended exploration) | | Resources | Preliminary compute/time budget is stated | | Prior Work | Relevant baselines and datasets have been identified |
If any gate fails, route back to hypothesis-generation or research-question refinement.
Produce a complete experiment design document that specifies:
| Signal in Hypothesis | Strategy | When to Use | |---------------------|----------|-------------| | "Factor X affects Y" | factor-level-design | Testing effects of specific variables | | "Component C contributes to performance" | ablation-design | Understanding component contributions | | "Method M outperforms baseline B" | comparison-design | Claiming superiority over existing work | | "Performance scales with resource R" | scaling-design | Understanding scaling behavior | | "Method works under condition C" | robustness-design | Testing failure boundaries |
Multiple strategies may be composed for complex hypotheses.
| Tier | GPU-hours | Max Factors | Max Runs | Strategy Constraint | |------|-----------|-------------|----------|-------------------| | Micro | < 10 | 3 | 20 | Fractional factorial or single ablation | | Small | 10-100 | 5 | 50 | Full factorial on key factors | | Medium | 100-1000 | 8 | 200 | Multi-strategy composition | | Large | > 1000 | Unlimited | Unlimited | Full design space exploration |
Every campaign invocation must produce at minimum:
研究过程经 context-management 落盘,与最终报告分属不同文件:
experiment-design,建立本 campaign 的过程 context 文件。init 幂等——同 Phase 重入返回原文件。
strategy 的过程与中间产出 append 进上一步的过程文件。
另起 experiment-design-report 文件落盘(见该 SOP)。
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | ablation-design | Design ablation studies to isolate component contributions in ML systems | | comparison-design | Design fair comparison experiments against baselines and competing methods | | experiment-execution-factor-level-design | Design factorial experiments to test how specific factors affect outcomes | | robustness-design | Design experiments to identify failure boundaries and robustness limits | | scaling-design | Design scaling experiments to characterize performance-resource relationships |
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | budget-constrained-design | Optimize experiment design under compute and time budget constraints | | reproducibility-protocol | Ensure experiment reproducibility through systematic environment and seed control | | statistical-method-selection | Select appropriate statistical methods for experiment analysis |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. | | design-synthesis | SOP: synthesize complete experiment design report | | experiment-execution-paper-overview | Import SOP: paper landscape scan (from literature-engine skill) | | experiment-execution-paper-research | Import SOP: paper full-text reading (from literature-engine skill) | | experiment-execution-paper-search | Import SOP: paper AI summary reading (from literature-engine skill) | | experiment-execution-quality-gate-check | Shared SOP: verify quality gate criteria are met before proceeding | | experiment-execution-saturation-detection | Shared SOP: detect information saturation — know when to stop searching/analyzing | | experiment-execution-web-research | Import SOP: deep full-page content analysis (from web-browsing skill) | | experiment-execution-web-search | Import SOP: quick web scan discovery (from web-browsing skill) |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 35,022 | 29,751 | -15% | 1 | 1 | 0% | 6,236 | 6,432 | +3% | 0 | 0 | — |
case-02 | fail→fail | 33,025 | 8,519 | -74% | 1 | 1 | 0% | 6,242 | 1,822 | -71% | 0 | 0 | — |
case-03 | fail→fail | 33,841 | 8,260 | -76% | 1 | 1 | 0% | 6,246 | 1,802 | -71% | 0 | 0 | — |
case-04 | fail→pass | 15,943 | 10,433 | -35% | 1 | 1 | 0% | 2,724 | 3,081 | +13% | 0 | 0 | — |
case-05 | pass→pass | 4,002 | 13,041 | +226% | 1 | 1 | 0% | 614 | 3,281 | +434% | 0 | 0 | — |
case-06 | fail→pass | 27,235 | 25,678 | -6% | 1 | 1 | 0% | 4,506 | 5,445 | +21% | 0 | 0 | — |
case-07 | fail→pass | 14,419 | 10,268 | -29% | 1 | 1 | 0% | 2,552 | 3,053 | +20% | 0 | 0 | — |
case-08 | fail→fail | 29,365 | 10,779 | -63% | 1 | 1 | 0% | 6,193 | 2,358 | -62% | 0 | 0 | — |
case-09 | fail→fail | 17,200 | 7,534 | -56% | 1 | 1 | 0% | 2,673 | 1,845 | -31% | 0 | 0 | — |
case-10 | fail→fail | 19,438 | 7,028 | -64% | 1 | 1 | 0% | 3,033 | 1,784 | -41% | 0 | 0 | — |
case-11 | fail→fail | 42,591 | 6,060 | -86% | 1 | 1 | 0% | 1,635 | 1,712 | +5% | 0 | 0 | — |
case-12 | fail→fail | 18,021 | 6,868 | -62% | 1 | 1 | 0% | 3,444 | 1,699 | -51% | 0 | 0 | — |
case-13 | pass→fail | 16,686 | 5,290 | -68% | 1 | 1 | 0% | 2,930 | 1,603 | -45% | 0 | 0 | — |
case-14 | fail→fail | 15,458 | 6,141 | -60% | 1 | 1 | 0% | 2,508 | 1,666 | -34% | 0 | 0 | — |
case-15 | pass→fail | 26,453 | 8,372 | -68% | 1 | 1 | 0% | 5,022 | 1,868 | -63% | 0 | 0 | — |
case-16 | fail→fail | 12,373 | 8,026 | -35% | 1 | 1 | 0% | 2,203 | 1,708 | -22% | 0 | 0 | — |
case-17 | fail→pass | 16,539 | 7,740 | -53% | 1 | 1 | 0% | 2,392 | 2,612 | +9% | 0 | 0 | — |
case-18 | fail→pass | 9,639 | 3,459 | -64% | 1 | 1 | 0% | 1,366 | 1,861 | +36% | 0 | 0 | — |
case-19 | pass→pass | 15,660 | 18,342 | +17% | 1 | 1 | 0% | 2,737 | 3,852 | +41% | 0 | 0 | — |
case-20 | fail→pass | 11,890 | 3,769 | -68% | 1 | 1 | 0% | 1,805 | 1,905 | +6% | 0 | 0 | — |
case-21 | pass→pass | 13,140 | 5,905 | -55% | 1 | 1 | 0% | 1,917 | 2,301 | +20% | 0 | 0 | — |
case-22 | fail→pass | 12,467 | 4,117 | -67% | 1 | 1 | 0% | 2,014 | 1,969 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 11 counted toward the lift figure. The other 11 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 11 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.