Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Structured hypothesis formulation: turn observations into testable hypotheses with predictions, propose mechanisms, design experiments. Follows the scientific method. Use scientific-brainstorming for open ideation; hypogenic for automated LLM hypothesis testing on datasets.
.claude/skills/jaechang-hits-hypothesis-generation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 152% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 118% | 0% |
Hypothesis generation is a systematic process for developing testable mechanistic explanations from observations. This knowhow covers the full cycle: from understanding a phenomenon through literature synthesis, generating competing hypotheses, evaluating hypothesis quality, designing experimental tests, and formulating testable predictions.
Good hypotheses are mechanistic (explain HOW/WHY), not descriptive (restate WHAT).
| Criterion | Definition | Example of Strong | Example of Weak | |-----------|-----------|-------------------|-----------------| | Testability | Can be empirically investigated | "Protein X binds to receptor Y" (can test with co-IP) | "Life force drives cellular growth" (untestable) | | Falsifiability | Specific observations would disprove it | "If X is absent, effect disappears" | "X contributes to the effect somehow" | | Parsimony | Simplest explanation fitting the evidence | Single mechanism | Multi-step chain without evidence | | Explanatory Power | Accounts for observed patterns | Explains dose-response and tissue specificity | Explains only one observation | | Scope | Range of phenomena covered | Applies across related systems | Limited to single dataset | | Consistency | Aligns with established knowledge | Consistent with known pathway biology | Contradicts thermodynamics | | Novelty | Offers new insight | Proposes unexplored mechanism | Restates established knowledge |
Hypotheses can operate at different scales. Strong hypothesis sets include explanations at multiple levels:
What is your starting point?
├── Specific observation / data → Follow the full 8-step Workflow below
├── Broad research question → Start with Step 2 (literature search) to narrow scope
├── Existing hypothesis to refine → Start at Step 5 (evaluate quality) and iterate
└── Need creative ideation first → Use scientific-brainstorming skill, then return here| Starting Situation | Approach | Key Steps | |-------------------|----------|-----------| | Unexpected experimental result | Phenomenon-driven | Steps 1→2→3→4 (focus on competing explanations) | | Literature gap identified | Gap-driven | Steps 2→3→4→5 (focus on novelty criterion) | | Cross-domain analogy noticed | Analogy-driven | Steps 1→4→5→6 (focus on translating mechanism) | | Contradictory findings in literature | Conflict-driven | Steps 2→3→4→7 (focus on discriminating predictions) | | Large dataset patterns | Data-driven | Use hypogenic first, then Steps 5→6→7 here |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | fail→fail | 14,494 | 40,604 | +180% | 1 | 1 | 0% | 2,033 | 8,256 | +306% | 0 | 0 | — |
case-06 | fail→pass | 18,453 | 33,131 | +80% | 1 | 1 | 0% | 2,733 | 6,885 | +152% | 0 | 0 | — |
case-01 | fail→fail | 34,424 | 47,074 | +37% | 1 | 1 | 0% | 5,048 | 7,791 | +54% | 0 | 0 | — |
case-02 | fail→pass | 20,649 | 21,780 | +5% | 1 | 1 | 0% | 2,965 | 5,186 | +75% | 0 | 0 | — |
case-03 | pass→pass | 11,157 | 17,784 | +59% | 1 | 1 | 0% | 1,774 | 4,856 | +174% | 0 | 0 | — |
case-04 | pass→pass | 17,216 | 18,953 | +10% | 1 | 1 | 0% | 2,721 | 4,843 | +78% | 0 | 0 | — |
case-05 | pass→pass | 8,142 | 9,988 | +23% | 1 | 1 | 0% | 1,394 | 3,594 | +158% | 0 | 0 | — |
case-07 | pass→pass | 18,417 | 16,710 | -9% | 1 | 1 | 0% | 2,535 | 4,502 | +78% | 0 | 0 | — |
case-08 | pass→pass | 13,542 | 26,777 | +98% | 1 | 1 | 0% | 2,229 | 6,517 | +192% | 0 | 0 | — |
case-09 | fail→pass | 14,973 | 16,420 | +10% | 1 | 1 | 0% | 2,276 | 4,799 | +111% | 0 | 0 | — |
case-10 | fail→fail | 17,577 | 14,302 | -19% | 1 | 1 | 0% | 2,693 | 4,184 | +55% | 0 | 0 | — |
case-11 | pass→pass | 17,330 | 11,768 | -32% | 1 | 1 | 0% | 2,334 | 3,788 | +62% | 0 | 0 | — |
case-12 | fail→pass | 17,805 | 20,410 | +15% | 1 | 1 | 0% | 3,131 | 5,473 | +75% | 0 | 0 | — |
case-14 | pass→fail | 12,465 | 15,459 | +24% | 1 | 1 | 0% | 2,017 | 4,682 | +132% | 0 | 0 | — |
case-15 | fail→fail | 13,199 | 24,928 | +89% | 1 | 1 | 0% | 2,055 | 6,181 | +201% | 0 | 0 | — |
case-16 | pass→pass | 14,274 | 16,825 | +18% | 1 | 1 | 0% | 1,994 | 4,486 | +125% | 0 | 0 | — |
case-17 | fail→pass | 15,569 | 21,761 | +40% | 1 | 1 | 0% | 2,445 | 5,328 | +118% | 0 | 0 | — |
case-18 | pass→pass | 14,162 | 19,769 | +40% | 1 | 1 | 0% | 2,200 | 5,338 | +143% | 0 | 0 | — |
case-19 | fail→pass | 17,399 | 26,289 | +51% | 1 | 1 | 0% | 2,562 | 6,062 | +137% | 0 | 0 | — |
case-20 | pass→pass | 10,698 | 11,691 | +9% | 1 | 1 | 0% | 2,149 | 4,360 | +103% | 0 | 0 | — |
case-21 | pass→pass | 11,767 | 11,477 | -2% | 1 | 1 | 0% | 2,045 | 3,896 | +91% | 0 | 0 | — |
case-22 | pass→pass | 10,621 | 7,496 | -29% | 1 | 1 | 0% | 1,681 | 3,218 | +91% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.