Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate testable hypotheses. Formulate from observations, design experiments, explore competing explanations, develop predictions, propose mechanisms, for scientific inquiry across domains.
.claude/skills/microck-hypothesis-generation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-18 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 86% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 73% | 0% |
| case-02 | ✓→✗ | ▼ Worse | -71% | 0% |
| case-05 | ✓→✗ | ▼ Worse | 71% | 0% |
Hypothesis generation is a systematic process for developing testable explanations. Formulate evidence-based hypotheses from observations, design experiments, explore competing explanations, and develop predictions. Apply this skill for scientific inquiry across domains.
This skill should be used when:
Follow this systematic process to generate robust scientific hypotheses:
Start by clarifying the observation, question, or phenomenon that requires explanation:
Search existing scientific literature to ground hypotheses in current evidence. Use both PubMed (for biomedical topics) and general web search (for broader scientific domains):
For biomedical topics:
For all scientific domains:
Search strategy:
references/literature_search_strategies.md for detailed search techniquesAnalyze and integrate findings from literature search:
Develop 3-5 distinct hypotheses that could explain the phenomenon. Each hypothesis should:
Strategies for generating hypotheses:
Assess each hypothesis against established quality criteria from references/hypothesis_quality_criteria.md:
Testability: Can the hypothesis be empirically tested? Falsifiability: What observations would disprove it? Parsimony: Is it the simplest explanation that fits the evidence? Explanatory Power: How much of the phenomenon does it explain? Scope: What range of observations does it cover? Consistency: Does it align with established principles? Novelty: Does it offer new insights beyond existing explanations?
Explicitly note the strengths and weaknesses of each hypothesis.
For each viable hypothesis, propose specific experiments or studies to test it. Consult references/experimental_design_patterns.md for common approaches:
Experimental design elements:
Consider multiple approaches:
For each hypothesis, generate specific, quantitative predictions:
Use the template in assets/hypothesis_output_template.md to present hypotheses in a clear, consistent format:
Standard structure:
Ensure all generated hypotheses meet these standards:
hypothesis_quality_criteria.md - Framework for evaluating hypothesis quality (testability, falsifiability, parsimony, explanatory power, scope, consistency)experimental_design_patterns.md - Common experimental approaches across domains (RCTs, observational studies, lab experiments, computational models)literature_search_strategies.md - Effective search techniques for PubMed and general scientific sourceshypothesis_output_template.md - Structured format for presenting hypotheses consistently with all required sections| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 38,306 | 8,305 | -78% | 1 | 1 | 0% | 6,239 | 1,737 | -72% | 0 | 0 | — |
case-02 | pass→fail | 37,929 | 8,617 | -77% | 1 | 1 | 0% | 5,940 | 1,704 | -71% | 0 | 0 | — |
case-07 | pass→pass | 29,986 | 37,676 | +26% | 1 | 1 | 0% | 4,639 | 7,448 | +61% | 0 | 0 | — |
case-03 | fail→fail | 29,602 | 44,378 | +50% | 1 | 1 | 0% | 5,072 | 1,755 | -65% | 0 | 0 | — |
case-04 | fail→fail | 4,130 | 15,632 | +278% | 1 | 1 | 0% | 560 | 4,095 | +631% | 0 | 0 | — |
case-05 | pass→fail | 25,946 | 35,874 | +38% | 1 | 1 | 0% | 4,342 | 7,413 | +71% | 0 | 0 | — |
case-06 | pass→fail | 29,015 | 47,601 | +64% | 1 | 1 | 0% | 4,764 | 7,415 | +56% | 0 | 0 | — |
case-08 | pass→pass | 28,123 | 43,001 | +53% | 1 | 1 | 0% | 4,523 | 7,444 | +65% | 0 | 0 | — |
case-09 | pass→pass | 26,557 | 38,626 | +45% | 1 | 1 | 0% | 4,406 | 7,425 | +69% | 0 | 0 | — |
case-10 | pass→pass | 23,568 | 36,392 | +54% | 1 | 1 | 0% | 3,674 | 7,258 | +98% | 0 | 0 | — |
case-11 | pass→fail | 28,042 | 10,395 | -63% | 1 | 1 | 0% | 4,149 | 1,743 | -58% | 0 | 0 | — |
case-12 | pass→pass | 27,089 | 39,409 | +45% | 1 | 1 | 0% | 4,402 | 7,431 | +69% | 0 | 0 | — |
case-13 | pass→pass | 33,493 | 34,607 | +3% | 1 | 1 | 0% | 4,213 | 7,426 | +76% | 0 | 0 | — |
case-14 | pass→pass | 30,729 | 37,661 | +23% | 1 | 1 | 0% | 4,731 | 7,417 | +57% | 0 | 0 | — |
case-15 | pass→pass | 32,307 | 46,528 | +44% | 1 | 1 | 0% | 3,970 | 7,415 | +87% | 0 | 0 | — |
case-16 | pass→pass | 25,872 | 45,729 | +77% | 1 | 1 | 0% | 4,091 | 7,422 | +81% | 0 | 0 | — |
case-17 | fail→fail | 23,984 | 16,875 | -30% | 1 | 1 | 0% | 3,564 | 2,492 | -30% | 0 | 0 | — |
case-18 | fail→pass | 29,259 | 64,250 | +120% | 1 | 1 | 0% | 4,427 | 7,421 | +68% | 0 | 0 | — |
case-19 | pass→pass | 28,944 | 40,847 | +41% | 1 | 1 | 0% | 4,395 | 7,418 | +69% | 0 | 0 | — |
case-20 | pass→pass | 30,280 | 39,291 | +30% | 1 | 1 | 0% | 4,607 | 7,252 | +57% | 0 | 0 | — |
case-21 | fail→pass | 26,301 | 37,900 | +44% | 1 | 1 | 0% | 4,003 | 7,427 | +86% | 0 | 0 | — |
case-22 | fail→pass | 41,280 | 45,264 | +10% | 1 | 1 | 0% | 4,294 | 7,421 | +73% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 17 comparable cases. 6 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.