Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Strategy: deduce hypotheses from existing theory
.claude/skills/yogsoth-ai-deductive-hypothesis-generation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 15% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 112% | 0% |
| case-20 | ✓→✓ | = Same ✓ | -18% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 83% | 0% |
Deduce hypotheses from existing theory: in domains with mature theory, transform theoretical propositions into specific testable predictions through an explicit reasoning chain.
Not applicable: emerging domains rich in data but lacking theory → use inductive-hypothesis-generation instead.
Theory → Mechanism → Variable Relationship → Testable Prediction
The core logic of deduction:
Every step must be traceable: each prediction traces back to a mechanism, each mechanism traces back to a theory. This is what distinguishes a deductive hypothesis from a guess.
Common pitfalls:
| Tier | Theory coverage | Mechanism extraction | Hypothesis output | Falsifiability | |------|---------|---------|---------|---------| | S | ≥2 named theories | ≥3 causal mechanisms | ≥2 structured hypotheses | 1 falsification scenario per hypothesis | | M | ≥3 named theories | ≥5 causal mechanisms | ≥3 structured hypotheses | ≥1 scenario + boundary conditions per hypothesis | | L | ≥5 named theories | ≥8 causal mechanisms | ≥5 structured hypotheses | full falsifiability audit + competing-theory comparison |
theory-identification SOP: scan the domain literature and list the named theories relevant to the gap and their core propositionsmechanism-extraction SOP (via the theory-mechanism-extraction tactic): extract causal mechanism chains from each theoryvariable-identification SOP: translate the constructs in the mechanisms into operational variablesrelationship-specification SOP: specify directional relationships among variables (including moderating/mediating structures)boundary-condition-specification SOP: identify the preconditions under which the theory applies (population, context, time range, etc.)falsifiability-check SOP (via the falsifiability-audit tactic): generate a falsification scenario for each hypothesisoperationalization SOP: provide draft measurement methods for the key variablesAfter each round, record:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | falsifiability-audit | Tactic: hypothesis quality assurance — check falsifiability, repair failing hypotheses, complete operationalization and boundary-condition specification | | theory-mechanism-extraction | Tactic: Core of the deductive path — start from theory to extract mechanisms, variables, and relationships, generating hypothesis candidates |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | pass→pass | 11,745 | 8,574 | -27% | 1 | 1 | 0% | 1,833 | 2,100 | +15% | 0 | 0 | — |
case-01 | pass→pass | 7,513 | 11,327 | +51% | 1 | 1 | 0% | 1,321 | 2,799 | +112% | 0 | 0 | — |
case-20 | pass→pass | 20,900 | 11,847 | -43% | 1 | 1 | 0% | 3,192 | 2,612 | -18% | 0 | 0 | — |
case-02 | pass→pass | 12,637 | 16,117 | +28% | 1 | 1 | 0% | 1,929 | 3,532 | +83% | 0 | 0 | — |
case-03 | pass→pass | 10,424 | 14,722 | +41% | 1 | 1 | 0% | 1,740 | 3,204 | +84% | 0 | 0 | — |
case-04 | pass→pass | 13,498 | 18,081 | +34% | 1 | 1 | 0% | 2,069 | 3,740 | +81% | 0 | 0 | — |
case-05 | pass→pass | 14,051 | 15,283 | +9% | 1 | 1 | 0% | 2,136 | 3,229 | +51% | 0 | 0 | — |
case-07 | pass→pass | 7,826 | 7,843 | +0% | 1 | 1 | 0% | 1,205 | 2,054 | +70% | 0 | 0 | — |
case-08 | fail→pass | 11,994 | 15,315 | +28% | 1 | 1 | 0% | 1,731 | 2,345 | +35% | 0 | 0 | — |
case-09 | pass→pass | 15,788 | 11,285 | -29% | 1 | 1 | 0% | 2,440 | 2,587 | +6% | 0 | 0 | — |
case-10 | fail→fail | 6,706 | 4,050 | -40% | 1 | 1 | 0% | 965 | 1,493 | +55% | 0 | 0 | — |
case-11 | pass→pass | 17,302 | 16,354 | -5% | 1 | 1 | 0% | 2,740 | 3,405 | +24% | 0 | 0 | — |
case-12 | pass→pass | 19,030 | 17,990 | -5% | 1 | 1 | 0% | 2,731 | 3,729 | +37% | 0 | 0 | — |
case-13 | pass→pass | 10,541 | 11,600 | +10% | 1 | 1 | 0% | 1,693 | 2,680 | +58% | 0 | 0 | — |
case-14 | pass→pass | 9,821 | 11,623 | +18% | 1 | 1 | 0% | 1,533 | 2,774 | +81% | 0 | 0 | — |
case-15 | pass→pass | 6,524 | 17,450 | +167% | 1 | 1 | 0% | 1,157 | 3,636 | +214% | 0 | 0 | — |
case-16 | pass→pass | 9,085 | 10,262 | +13% | 1 | 1 | 0% | 1,434 | 2,491 | +74% | 0 | 0 | — |
case-17 | fail→fail | 7,626 | 2,419 | -68% | 1 | 1 | 0% | 1,045 | 1,254 | +20% | 0 | 0 | — |
case-18 | pass→pass | 9,859 | 17,146 | +74% | 1 | 1 | 0% | 1,622 | 3,721 | +129% | 0 | 0 | — |
case-19 | pass→pass | 14,632 | 13,309 | -9% | 1 | 1 | 0% | 2,180 | 2,880 | +32% | 0 | 0 | — |
case-21 | pass→pass | 16,668 | 24,633 | +48% | 1 | 1 | 0% | 2,600 | 4,665 | +79% | 0 | 0 | — |
case-22 | fail→fail | 14,752 | 23,146 | +57% | 1 | 1 | 0% | 2,644 | 4,558 | +72% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.