Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Strategy: induce and distill hypotheses from data/observations
.claude/skills/yogsoth-ai-inductive-hypothesis-generation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 104% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 6% | 0% |
Induce and distill hypotheses from data/observations: in domains with theoretical gaps or insufficient theory, distill regularities from empirical patterns and cautiously generalize them into testable propositions.
Not applicable: domains that already have a clear theoretical framework → use deductive-hypothesis-generation instead.
Observe patterns → Extract regularity → Generalize cautiously → Formulate testable claim
The core logic of induction:
The core risk of induction: over-generalization (jumping from a limited sample to a universal law). Each inductive hypothesis must make explicit:
| Tier | Pattern coverage | Regularity extraction | Hypothesis yield | Generalization boundary | |------|---------|---------|---------|---------| | S | ≥3 independent observation patterns | ≥2 regularities | ≥2 structured hypotheses | Each hypothesis specifies its sample source | | M | ≥5 independent observation patterns | ≥3 regularities | ≥3 structured hypotheses | Generalization boundary + falsification scenario | | L | ≥8 independent observation patterns | ≥5 regularities | ≥4 structured hypotheses | Complete generalization boundary + comparison of competing regularities |
anomaly-characterization SOP: systematically organize the patterns in existing observations/data (including frequency, conditions, exceptions)explanation-generation SOP (via the anomaly-driven-abduction tactic): generate candidate regularity explanations for each patternvariable-identification SOP: turn the constructs in the regularities into operationalizable variablesrelationship-specification SOP: specify the directional relationships between variables (including moderating conditions)falsifiability-check SOP (via the falsifiability-audit tactic): generate a falsification scenario + generalization boundary for each hypothesisRecord after each round:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | anomaly-driven-abduction | Tactic: Inductive/abductive path — describe anomalous phenomena, generate candidate explanations, rank by plausibility | | falsifiability-audit | Tactic: hypothesis quality assurance — check falsifiability, repair failing hypotheses, complete operationalization and boundary-condition specification |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | hypothesis-formation-variable-identification | SOP: identify variables and their roles within a hypothesis | | relationship-specification | SOP: specify the direction and form of relationships between variables |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 35,391 | 39,352 | +11% | 1 | 1 | 0% | 5,741 | 7,153 | +25% | 0 | 0 | — |
case-02 | fail→fail | 39,381 | 70,156 | +78% | 1 | 1 | 0% | 5,371 | 6,696 | +25% | 0 | 0 | — |
case-03 | pass→fail | 19,139 | 46,144 | +141% | 1 | 1 | 0% | 2,455 | 7,786 | +217% | 0 | 0 | — |
case-04 | pass→fail | 20,389 | 42,622 | +109% | 1 | 1 | 0% | 2,674 | 8,062 | +201% | 0 | 0 | — |
case-05 | pass→fail | 17,295 | 49,511 | +186% | 1 | 1 | 0% | 2,946 | 9,110 | +209% | 0 | 0 | — |
case-06 | pass→pass | 21,136 | 41,196 | +95% | 1 | 1 | 0% | 3,330 | 6,805 | +104% | 0 | 0 | — |
case-07 | fail→pass | 20,513 | 29,085 | +42% | 1 | 1 | 0% | 2,876 | 4,762 | +66% | 0 | 0 | — |
case-08 | fail→pass | 53,879 | 37,022 | -31% | 1 | 1 | 0% | 8,243 | 6,105 | -26% | 0 | 0 | — |
case-09 | pass→pass | 21,759 | 18,629 | -14% | 1 | 1 | 0% | 2,614 | 4,083 | +56% | 0 | 0 | — |
case-10 | pass→pass | 19,758 | 31,413 | +59% | 1 | 1 | 0% | 2,397 | 3,822 | +59% | 0 | 0 | — |
case-11 | pass→pass | 22,937 | 21,762 | -5% | 1 | 1 | 0% | 2,616 | 4,310 | +65% | 0 | 0 | — |
case-12 | fail→pass | 20,197 | 13,699 | -32% | 1 | 1 | 0% | 2,288 | 3,134 | +37% | 0 | 0 | — |
case-13 | pass→pass | 22,393 | 28,963 | +29% | 1 | 1 | 0% | 2,605 | 4,609 | +77% | 0 | 0 | — |
case-14 | pass→pass | 16,340 | 22,271 | +36% | 1 | 1 | 0% | 1,747 | 3,278 | +88% | 0 | 0 | — |
case-15 | pass→pass | 18,976 | 26,617 | +40% | 1 | 1 | 0% | 2,350 | 4,263 | +81% | 0 | 0 | — |
case-16 | fail→pass | 15,452 | 21,603 | +40% | 1 | 1 | 0% | 1,680 | 3,430 | +104% | 0 | 0 | — |
case-17 | fail→pass | 16,715 | 11,652 | -30% | 1 | 1 | 0% | 1,870 | 1,985 | +6% | 0 | 0 | — |
case-18 | fail→fail | 12,775 | 14,117 | +11% | 1 | 1 | 0% | 1,138 | 2,049 | +80% | 0 | 0 | — |
case-19 | pass→pass | 14,642 | 30,844 | +111% | 1 | 1 | 0% | 2,310 | 4,912 | +113% | 0 | 0 | — |
case-20 | fail→fail | 8,645 | 32,222 | +273% | 1 | 1 | 0% | 1,424 | 3,295 | +131% | 0 | 0 | — |
case-21 | fail→fail | 19,682 | 14,042 | -29% | 1 | 1 | 0% | 2,088 | 2,393 | +15% | 0 | 0 | — |
case-22 | pass→pass | 8,393 | 13,107 | +56% | 1 | 1 | 0% | 1,413 | 2,964 | +110% | 0 | 0 | — |
case-23 | fail→pass | 10,512 | 18,356 | +75% | 1 | 1 | 0% | 1,732 | 3,889 | +125% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +13 percentage points is the difference between those two pass rates over the 23 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.