Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Tactic: hypothesis quality assurance — check falsifiability, repair failing hypotheses, complete operationalization and boundary-condition specification
.claude/skills/yogsoth-ai-falsifiability-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -17% | 0% |
Hypothesis quality assurance — check each hypothesis in the set for falsifiability, repair failing hypotheses, complete operational definitions, and specify boundary conditions, ensuring every hypothesis can be tested by experiment or observation.
The final gate of hypothesis formation. No matter how good the theoretical foundation, if the resulting hypotheses cannot be falsified, they are not scientific hypotheses. This tactic's job is "quality control," not "generation": the input is a set of existing hypothesis candidates, and the output is the proof document that each hypothesis has passed QC.
The three SOPs form a pipeline: falsifiability-check identifies problems → operationalization converts variables into measurable form → boundary-condition-specification delimits the range within which the hypothesis holds. No step skipping is allowed, and advancing directly when falsifiability-check fails is not allowed.
| SOP | Responsibility | When to call | |-----|------|---------| | falsifiability-check | Run a falsifiability check on each hypothesis, identify unfalsifiable hypotheses, and propose fixes | Required in all modes, executed first; if any hypothesis fails, iterate the fix and re-check | | operationalization | Provide operational definitions for each hypothesis's variables — how to measure, with what instrument, under what conditions | Required in all modes, executed after falsifiability-check passes | | boundary-condition-specification | Specify the preconditions, scope of applicability, and known limitations under which the hypothesis holds | Required in all modes, executed last |
Simplified (S tier, ≤3 hypotheses)
Standard (M tier, 4-6 hypotheses)
Deep (L tier, ≥7 hypotheses)
After execution, report to the calling strategy:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | boundary-condition-specification | SOP: Specify the boundary conditions under which a hypothesis holds | | falsifiability-check | SOP: check whether a hypothesis meets the falsifiability criterion | | operationalization | SOP: operationalize abstract concepts into measurable indicators and methods |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→fail | 16,793 | 20,055 | +19% | 1 | 1 | 0% | 2,653 | 4,132 | +56% | 0 | 0 | — |
case-01 | fail→pass | 16,465 | 25,365 | +54% | 1 | 1 | 0% | 2,811 | 4,812 | +71% | 0 | 0 | — |
case-02 | fail→pass | 28,023 | 23,382 | -17% | 1 | 1 | 0% | 4,943 | 5,144 | +4% | 0 | 0 | — |
case-03 | fail→fail | 35,029 | 36,649 | +5% | 1 | 1 | 0% | 6,227 | 7,086 | +14% | 0 | 0 | — |
case-05 | fail→fail | 4,077 | 19,432 | +377% | 1 | 1 | 0% | 781 | 4,490 | +475% | 0 | 0 | — |
case-06 | pass→pass | 16,243 | 22,182 | +37% | 1 | 1 | 0% | 2,720 | 4,407 | +62% | 0 | 0 | — |
case-07 | fail→pass | 15,656 | 20,104 | +28% | 1 | 1 | 0% | 2,684 | 4,539 | +69% | 0 | 0 | — |
case-08 | fail→pass | 11,022 | 7,586 | -31% | 1 | 1 | 0% | 1,720 | 2,189 | +27% | 0 | 0 | — |
case-09 | fail→fail | 13,269 | 11,212 | -16% | 1 | 1 | 0% | 2,107 | 2,796 | +33% | 0 | 0 | — |
case-10 | pass→pass | 16,488 | 6,580 | -60% | 1 | 1 | 0% | 2,461 | 1,960 | -20% | 0 | 0 | — |
case-11 | fail→pass | 12,129 | 4,027 | -67% | 1 | 1 | 0% | 1,924 | 1,589 | -17% | 0 | 0 | — |
case-12 | fail→pass | 16,423 | 28,619 | +74% | 1 | 1 | 0% | 2,396 | 5,576 | +133% | 0 | 0 | — |
case-13 | fail→fail | 15,735 | 3,219 | -80% | 1 | 1 | 0% | 2,273 | 1,368 | -40% | 0 | 0 | — |
case-14 | pass→pass | 12,876 | 12,401 | -4% | 1 | 1 | 0% | 1,741 | 2,173 | +25% | 0 | 0 | — |
case-15 | pass→pass | 11,976 | 17,066 | +43% | 1 | 1 | 0% | 1,920 | 3,506 | +83% | 0 | 0 | — |
case-16 | fail→pass | 7,797 | 13,972 | +79% | 1 | 1 | 0% | 1,246 | 3,207 | +157% | 0 | 0 | — |
case-17 | fail→pass | 13,060 | 4,264 | -67% | 1 | 1 | 0% | 2,006 | 1,558 | -22% | 0 | 0 | — |
case-18 | fail→fail | 8,569 | 3,321 | -61% | 1 | 1 | 0% | 1,241 | 1,454 | +17% | 0 | 0 | — |
case-19 | pass→pass | 14,476 | 4,046 | -72% | 1 | 1 | 0% | 2,275 | 1,575 | -31% | 0 | 0 | — |
case-20 | pass→pass | 11,506 | 7,744 | -33% | 1 | 1 | 0% | 1,700 | 1,993 | +17% | 0 | 0 | — |
case-21 | fail→pass | 10,536 | 4,474 | -58% | 1 | 1 | 0% | 1,608 | 1,523 | -5% | 0 | 0 | — |
case-22 | fail→pass | 7,798 | 4,205 | -46% | 1 | 1 | 0% | 1,265 | 1,530 | +21% | 0 | 0 | — |
case-23 | fail→pass | 8,656 | 5,709 | -34% | 1 | 1 | 0% | 1,408 | 1,944 | +38% | 0 | 0 | — |
case-24 | fail→pass | 11,908 | 7,356 | -38% | 1 | 1 | 0% | 1,983 | 1,997 | +1% | 0 | 0 | — |
case-25 | pass→pass | 12,648 | 5,561 | -56% | 1 | 1 | 0% | 1,902 | 1,855 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +44 percentage points is the difference between those two pass rates over the 25 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.