Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Tactic: Multi-hypothesis management — generate competing hypotheses, design discriminating predictions, build a structured comparison matrix
.claude/skills/yogsoth-ai-competing-hypothesis-matrix/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-18 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 5% | 0% |
Multi-hypothesis management — systematically generate alternative explanations competing with the primary hypothesis, design key predictions that can distinguish them, and build a structured comparison matrix to avoid confirmation bias.
The most dangerous bias in scientific reasoning is holding only one hypothesis. This tactic forces CC to first systematically construct competing hypotheses before settling on the "best" hypothesis, then design predictions that can distinguish them.
The three steps cannot be reordered: first generate competing hypotheses (skipping not allowed), then design discriminating predictions (not allowed to only compare without testing), and finally build the comparison matrix (not allowed to only enumerate without quantifying). The final output is not "which hypothesis is correct" but "what experiment can distinguish them."
| SOP | Responsibility | When to call | |-----|------|---------| | competing-hypothesis-generation | Based on the primary hypothesis, generate ≥3 alternative hypotheses competing with it (different mechanisms, same or similar phenomenon prediction range) | Required in all modes, execute first | | discriminating-prediction-design | Design discriminating predictions for each pair of competing hypotheses — find an observable result for which the two hypotheses predict differently | Required in all modes, after competing-hypothesis-generation | | hypothesis-comparison-matrix | Assemble all hypotheses and discriminating predictions into a structured comparison matrix, annotating each hypothesis's expected outcome for each prediction | Required in all modes, execute last |
Simplified (S tier, 1 primary hypothesis)
Standard (M tier, 2-3 primary hypotheses)
Deep (L tier, complex hypothesis set)
Report to the calling strategy after execution:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | competing-hypothesis-generation | SOP: Generate mechanistically distinct competing hypotheses for the same phenomenon | | discriminating-prediction-design | SOP: design key predictions and observation plans that can distinguish competing hypotheses | | hypothesis-comparison-matrix | SOP: Build a multi-dimensional comparison matrix of competing hypotheses |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-17 | pass→pass | 34,164 | 33,061 | -3% | 1 | 1 | 0% | 5,435 | 6,003 | +10% | 0 | 0 | — |
case-18 | fail→pass | 32,479 | 25,947 | -20% | 1 | 1 | 0% | 4,753 | 4,780 | +1% | 0 | 0 | — |
case-19 | fail→pass | 24,603 | 21,333 | -13% | 1 | 1 | 0% | 3,549 | 4,156 | +17% | 0 | 0 | — |
case-20 | pass→fail | 15,750 | 29,766 | +89% | 1 | 1 | 0% | 2,762 | 6,079 | +120% | 0 | 0 | — |
case-21 | pass→fail | 14,434 | 30,728 | +113% | 1 | 1 | 0% | 3,395 | 7,009 | +106% | 0 | 0 | — |
case-22 | pass→pass | 14,641 | 20,782 | +42% | 1 | 1 | 0% | 2,374 | 4,102 | +73% | 0 | 0 | — |
case-01 | fail→pass | 15,267 | 17,924 | +17% | 1 | 1 | 0% | 2,553 | 3,791 | +48% | 0 | 0 | — |
case-02 | fail→pass | 23,957 | 25,613 | +7% | 1 | 1 | 0% | 3,987 | 5,178 | +30% | 0 | 0 | — |
case-03 | fail→fail | 26,005 | 25,297 | -3% | 1 | 1 | 0% | 4,124 | 5,043 | +22% | 0 | 0 | — |
case-04 | fail→fail | 27,753 | 36,547 | +32% | 1 | 1 | 0% | 4,527 | 6,973 | +54% | 0 | 0 | — |
case-05 | fail→pass | 38,023 | 33,655 | -11% | 1 | 1 | 0% | 6,190 | 6,528 | +5% | 0 | 0 | — |
case-06 | pass→pass | 20,076 | 25,317 | +26% | 1 | 1 | 0% | 3,466 | 5,132 | +48% | 0 | 0 | — |
case-07 | fail→pass | 28,654 | 27,674 | -3% | 1 | 1 | 0% | 4,376 | 5,366 | +23% | 0 | 0 | — |
case-08 | fail→pass | 24,003 | 24,704 | +3% | 1 | 1 | 0% | 3,582 | 4,590 | +28% | 0 | 0 | — |
case-09 | fail→pass | 23,549 | 28,951 | +23% | 1 | 1 | 0% | 3,236 | 5,187 | +60% | 0 | 0 | — |
case-10 | fail→pass | 25,411 | 31,565 | +24% | 1 | 1 | 0% | 3,682 | 5,469 | +49% | 0 | 0 | — |
case-11 | fail→pass | 25,283 | 30,986 | +23% | 1 | 1 | 0% | 3,969 | 5,649 | +42% | 0 | 0 | — |
case-12 | fail→pass | 25,001 | 17,653 | -29% | 1 | 1 | 0% | 3,733 | 3,535 | -5% | 0 | 0 | — |
case-13 | fail→pass | 24,406 | 20,647 | -15% | 1 | 1 | 0% | 3,473 | 4,092 | +18% | 0 | 0 | — |
case-14 | fail→pass | 29,878 | 21,791 | -27% | 1 | 1 | 0% | 4,239 | 3,936 | -7% | 0 | 0 | — |
case-15 | pass→pass | 22,853 | 34,576 | +51% | 1 | 1 | 0% | 3,316 | 5,952 | +79% | 0 | 0 | — |
case-16 | fail→pass | 23,510 | 23,142 | -2% | 1 | 1 | 0% | 3,607 | 4,602 | +28% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.