Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Campaign: Systematically assess and rank research gaps, determining the targets most worth attacking
.claude/skills/yogsoth-ai-hypothesis-formation-gap-prioritization/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -60% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -10% | 0% |
Systematically assess and rank research gaps — answering "which gaps are most worth attacking?"
<HARD-GATE> Preconditions (all must hold before starting):
Not met → stop, and inform the user that upstream work must be completed first. </HARD-GATE>
Turn a batch of unranked research gaps into a prioritized attack list. The output is not "discovering new gaps" but "making decisions about existing gaps."
| Strategy | When to use | Default | |----------|---------|------| | multi-criteria-ranking | Moderate number of gaps (5-20), systematic assessment needed | ✓ | | evidence-based-prioritization | Gaps from rigorous literature review, with ample supporting evidence | | | stakeholder-weighted-ranking | Research involves multiple stakeholders | | | portfolio-optimization | Large number of gaps (20+), portfolio-level decisions needed | | | rapid-triage | Very large number of gaps (50+), rapid coarse screening needed first | |
CC autonomously selects a strategy based on gap count, evidence sufficiency, and stakeholder complexity. Strategies can be combined (e.g., rapid-triage to screen first → multi-criteria-ranking for fine ranking).
| Tier | Number of gaps | Assessment dimensions | Scoring rounds | Final output | |------|---------|---------|---------|---------| | S | 3-5 | ≥3 | 1 | Ranking + top 2 attack suggestions | | M | 6-20 | ≥4 | ≥2 (incl. sensitivity testing) | Ranking + top 3-5 attack suggestions | | L | 20+ | ≥5 | ≥3 (incl. portfolio optimization) | Ranking + top 5-8 attack suggestions + portfolio analysis |
Each campaign execution must produce:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | evidence-based-prioritization | Strategy: evidence-strength-based AHRQ PiCMe assessment — drive gap prioritization with the quality of literature evidence | | hypothesis-formation-portfolio-optimization | Strategy: Treat the gap set as an investment portfolio — use risk/return/diversity optimization to select the optimal gap portfolio | | multi-criteria-ranking | Strategy: multi-dimensional weighted scoring and ranking — decompose a gap into independent sub-questions, then recombine into a priority list | | rapid-triage | Strategy: rapid coarse screening — two filtering rounds compress a large set of gaps into a fine-rankable candidate set | | stakeholder-weighted-ranking | Strategy: Weight by stakeholder perspective — the same gap carries different weight under different perspectives; take the consensus ranking at the end |
Optional, no fixed order; the final leaf is always a sop.
| Tactic | When to use | | --- | --- | | pairwise-comparison | Tactic: rank gaps through relative comparison rather than absolute scoring, suited to hard-to-quantify situations |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. | | hypothesis-formation-quality-gate-check | Shared SOP: General quality-gate check (format completeness, logical consistency) | | hypothesis-formation-saturation-detection | Shared SOP: judge whether the current activity has reached information saturation |
Optional, no fixed order; the final leaf is always a sop.
| Campaign | When to use | | --- | --- | | hypothesis-formulation | Campaign: transform insights and gaps into structured testable hypotheses |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 27,424 | 27,679 | +1% | 1 | 1 | 0% | 4,132 | 1,901 | -54% | 0 | 0 | — |
case-07 | fail→pass | 19,501 | 5,001 | -74% | 1 | 1 | 0% | 2,986 | 2,107 | -29% | 0 | 0 | — |
case-02 | fail→fail | 41,375 | 7,522 | -82% | 1 | 1 | 0% | 6,409 | 1,909 | -70% | 0 | 0 | — |
case-03 | fail→pass | 35,973 | 106,861 | +197% | 1 | 1 | 0% | 6,475 | 6,866 | +6% | 0 | 0 | — |
case-04 | fail→pass | 15,725 | 6,349 | -60% | 1 | 1 | 0% | 1,976 | 2,107 | +7% | 0 | 0 | — |
case-05 | fail→fail | 24,211 | 29,533 | +22% | 1 | 1 | 0% | 3,989 | 5,793 | +45% | 0 | 0 | — |
case-06 | fail→pass | 41,326 | 9,163 | -78% | 1 | 1 | 0% | 6,181 | 2,486 | -60% | 0 | 0 | — |
case-08 | fail→fail | 9,026 | 19,681 | +118% | 1 | 1 | 0% | 1,245 | 3,761 | +202% | 0 | 0 | — |
case-09 | fail→fail | 17,335 | 8,175 | -53% | 1 | 1 | 0% | 2,567 | 1,578 | -39% | 0 | 0 | — |
case-10 | fail→pass | 16,861 | 7,819 | -54% | 1 | 1 | 0% | 2,647 | 2,395 | -10% | 0 | 0 | — |
case-11 | fail→pass | 17,168 | 4,911 | -71% | 1 | 1 | 0% | 2,677 | 1,951 | -27% | 0 | 0 | — |
case-12 | pass→fail | 15,150 | 7,974 | -47% | 1 | 1 | 0% | 2,159 | 2,390 | +11% | 0 | 0 | — |
case-13 | fail→fail | 21,919 | 8,445 | -61% | 1 | 1 | 0% | 2,216 | 2,442 | +10% | 0 | 0 | — |
case-14 | fail→fail | 12,880 | 7,619 | -41% | 1 | 1 | 0% | 1,927 | 1,659 | -14% | 0 | 0 | — |
case-15 | fail→fail | 30,998 | 3,076 | -90% | 1 | 1 | 0% | 1,143 | 1,603 | +40% | 0 | 0 | — |
case-16 | pass→fail | 30,106 | 10,613 | -65% | 1 | 1 | 0% | 4,374 | 1,705 | -61% | 0 | 0 | — |
case-17 | fail→fail | 31,363 | 8,742 | -72% | 1 | 1 | 0% | 4,462 | 2,561 | -43% | 0 | 0 | — |
case-18 | fail→fail | 11,528 | 4,092 | -65% | 1 | 1 | 0% | 1,520 | 1,834 | +21% | 0 | 0 | — |
case-19 | fail→fail | 11,078 | 2,755 | -75% | 1 | 1 | 0% | 1,508 | 1,586 | +5% | 0 | 0 | — |
case-20 | pass→pass | 13,924 | 7,650 | -45% | 1 | 1 | 0% | 1,926 | 2,200 | +14% | 0 | 0 | — |
case-21 | pass→pass | 16,400 | 13,028 | -21% | 1 | 1 | 0% | 2,429 | 3,330 | +37% | 0 | 0 | — |
case-22 | fail→fail | 6,830 | 8,348 | +22% | 1 | 1 | 0% | 941 | 1,519 | +61% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 15 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.