Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Mechanism-Gap Hunting Campaign — hunt the specific link where scientific progress is BLOCKED, not where literature is empty. Use this instead of gap-analysis whenever the goal is truth-seeking / AI4S research rather than finding a publishable white-space — i.e. when the user asks "where is progress actually stuck", "what mechanism is blocking this", or wants a complete blocker set rather than a ranked gap list. 4 stages over reused deep-insight SOPs, with injected per-stage directives. Exhaustiv
.claude/skills/yogsoth-ai-mechanism-gap-hunting/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 216% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 76% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 94% | 0% |
Hunt the specific link where scientific progress is blocked, not where literature is empty.
This campaign is a strategy book — CC reads, internalizes, and autonomously constructs an approach. The SKILL.md files are textbooks, not scripts. CC decides execution order, depth, and iteration based on research context.
This campaign is the deliberate inverse of gap-analysis. Where gap-analysis asks "what is missing in the literature", mechanism-gap-hunting asks "what specific mechanism is blocking progress right now". The unit of output is a blocker, not a gap. Use this skill when the user wants truth-seeking over publishable white-space discovery.
A mechanism gap is the specific link blocking a target scientific problem from moving forward.
Three admission criteria — all required:
The three layers are parameterizable: swap the layer set to reuse this campaign on problems outside the default domain.
| Stage | Reused skill | Layer | Injected directive | | --- | --- | --- | --- | | ① Blocker-symptom extraction | assumption-audit → deep-insight-assumption-surfacing | strategy → sop | Collect NOT "future work / remains to be studied" (white-space) but the authors' load-bearing concessions: we (must) assume, cannot be measured (in vitro), as a first step, the relationship … is unknown, for simplicity / we approximate. Each concession = one candidate symptom; record verbatim quote + citation (= E). | | ② Three-layer localization | dialectical-reformulation + abstraction-laddering | strategy + sop | Pin each symptom to exactly one layer and state what is stuck there (theory / identifiability / container as parameterized in Core Definition). Symptoms fitting no layer → tag "suspected out-of-scope", keep in a side list, do NOT discard. | | ③ Root-cause drilling | five-whys-drilling + current-reality-tree | sop + sop | For each localized symptom, drill "why does this block progress" with TOC sufficient-cause logic (current-reality-tree), chaining symptom→mechanism root cause. Stop only when the root cause satisfies the Blocking criterion. Output the S→W segment per blocker. | | ④ Block-verification (truth filter) | failure-mode-analysis → failure-clustering | strategy → sop | Verification standard = "does this mechanism truly impede progress", NOT "is it confirmed empty across databases". Filter the disguise: things that look like blockers but are publishable points wearing novelty packaging — if filling a "gap" only increases publishability without advancing characterization or recovery, mark pseudo-blocker, remove from the set (keep archived for audit). |
This campaign produces a complete blocker set. Every candidate symptom extracted in Stage ① is tracked through all four stages. There is no scoring, no ranking, and no elimination except Stage ④'s pseudo-blocker filter.
gap-prioritization and multi-criteria-scoring MUST NOT appear in this campaign. This is the deliberate inverse of gap-analysis, whose prioritization is central. If the user later wants a ranked subset, they route to gap-analysis after this campaign finishes.
Each confirmed blocker is locked to this block:
[layer: theory/identifiability/container] <blocker name>
| E: verbatim quote + citation
| S: symptom
| W: root-cause mechanism
| verify: how the Blocking criterion is satisfiedPseudo-blockers are archived separately with a one-line note: why they were disqualified (publishability packaging vs. genuine blocking).
| Stage | SOP (real flat-body name) | Role in this campaign | | --- | --- | --- | | ① | assumption-audit | Entry strategy: surface load-bearing concessions from the literature | | ① | deep-insight-assumption-surfacing | Sop: extract and record each concession as a candidate E with verbatim quote + citation | | ② | dialectical-reformulation | Strategy: challenge each symptom's layer attribution — stress-test the localization | | ② | abstraction-laddering | Sop: move up/down abstraction levels to pin the symptom to exactly one layer | | ③ | five-whys-drilling | Sop: iterative causal drilling — why does this symptom exist? | | ③ | current-reality-tree | Sop: TOC sufficient-cause logic — chain symptom → root cause until Blocking criterion is met | | ④ | failure-mode-analysis | Strategy: evaluate whether each root cause is a true failure mode for the field | | ④ | failure-clustering | Sop: cluster candidate blockers, detect publishability-packaging disguise, output final set |
context-init is called once at campaign start to initialize the context file for this run.
context-checkpoint is called after each stage completes. Each checkpoint captures:
Accumulated state persists across stages within a campaign run.
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | assumption-audit | Surface all assumptions, classify by vulnerability (load-bearing × likely-false), validate causal logic. Focus on dangerous assumptions — high load-bearing + non-explicit. | | dialectical-reformulation | Surface Argyris governing variables and test whether the problem dissolves under alternative governing variables (double-loop learning). | | failure-mode-analysis | Systematically catalog failure modes — generate edge cases, observe failures, cluster by mechanism, identify triggers and frequency. |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | abstraction-laddering | Move between concrete and abstract framings — 3 levels up (Why?) and 3 levels down (How?) to find the most productive research level. | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. | | current-reality-tree | Build TOC Current Reality Trees — connect Undesirable Effects via sufficient-cause logic to identify 1-3 root causes. | | deep-insight-assumption-surfacing | Systematically extract implicit assumptions from methods, frameworks, or arguments. Identifies what is taken for granted without explicit justification. | | failure-clustering | Group observed failures by mechanism (not symptom), identify common triggers per cluster, estimate frequency and severity. | | five-whys-drilling | Iterative "Why?" questioning (5+ levels) to drill from surface phenomenon to actionable root cause. Each level verified against evidence. |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | fail→pass | 21,563 | 11,554 | -46% | 1 | 1 | 0% | 2,441 | 3,846 | +58% | 0 | 0 | — |
case-01 | fail→fail | 31,946 | 47,622 | +49% | 1 | 1 | 0% | 5,370 | 8,761 | +63% | 0 | 0 | — |
case-02 | fail→fail | 13,937 | 25,179 | +81% | 1 | 1 | 0% | 1,461 | 6,154 | +321% | 0 | 0 | — |
case-03 | pass→fail | 35,261 | 39,377 | +12% | 1 | 1 | 0% | 3,250 | 8,543 | +163% | 0 | 0 | — |
case-04 | fail→fail | 13,102 | 17,999 | +37% | 1 | 1 | 0% | 1,176 | 3,969 | +238% | 0 | 0 | — |
case-05 | fail→fail | 26,469 | 30,924 | +17% | 1 | 1 | 0% | 2,971 | 5,951 | +100% | 0 | 0 | — |
case-06 | fail→pass | 18,003 | 30,049 | +67% | 1 | 1 | 0% | 1,837 | 5,796 | +216% | 0 | 0 | — |
case-07 | fail→pass | 20,857 | 14,156 | -32% | 1 | 1 | 0% | 2,619 | 4,087 | +56% | 0 | 0 | — |
case-08 | pass→pass | 21,931 | 9,953 | -55% | 1 | 1 | 0% | 2,591 | 2,649 | +2% | 0 | 0 | — |
case-09 | pass→pass | 19,218 | 12,737 | -34% | 1 | 1 | 0% | 1,806 | 3,166 | +75% | 0 | 0 | — |
case-11 | fail→pass | 17,696 | 8,196 | -54% | 1 | 1 | 0% | 1,845 | 3,242 | +76% | 0 | 0 | — |
case-12 | fail→pass | 14,202 | 12,134 | -15% | 1 | 1 | 0% | 1,571 | 3,044 | +94% | 0 | 0 | — |
case-13 | fail→pass | 25,334 | 13,653 | -46% | 1 | 1 | 0% | 3,580 | 3,620 | +1% | 0 | 0 | — |
case-14 | pass→pass | 25,846 | 27,962 | +8% | 1 | 1 | 0% | 3,207 | 5,462 | +70% | 0 | 0 | — |
case-15 | fail→pass | 36,826 | 9,426 | -74% | 1 | 1 | 0% | 1,652 | 2,660 | +61% | 0 | 0 | — |
case-16 | pass→pass | 13,482 | 18,052 | +34% | 1 | 1 | 0% | 2,108 | 3,745 | +78% | 0 | 0 | — |
case-17 | pass→pass | 20,277 | 8,351 | -59% | 1 | 1 | 0% | 2,333 | 2,508 | +8% | 0 | 0 | — |
case-18 | fail→pass | 21,600 | 8,142 | -62% | 1 | 1 | 0% | 2,436 | 2,367 | -3% | 0 | 0 | — |
case-19 | pass→pass | 10,631 | 2,359 | -78% | 1 | 1 | 0% | 858 | 2,247 | +162% | 0 | 0 | — |
case-20 | fail→pass | 37,454 | 9,248 | -75% | 1 | 1 | 0% | 1,398 | 2,554 | +83% | 0 | 0 | — |
case-21 | fail→pass | 17,588 | 14,811 | -16% | 1 | 1 | 0% | 1,972 | 2,697 | +37% | 0 | 0 | — |
case-22 | fail→fail | 11,242 | 3,064 | -73% | 1 | 1 | 0% | 1,763 | 2,285 | +30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.