Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Campaign: Refine hypotheses into precise, framed research questions
.claude/skills/yogsoth-ai-research-question/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 77% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -10% | 0% |
Refine hypotheses into precise research questions — answering "how do we refine a hypothesis into a precise research question?"
<HARD-GATE> Preconditions (all must hold before starting):
Not satisfied → stop, recommend completing hypothesis-formulation or clarifying the research direction first. </HARD-GATE>
Transform a "testable hypothesis" into a "precise research question" — the question must have a clear scope, measurable success criteria, and decomposable sub-questions. The output is a research question that can directly guide experiment design or literature review.
| Strategy | When to Use | Core action | |----------|---------|---------| | framework-guided-formulation | Research type is clear, with a corresponding standard framework | Select framework → fill in | | scope-calibration | Question is too broad or too narrow | zoom in/out | | decomposition-formulation | High question complexity, not answerable by a single experiment | Decompose | | comparative-formulation | A comparison of A vs B is needed | Construct the comparison | | feasibility-constrained-formulation | The ideal question exceeds available resources | Pragmatic adjustment |
The CC selects autonomously based on hypothesis characteristics and constraints. A common combination: framework-guided → scope-calibration → decomposition.
| Tier | Framework assessment | RQ output | FINER check | Sub-questions | |------|---------|---------|-----------|--------------| | S | ≥2 framework comparisons | ≥1 precise RQ | All 5 items pass | Optional | | M | ≥3 framework comparisons | ≥2 precise RQs | All 5 items pass + success criteria | ≥3 sub-questions | | L | ≥4 framework comparisons | ≥3 precise RQs | All 5 items pass + criteria + answering sequence | ≥5 sub-questions + dependency graph |
Every RQ must contain:
Each campaign execution must produce:
<!-- BEGIN available-tables (generated) -->
Optional, no fixed order; the final leaf is always a sop.
| Strategy | When to use | | --- | --- | | comparative-formulation | Strategy: Construct comparative research questions — systematic comparison of A vs B | | decomposition-formulation | Strategy: decompose a complex research question into a hierarchy of independently answerable sub-questions | | feasibility-constrained-formulation | Strategy: reshape a research question under resource constraints — pragmatic adjustment that preserves core value | | framework-guided-formulation | Strategy: Select an RQ framework (PICO/SPIDER/SPICE/ECLIPSE) and apply it systematically | | scope-calibration | Strategy: Adjust research question scope — zoom in/out until the scope is appropriate |
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | context-checkpoint | Append research process and results to the current Phase's context file. Covers both process and results with genuine substance. Use this skill at plan-designated checkpoint points — typically after each strategy completes or at key decision nodes within a research Phase. | | context-init | Create a new context file for a research Phase. Called once at Phase start to initialize the file that subsequent context-checkpoint calls will append to. Use this skill whenever a new research Phase begins and a fresh context file is needed. | | hypothesis-formation-quality-gate-check | Shared SOP: General quality-gate check (format completeness, logical consistency) | | hypothesis-formation-saturation-detection | Shared SOP: judge whether the current activity has reached information saturation | | question-synthesis | SOP: synthesize all intermediate products into a final research-question set |
Optional, no fixed order; the final leaf is always a sop.
| Campaign | When to use | | --- | --- | | hypothesis-formulation | Campaign: transform insights and gaps into structured testable hypotheses |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,741 | 18,139 | -27% | 1 | 1 | 0% | 3,371 | 1,868 | -45% | 0 | 0 | — |
case-02 | fail→pass | 133,938 | 55,465 | -59% | 1 | 1 | 0% | 3,966 | 7,038 | +77% | 0 | 0 | — |
case-03 | fail→fail | 45,137 | 20,581 | -54% | 1 | 1 | 0% | 7,851 | 1,964 | -75% | 0 | 0 | — |
case-04 | pass→pass | 19,326 | 63,589 | +229% | 1 | 1 | 0% | 2,735 | 7,482 | +174% | 0 | 0 | — |
case-05 | pass→pass | 18,007 | 25,780 | +43% | 1 | 1 | 0% | 2,314 | 5,206 | +125% | 0 | 0 | — |
case-06 | pass→pass | 13,675 | 20,567 | +50% | 1 | 1 | 0% | 2,152 | 3,583 | +66% | 0 | 0 | — |
case-07 | fail→pass | 13,680 | 9,098 | -33% | 1 | 1 | 0% | 2,071 | 2,504 | +21% | 0 | 0 | — |
case-08 | fail→pass | 18,708 | 14,482 | -23% | 1 | 1 | 0% | 2,137 | 2,683 | +26% | 0 | 0 | — |
case-09 | pass→fail | 25,689 | 63,559 | +147% | 1 | 1 | 0% | 3,555 | 1,720 | -52% | 0 | 0 | — |
case-10 | fail→pass | 54,688 | 43,828 | -20% | 1 | 1 | 0% | 8,260 | 7,852 | -5% | 0 | 0 | — |
case-11 | pass→fail | 20,679 | 17,490 | -15% | 1 | 1 | 0% | 2,383 | 1,509 | -37% | 0 | 0 | — |
case-12 | fail→fail | 10,754 | 18,422 | +71% | 1 | 1 | 0% | 1,018 | 1,579 | +55% | 0 | 0 | — |
case-13 | fail→fail | 34,190 | 19,326 | -43% | 1 | 1 | 0% | 4,910 | 1,653 | -66% | 0 | 0 | — |
case-14 | pass→pass | 18,516 | 47,813 | +158% | 1 | 1 | 0% | 3,080 | 8,372 | +172% | 0 | 0 | — |
case-15 | pass→fail | 20,490 | 20,127 | -2% | 1 | 1 | 0% | 3,183 | 1,840 | -42% | 0 | 0 | — |
case-16 | fail→fail | 9,906 | 17,726 | +79% | 1 | 1 | 0% | 1,779 | 1,745 | -2% | 0 | 0 | — |
case-17 | pass→pass | 13,289 | 23,214 | +75% | 1 | 1 | 0% | 1,261 | 5,189 | +311% | 0 | 0 | — |
case-18 | pass→fail | 14,156 | 20,491 | +45% | 1 | 1 | 0% | 1,516 | 1,794 | +18% | 0 | 0 | — |
case-19 | pass→pass | 14,227 | 35,435 | +149% | 1 | 1 | 0% | 1,615 | 6,368 | +294% | 0 | 0 | — |
case-20 | pass→fail | 13,614 | 12,178 | -11% | 1 | 1 | 0% | 1,561 | 1,555 | -0% | 0 | 0 | — |
case-21 | pass→fail | 23,582 | 22,511 | -5% | 1 | 1 | 0% | 3,067 | 2,231 | -27% | 0 | 0 | — |
case-22 | fail→pass | 20,096 | 11,435 | -43% | 1 | 1 | 0% | 2,352 | 2,127 | -10% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 11 counted toward the lift figure. The other 11 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 11 comparable cases. 9 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.