Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate structured research questions, testable hypotheses, and candidate empirical strategies from a topic, phenomenon, or dataset description. Use when user says "give me research ideas on X", "brainstorm questions about Y", "what could I study with this data?", "I'm looking for a paper idea on...", "generate hypotheses for...". One-shot generation, not multi-turn. For idea-refinement use `/interview-me`.
.claude/skills/pedrohcgs-research-ideation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 197% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 266% | 0% |
| case-08 | ✓→✗ | ▼ Worse | 100% | 0% |
| case-16 | ✓→✗ | ▼ Worse | 102% | 0% |
Generate structured research questions, testable hypotheses, and empirical strategies from a topic, phenomenon, or dataset.
Input: $ARGUMENTS — a topic (e.g., "minimum wage effects on employment"), a phenomenon (e.g., "why do firms cluster geographically?"), or a dataset description (e.g., "panel of US counties with pollution and health outcomes, 2000-2020").
$ARGUMENTS and any referenced files. Check master_supporting_docs/ for related papers. Check .claude/rules/ for domain conventions.methods-referee.md):reduced-form (DiD, IV, RD, event study, synthetic control)structural (estimation of a fully-specified model)theory+empirics (formal model + empirical test of its predictions)descriptive (measurement, data construction, pattern documentation)formal-theory (pure theory, no empirical test in this paper)survey-experiment (vignette, conjoint, list-experiment)unsure (when multiple types are plausible — the user can pick later via /interview-me)Use .claude/references/discipline-cards.md to bias the distribution by field (econ vs poli-sci default frequencies differ — e.g., poli-sci skews more toward survey-experiment and formal-theory than econ does).
quality_reports/research_ideation_[sanitized_topic].mdmarkdown# Research Ideation: [Topic] **Date:** [YYYY-MM-DD] **Input:** [Original input] ## Overview [1-2 paragraphs situating the topic and why it matters] ## Research Questions ### RQ1: [Question] (Feasibility: High/Medium/Low) **Type:** Descriptive / Correlational / Causal / Mechanism / Policy **Paper type:** reduced-form / structural / theory+empirics / descriptive / formal-theory / survey-experiment / unsure **Hypothesis:** [Testable prediction] **Identification Strategy:** - **Method:** [the identification approach you would defend in a seminar] - **Treatment:** [What varies and when] - **Control group:** [Comparison units] - **Key assumption:** [the assumption the method's validity rests on, stated so it can be attacked] **Data Requirements:** - [Dataset 1 — what it provides] - [Dataset 2 — what it provides] **Potential Pitfalls:** 1. [Threat 1 and possible mitigation] 2. [Threat 2 and possible mitigation] **Related Work:** [Author (Year)], [Author (Year)] --- [Repeat for RQ2-RQ5] ## Ranking | RQ | Feasibility | Contribution | Priority | |----|-------------|-------------|----------| | 1 | High | Medium | ... | | 2 | Medium | High | ... | ## Suggested Next Steps 1. [Most promising direction and immediate action] 2. [Data to obtain] 3. [Literature to review deeper]
Before returning the ideation report, run the Post-Flight Verification protocol from .claude/rules/post-flight-verification.md. Research ideation is hallucination-prone in three specific ways:
educ_attain" can be confidently wrong about variable names, coverage years, or restricted-access status.educ_attain variable 1990–2024?"claim-verifier via the Agent tool with subagent_type=claim-verifier and context=fork. Hand it claims + questions + source pointers (WebSearch allowed, NBER/SSRN URLs preferred, dataset codebooks preferred). Do NOT include the draft.--no-verify flag| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | pass→pass | 26,906 | 54,254 | +102% | 1 | 1 | 0% | 2,017 | 7,963 | +295% | 0 | 0 | — |
case-01 | fail→fail | 49,331 | 37,126 | -25% | 1 | 1 | 0% | 7,737 | 1,908 | -75% | 0 | 0 | — |
case-02 | fail→fail | 49,836 | 6,153 | -88% | 1 | 1 | 0% | 8,319 | 2,013 | -76% | 0 | 0 | — |
case-03 | fail→fail | 53,464 | 49,829 | -7% | 1 | 1 | 0% | 8,296 | 8,680 | +5% | 0 | 0 | — |
case-04 | fail→fail | 19,091 | 38,664 | +103% | 1 | 1 | 0% | 2,986 | 7,701 | +158% | 0 | 0 | — |
case-05 | fail→fail | 30,416 | 55,805 | +83% | 1 | 1 | 0% | 3,340 | 9,040 | +171% | 0 | 0 | — |
case-06 | fail→fail | 45,040 | 58,120 | +29% | 1 | 1 | 0% | 7,002 | 9,762 | +39% | 0 | 0 | — |
case-07 | fail→fail | 30,878 | 34,001 | +10% | 1 | 1 | 0% | 5,253 | 7,081 | +35% | 0 | 0 | — |
case-08 | pass→fail | 29,395 | 49,560 | +69% | 1 | 1 | 0% | 4,635 | 9,283 | +100% | 0 | 0 | — |
case-09 | pass→pass | 18,822 | 42,208 | +124% | 1 | 1 | 0% | 2,991 | 8,415 | +181% | 0 | 0 | — |
case-10 | fail→fail | 48,323 | 51,077 | +6% | 1 | 1 | 0% | 7,250 | 2,489 | -66% | 0 | 0 | — |
case-12 | pass→pass | 20,820 | 47,916 | +130% | 1 | 1 | 0% | 3,268 | 9,722 | +197% | 0 | 0 | — |
case-13 | pass→pass | 42,685 | 52,461 | +23% | 1 | 1 | 0% | 7,335 | 9,738 | +33% | 0 | 0 | — |
case-14 | fail→fail | 21,821 | 13,043 | -40% | 1 | 1 | 0% | 3,650 | 1,656 | -55% | 0 | 0 | — |
case-15 | fail→pass | 21,089 | 37,789 | +79% | 1 | 1 | 0% | 3,556 | 7,031 | +98% | 0 | 0 | — |
case-16 | pass→fail | 10,337 | 18,479 | +79% | 1 | 1 | 0% | 1,755 | 3,547 | +102% | 0 | 0 | — |
case-17 | pass→fail | 18,965 | 5,747 | -70% | 1 | 1 | 0% | 3,205 | 2,104 | -34% | 0 | 0 | — |
case-18 | fail→pass | 12,179 | 57,384 | +371% | 1 | 1 | 0% | 2,206 | 6,555 | +197% | 0 | 0 | — |
case-19 | fail→pass | 12,824 | 65,238 | +409% | 1 | 1 | 0% | 2,139 | 7,820 | +266% | 0 | 0 | — |
case-20 | fail→fail | 14,665 | 47,090 | +221% | 1 | 1 | 0% | 2,920 | 9,787 | +235% | 0 | 0 | — |
case-21 | fail→fail | 36,044 | 5,272 | -85% | 1 | 1 | 0% | 5,658 | 1,973 | -65% | 0 | 0 | — |
case-22 | fail→fail | 34,596 | 41,364 | +20% | 1 | 1 | 0% | 8,245 | 9,292 | +13% | 0 | 0 | — |
case-23 | fail→fail | 49,607 | 39,652 | -20% | 1 | 1 | 0% | 8,242 | 8,199 | -1% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 16 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 16 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/13/2026 | +40% |
Other measured skills in the registry, with their headline benchmark lift.