Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Write a structured research prompt document — questions only, no answers
.claude/skills/sterlingcrispin-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 113% | 0% |
Write a research prompt document — a structured set of questions designed to guide deep technical investigation of a domain. The output is a markdown file that can be fed to an AI (or used by the agent itself) for thorough research.
You are writing QUESTIONS, not answers. Do not research the topic yourself. Do not include findings, recommendations, or pre-baked conclusions. The entire point is to produce a document that drives research — not to do the research.
Ask the user what they want to research if it's not clear. Understand:
Look for an existing docs/research_prompts/ folder in the project. If it exists, read 1-2 existing prompts to match the project's established format and voice. If no folder exists, create docs/research_prompts/.
Follow the format below exactly. Write it to docs/research_prompts/<topic-slug>.md.
Every research prompt follows this structure:
markdown# Research Prompt: <Clear Descriptive Title> ## Objective 1-2 paragraphs. What are we researching and why. State the core question in bold. Frame it as exploration, not as a spec. Don't prescribe the answer. ## Context ### Current State / Pain Point What exists today. What's broken or insufficient. Be specific — reference actual code, tools, numbers. This grounds the research in reality. ### What We Want Bullet list of desired capabilities. What the end state looks like. Keep it high-level — the research should figure out HOW, not be told how. ## Key Questions ### 0. <First Major Topic Area> 0a. **<Specific sub-question>** - Concrete question that demands investigation - Follow-up angle or related concern - "How does X handle this?" type comparative question 0b. **<Next sub-question>** - ... ### 1. <Second Major Topic Area> 1a. **<Sub-question>** - ... ### 2. <Third Major Topic Area> ... (Continue with as many sections as the topic demands. Typical range: 4-8 sections.) ## Desired Output What the research should produce. NOT the answers — but the SHAPE of the answers. Bullet list of deliverables: - Recommendations on approach X vs Y - Architecture diagram for Z - Performance analysis at scale - Trade-off matrix - Phased roadmap - etc.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 13,109 | 25,333 | +93% | 1 | 1 | 0% | 2,091 | 4,774 | +128% | 0 | 0 | — |
case-02 | fail→pass | 14,986 | 21,623 | +44% | 1 | 1 | 0% | 2,282 | 4,442 | +95% | 0 | 0 | — |
case-03 | fail→pass | 24,116 | 25,082 | +4% | 1 | 1 | 0% | 3,495 | 4,889 | +40% | 0 | 0 | — |
case-04 | pass→fail | 24,570 | 27,717 | +13% | 1 | 1 | 0% | 4,256 | 6,182 | +45% | 0 | 0 | — |
case-05 | pass→fail | 22,122 | 21,314 | -4% | 1 | 1 | 0% | 3,492 | 4,289 | +23% | 0 | 0 | — |
case-06 | pass→fail | 19,226 | 21,839 | +14% | 1 | 1 | 0% | 2,878 | 4,513 | +57% | 0 | 0 | — |
case-07 | fail→pass | 17,953 | 18,821 | +5% | 1 | 1 | 0% | 2,700 | 4,347 | +61% | 0 | 0 | — |
case-08 | pass→pass | 17,082 | 17,976 | +5% | 1 | 1 | 0% | 2,470 | 3,957 | +60% | 0 | 0 | — |
case-09 | fail→pass | 22,726 | 21,408 | -6% | 1 | 1 | 0% | 3,372 | 4,273 | +27% | 0 | 0 | — |
case-10 | fail→pass | 13,698 | 21,316 | +56% | 1 | 1 | 0% | 2,088 | 4,449 | +113% | 0 | 0 | — |
case-11 | pass→pass | 15,954 | 16,385 | +3% | 1 | 1 | 0% | 2,408 | 3,784 | +57% | 0 | 0 | — |
case-12 | fail→pass | 23,075 | 19,954 | -14% | 1 | 1 | 0% | 3,517 | 4,433 | +26% | 0 | 0 | — |
case-13 | fail→pass | 14,629 | 19,029 | +30% | 1 | 1 | 0% | 2,286 | 3,939 | +72% | 0 | 0 | — |
case-14 | pass→pass | 17,229 | 17,578 | +2% | 1 | 1 | 0% | 2,527 | 3,790 | +50% | 0 | 0 | — |
case-15 | pass→pass | 20,335 | 18,148 | -11% | 1 | 1 | 0% | 2,963 | 3,989 | +35% | 0 | 0 | — |
case-16 | fail→pass | 14,193 | 20,056 | +41% | 1 | 1 | 0% | 2,136 | 4,224 | +98% | 0 | 0 | — |
case-17 | pass→pass | 14,749 | 18,338 | +24% | 1 | 1 | 0% | 2,216 | 3,831 | +73% | 0 | 0 | — |
case-18 | fail→fail | 18,949 | 20,745 | +9% | 1 | 1 | 0% | 3,012 | 4,156 | +38% | 0 | 0 | — |
case-19 | pass→pass | 18,036 | 19,110 | +6% | 1 | 1 | 0% | 2,944 | 3,889 | +32% | 0 | 0 | — |
case-20 | pass→pass | 15,974 | 20,252 | +27% | 1 | 1 | 0% | 2,471 | 4,174 | +69% | 0 | 0 | — |
case-21 | pass→pass | 18,609 | 26,470 | +42% | 1 | 1 | 0% | 3,113 | 5,279 | +70% | 0 | 0 | — |
case-22 | fail→fail | 9,460 | 17,907 | +89% | 1 | 1 | 0% | 1,323 | 4,020 | +204% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.