Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Research how to implement a phase (standalone - usually use /gsd-plan-phase instead)
.claude/skills/coco-research-gsd-research-phase/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 17% | 0% |
<objective> Research how to implement a phase. Spawns gsd-phase-researcher agent with phase context.
Note: This is a standalone research command. For most workflows, use /gsd-plan-phase which integrates research automatically.
Use this command when:
Orchestrator role: Parse phase, validate against roadmap, check existing research, gather context, spawn researcher agent, present results.
Why subagent: Research burns context fast (WebSearch, Context7 queries, source verification). Fresh 200k context for investigation. Main context stays lean for user interaction. </objective>
<available_agent_types> Valid GSD subagent types (use exact names — do not fall back to 'general-purpose'):
</available_agent_types>
<context> Phase number: $ARGUMENTS (required)
Normalize phase input in step 1 before any directory lookups. </context>
<process>
bashINIT=$(node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs" init phase-op "$ARGUMENTS") if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi
Extract from init JSON: phase_dir, phase_number, phase_name, phase_found, commit_docs, has_research, state_path, requirements_path, context_path, research_path.
Resolve researcher model:
bashRESEARCHER_MODEL=$(node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs" resolve-model gsd-phase-researcher --raw)
bashPHASE_INFO=$(node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs" roadmap get-phase "${phase_number}")
If found is false: Error and exit. If found is true: Extract phase_number, phase_name, goal from JSON.
bashls .planning/phases/${PHASE}-*/RESEARCH.md 2>/dev/null
If exists: Offer: 1) Update research, 2) View existing, 3) Skip. Wait for response.
If doesn't exist: Continue.
Use paths from INIT (do not inline file contents in orchestrator context):
requirements_pathcontext_pathstate_pathPresent summary with phase description and what files the researcher will load.
Research modes: ecosystem (default), feasibility, implementation, comparison.
markdown<research_type> Phase Research — investigating HOW to implement a specific phase well. </research_type> <key_insight> The question is NOT "which library should I use?" The question is: "What do I not know that I don't know?" For this phase, discover: - What's the established architecture pattern? - What libraries form the standard stack? - What problems do people commonly hit? - What's SOTA vs what Claude's training thinks is SOTA? - What should NOT be hand-rolled? </key_insight> <objective> Research implementation approach for Phase {phase_number}: {phase_name} Mode: ecosystem </objective> <files_to_read> - {requirements_path} (Requirements) - {context_path} (Phase context from discuss-phase, if exists) - {state_path} (Prior project decisions and blockers) </files_to_read> <additional_context> **Phase description:** {phase_description} </additional_context> <downstream_consumer> Your RESEARCH.md will be loaded by `/gsd-plan-phase` which uses specific sections: - `## Standard Stack` → Plans use these libraries - `## Architecture Patterns` → Task structure follows these - `## Don't Hand-Roll` → Tasks NEVER build custom solutions for listed problems - `## Common Pitfalls` → Verification steps check for these - `## Code Examples` → Task actions reference these patterns Be prescriptive, not exploratory. "Use X" not "Consider X or Y." </downstream_consumer> <quality_gate> Before declaring complete, verify: - [ ] All domains investigated (not just some) - [ ] Negative claims verified with official docs - [ ] Multiple sources for critical claims - [ ] Confidence levels assigned honestly - [ ] Section names match what plan-phase expects </quality_gate> <output> Write to: .planning/phases/${PHASE}-{slug}/${PHASE}-RESEARCH.md </output>
Task(
prompt=filled_prompt,
subagent_type="gsd-phase-researcher",
model="{researcher_model}",
description="Research Phase {phase}"
)## RESEARCH COMPLETE: Display summary, offer: Plan phase, Dig deeper, Review full, Done.
## CHECKPOINT REACHED: Present to user, get response, spawn continuation.
## RESEARCH INCONCLUSIVE: Show what was attempted, offer: Add context, Try different mode, Manual.
markdown<objective> Continue research for Phase {phase_number}: {phase_name} </objective> <prior_state> <files_to_read> - .planning/phases/${PHASE}-{slug}/${PHASE}-RESEARCH.md (Existing research) </files_to_read> </prior_state> <checkpoint_response> **Type:** {checkpoint_type} **Response:** {user_response} </checkpoint_response>
Task(
prompt=continuation_prompt,
subagent_type="gsd-phase-researcher",
model="{researcher_model}",
description="Continue research Phase {phase}"
)</process>
<success_criteria>
</success_criteria>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 31,553 | 31,657 | +0% | 1 | 1 | 0% | 4,105 | 2,900 | -29% | 0 | 0 | — |
case-02 | fail→fail | 29,112 | 10,051 | -65% | 1 | 1 | 0% | 3,754 | 1,781 | -53% | 0 | 0 | — |
case-03 | fail→fail | 22,484 | 5,206 | -77% | 1 | 1 | 0% | 3,662 | 1,726 | -53% | 0 | 0 | — |
case-04 | fail→fail | 13,137 | 11,766 | -10% | 1 | 1 | 0% | 1,574 | 2,120 | +35% | 0 | 0 | — |
case-05 | fail→pass | 12,649 | 3,856 | -70% | 1 | 1 | 0% | 1,862 | 2,118 | +14% | 0 | 0 | — |
case-06 | fail→pass | 13,261 | 3,341 | -75% | 1 | 1 | 0% | 1,248 | 1,937 | +55% | 0 | 0 | — |
case-07 | fail→pass | 15,543 | 9,136 | -41% | 1 | 1 | 0% | 1,561 | 2,135 | +37% | 0 | 0 | — |
case-08 | fail→pass | 20,238 | 11,024 | -46% | 1 | 1 | 0% | 2,465 | 2,369 | -4% | 0 | 0 | — |
case-09 | pass→pass | 16,901 | 8,120 | -52% | 1 | 1 | 0% | 1,760 | 1,995 | +13% | 0 | 0 | — |
case-10 | fail→pass | 19,686 | 12,499 | -37% | 1 | 1 | 0% | 2,198 | 2,573 | +17% | 0 | 0 | — |
case-11 | fail→pass | 6,903 | 8,897 | +29% | 1 | 1 | 0% | 1,131 | 2,015 | +78% | 0 | 0 | — |
case-12 | fail→pass | 9,270 | 9,303 | +0% | 1 | 1 | 0% | 1,419 | 2,289 | +61% | 0 | 0 | — |
case-13 | fail→fail | 10,182 | 3,404 | -67% | 1 | 1 | 0% | 1,568 | 2,057 | +31% | 0 | 0 | — |
case-14 | pass→pass | 10,344 | 8,625 | -17% | 1 | 1 | 0% | 1,730 | 2,143 | +24% | 0 | 0 | — |
case-15 | pass→pass | 17,727 | 3,961 | -78% | 1 | 1 | 0% | 2,044 | 2,074 | +1% | 0 | 0 | — |
case-16 | pass→pass | 16,049 | 7,698 | -52% | 1 | 1 | 0% | 1,994 | 1,931 | -3% | 0 | 0 | — |
case-17 | pass→pass | 18,178 | 9,186 | -49% | 1 | 1 | 0% | 2,041 | 2,886 | +41% | 0 | 0 | — |
case-18 | fail→fail | 12,581 | 4,337 | -66% | 1 | 1 | 0% | 1,219 | 2,038 | +67% | 0 | 0 | — |
case-19 | fail→pass | 19,679 | 8,470 | -57% | 1 | 1 | 0% | 3,122 | 1,973 | -37% | 0 | 0 | — |
case-20 | fail→fail | 21,989 | 13,487 | -39% | 1 | 1 | 0% | 3,777 | 1,997 | -47% | 0 | 0 | — |
case-21 | pass→fail | 7,432 | 12,277 | +65% | 1 | 1 | 0% | 1,003 | 2,008 | +100% | 0 | 0 | — |
case-22 | fail→fail | 4,848 | 7,620 | +57% | 1 | 1 | 0% | 770 | 1,897 | +146% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 15 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.