Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Systematic debugging with persistent state across context resets
.claude/skills/coco-research-gsd-debug/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 11% | 0% |
<objective> Debug issues using scientific method with subagent isolation.
Orchestrator role: Gather symptoms, spawn gsd-debugger agent, handle checkpoints, spawn continuations.
Why subagent: Investigation burns context fast (reading files, forming hypotheses, testing). Fresh 200k context per investigation. Main context stays lean for user interaction.
Flags:
--diagnose — Diagnose only. Find root cause without applying a fix. Returns a structured Root Cause Report. Use when you want to validate the diagnosis before committing to a fix.</objective>
<available_agent_types> Valid GSD subagent types (use exact names — do not fall back to 'general-purpose'):
</available_agent_types>
<context> User's issue: $ARGUMENTS
Parse flags from $ARGUMENTS:
--diagnose is present, set diagnose_only=true and remove the flag from the issue description.diagnose_only=false.Check for active sessions:
bashls .planning/debug/*.md 2>/dev/null | grep -v resolved | head -5
</context>
<process>
bashINIT=$(node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs" state load) if [[ "$INIT" == @file:* ]]; then INIT=$(cat "${INIT#@file:}"); fi
Extract commit_docs from init JSON. Resolve debugger model:
bashdebugger_model=$(node "$HOME/.claude/get-shit-done/bin/gsd-tools.cjs" resolve-model gsd-debugger --raw)
If active sessions exist AND no $ARGUMENTS:
If $ARGUMENTS provided OR user describes new issue:
Use AskUserQuestion for each:
After all gathered, confirm ready to investigate.
Fill prompt and spawn:
markdown<objective> Investigate issue: {slug} **Summary:** {trigger} </objective> <symptoms> expected: {expected} actual: {actual} errors: {errors} reproduction: {reproduction} timeline: {timeline} </symptoms> <mode> symptoms_prefilled: true goal: {if diagnose_only: "find_root_cause_only", else: "find_and_fix"} </mode> <debug_file> Create: .planning/debug/{slug}.md </debug_file>
Task(
prompt=filled_prompt,
subagent_type="gsd-debugger",
model="{debugger_model}",
description="Debug {slug}"
)If ## ROOT CAUSE FOUND (diagnose-only mode):
goal: find_and_fix to apply the fix (see step 5)/gsd-plan-phase --gapsIf ## DEBUG COMPLETE (find_and_fix mode):
/gsd-plan-phase --gaps if further work neededIf ## CHECKPOINT REACHED:
human-verify:If ## INVESTIGATION INCONCLUSIVE:
When user responds to checkpoint OR selects "Fix now" from diagnose-only results, spawn fresh agent:
markdown<objective> Continue debugging {slug}. Evidence is in the debug file. </objective> <prior_state> <files_to_read> - .planning/debug/{slug}.md (Debug session state) </files_to_read> </prior_state> <checkpoint_response> **Type:** {checkpoint_type} **Response:** {user_response} </checkpoint_response> <mode> goal: find_and_fix </mode>
Task(
prompt=continuation_prompt,
subagent_type="gsd-debugger",
model="{debugger_model}",
description="Continue debug {slug}"
)</process>
<success_criteria>
</success_criteria>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 13,599 | 11,846 | -13% | 1 | 1 | 0% | 233 | 1,582 | +579% | 0 | 0 | — |
case-01 | fail→fail | 26,135 | 21,779 | -17% | 1 | 1 | 0% | 2,977 | 1,951 | -34% | 0 | 0 | — |
case-02 | fail→fail | 7,989 | 16,262 | +104% | 1 | 1 | 0% | 1,216 | 1,694 | +39% | 0 | 0 | — |
case-04 | fail→fail | 2,556 | 12,210 | +378% | 1 | 1 | 0% | 360 | 1,762 | +389% | 0 | 0 | — |
case-05 | fail→fail | 17,578 | 18,863 | +7% | 1 | 1 | 0% | 2,575 | 1,739 | -32% | 0 | 0 | — |
case-06 | pass→fail | 27,359 | 17,854 | -35% | 1 | 1 | 0% | 3,370 | 1,577 | -53% | 0 | 0 | — |
case-07 | fail→fail | 12,587 | 7,815 | -38% | 1 | 1 | 0% | 1,273 | 1,637 | +29% | 0 | 0 | — |
case-17 | fail→pass | 13,364 | 11,940 | -11% | 1 | 1 | 0% | 1,844 | 2,305 | +25% | 0 | 0 | — |
case-08 | fail→pass | 15,067 | 7,638 | -49% | 1 | 1 | 0% | 1,541 | 1,751 | +14% | 0 | 0 | — |
case-09 | fail→pass | 22,548 | 10,005 | -56% | 1 | 1 | 0% | 2,833 | 2,062 | -27% | 0 | 0 | — |
case-10 | fail→pass | 17,227 | 7,265 | -58% | 1 | 1 | 0% | 1,970 | 1,580 | -20% | 0 | 0 | — |
case-11 | fail→pass | 13,683 | 10,227 | -25% | 1 | 1 | 0% | 2,164 | 2,403 | +11% | 0 | 0 | — |
case-12 | fail→pass | 15,698 | 7,328 | -53% | 1 | 1 | 0% | 1,518 | 1,657 | +9% | 0 | 0 | — |
case-13 | fail→pass | 13,894 | 8,956 | -36% | 1 | 1 | 0% | 2,091 | 1,891 | -10% | 0 | 0 | — |
case-14 | fail→pass | 15,262 | 3,452 | -77% | 1 | 1 | 0% | 1,578 | 1,838 | +16% | 0 | 0 | — |
case-15 | pass→pass | 9,589 | 4,815 | -50% | 1 | 1 | 0% | 676 | 2,217 | +228% | 0 | 0 | — |
case-16 | fail→pass | 16,530 | 10,578 | -36% | 1 | 1 | 0% | 1,824 | 2,301 | +26% | 0 | 0 | — |
case-18 | fail→pass | 20,803 | 5,523 | -73% | 1 | 1 | 0% | 2,145 | 2,291 | +7% | 0 | 0 | — |
case-19 | fail→pass | 20,371 | 10,440 | -49% | 1 | 1 | 0% | 2,311 | 2,329 | +1% | 0 | 0 | — |
case-20 | fail→pass | 12,320 | 8,031 | -35% | 1 | 1 | 0% | 1,068 | 1,851 | +73% | 0 | 0 | — |
case-21 | fail→pass | 8,689 | 4,362 | -50% | 1 | 1 | 0% | 1,469 | 1,991 | +36% | 0 | 0 | — |
case-22 | pass→pass | 16,107 | 10,463 | -35% | 1 | 1 | 0% | 2,139 | 2,738 | +28% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 16 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 16 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.