Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute all plans in a phase with wave-based parallelization
.claude/skills/coco-research-gsd-execute-phase/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | -38% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 3% | 0% |
<objective> Execute all plans in a phase using wave-based parallel execution.
Orchestrator stays lean: discover plans, analyze dependencies, group into waves, spawn subagents, collect results. Each subagent loads the full execute-plan context and handles its own plan.
Optional wave filter:
--wave N executes only Wave N for pacing, quota management, or staged rolloutFlag handling rule:
$ARGUMENTS$ARGUMENTS, treat it as inactiveContext budget: ~15% orchestrator, 100% fresh per subagent. </objective>
<execution_context> @$HOME/.claude/get-shit-done/workflows/execute-phase.md @$HOME/.claude/get-shit-done/references/ui-brand.md </execution_context>
<runtime_note> Copilot (VS Code): Use vscode_askquestions wherever this workflow calls AskUserQuestion. They are equivalent — vscode_askquestions is the VS Code Copilot implementation of the same interactive question API. </runtime_note>
<context> Phase: $ARGUMENTS
Available optional flags (documentation only — not automatically active):
--wave N — Execute only Wave N in the phase. Use when you want to pace execution or stay inside usage limits.--gaps-only — Execute only gap closure plans (plans with gap_closure: true in frontmatter). Use after verify-work creates fix plans.--interactive — Execute plans sequentially inline (no subagents) with user checkpoints between tasks. Lower token usage, pair-programming style. Best for small phases, bug fixes, and verification gaps.Active flags must be derived from $ARGUMENTS:
--wave N is active only if the literal --wave token is present in $ARGUMENTS--gaps-only is active only if the literal --gaps-only token is present in $ARGUMENTS--interactive is active only if the literal --interactive token is present in $ARGUMENTSContext files are resolved inside the workflow via gsd-tools init execute-phase and per-subagent <files_to_read> blocks. </context>
<process> Execute the execute-phase workflow from @$HOME/.claude/get-shit-done/workflows/execute-phase.md end-to-end. Preserve all workflow gates (wave execution, checkpoint handling, verification, state updates, routing). </process>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,961 | 11,466 | +28% | 1 | 1 | 0% | 533 | 1,011 | +90% | 0 | 0 | — |
case-02 | fail→fail | 9,378 | 45,098 | +381% | 1 | 1 | 0% | 825 | 976 | +18% | 0 | 0 | — |
case-03 | fail→fail | 11,380 | 17,074 | +50% | 1 | 1 | 0% | 847 | 915 | +8% | 0 | 0 | — |
case-04 | pass→pass | 16,642 | 8,844 | -47% | 1 | 1 | 0% | 1,816 | 1,372 | -24% | 0 | 0 | — |
case-05 | fail→fail | 17,968 | 8,420 | -53% | 1 | 1 | 0% | 1,935 | 1,038 | -46% | 0 | 0 | — |
case-06 | pass→pass | 7,463 | 9,001 | +21% | 1 | 1 | 0% | 1,266 | 1,368 | +8% | 0 | 0 | — |
case-07 | pass→pass | 24,129 | 10,024 | -58% | 1 | 1 | 0% | 3,170 | 2,371 | -25% | 0 | 0 | — |
case-08 | fail→fail | 13,038 | 37,916 | +191% | 1 | 1 | 0% | 2,079 | 6,043 | +191% | 0 | 0 | — |
case-09 | fail→pass | 10,772 | 2,420 | -78% | 1 | 1 | 0% | 1,702 | 1,062 | -38% | 0 | 0 | — |
case-10 | fail→pass | 10,033 | 7,719 | -23% | 1 | 1 | 0% | 1,663 | 2,141 | +29% | 0 | 0 | — |
case-11 | fail→pass | 10,211 | 1,927 | -81% | 1 | 1 | 0% | 1,321 | 985 | -25% | 0 | 0 | — |
case-12 | fail→pass | 15,352 | 12,999 | -15% | 1 | 1 | 0% | 1,679 | 2,045 | +22% | 0 | 0 | — |
case-13 | pass→pass | 10,375 | 4,442 | -57% | 1 | 1 | 0% | 813 | 1,250 | +54% | 0 | 0 | — |
case-14 | pass→fail | 13,605 | 12,380 | -9% | 1 | 1 | 0% | 2,104 | 1,473 | -30% | 0 | 0 | — |
case-15 | fail→pass | 9,581 | 10,639 | +11% | 1 | 1 | 0% | 1,594 | 1,642 | +3% | 0 | 0 | — |
case-16 | fail→pass | 10,620 | 11,445 | +8% | 1 | 1 | 0% | 1,686 | 1,625 | -4% | 0 | 0 | — |
case-17 | fail→pass | 14,538 | 2,870 | -80% | 1 | 1 | 0% | 1,573 | 950 | -40% | 0 | 0 | — |
case-18 | pass→pass | 17,643 | 16,685 | -5% | 1 | 1 | 0% | 1,858 | 2,344 | +26% | 0 | 0 | — |
case-19 | pass→fail | 46,308 | 5,217 | -89% | 1 | 1 | 0% | 6,214 | 868 | -86% | 0 | 0 | — |
case-20 | pass→fail | 17,289 | 50,636 | +193% | 1 | 1 | 0% | 1,826 | 1,107 | -39% | 0 | 0 | — |
case-21 | pass→fail | 20,888 | 50,191 | +140% | 1 | 1 | 0% | 2,958 | 1,062 | -64% | 0 | 0 | — |
case-22 | pass→pass | 17,117 | 19,342 | +13% | 1 | 1 | 0% | 2,329 | 3,183 | +37% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 15 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.