Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Archive completed milestone and prepare for next version
.claude/skills/davepoon-gsd-complete-milestone/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 248% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 37% | 0% |
<objective> Mark milestone {{version}} complete, archive to milestones/, and update ROADMAP.md and REQUIREMENTS.md.
Purpose: Create historical record of shipped version, archive milestone artifacts (roadmap + requirements), and prepare for next milestone. Output: Milestone archived (roadmap + requirements), PROJECT.md evolved, git tagged. </objective>
<execution_context> Load these files NOW (before proceeding):
</execution_context>
<context> Project files:
.planning/ROADMAP.md.planning/REQUIREMENTS.md.planning/STATE.md.planning/PROJECT.mdUser input:
</context>
<process>
Follow complete-milestone.md workflow:
.planning/v{{version}}-MILESTONE-AUDIT.md/gsd:audit-milestone firstgaps_found: recommend /gsd:plan-milestone-gaps firstpassed: proceed to step 1markdown ## Pre-flight Check
{If no v{{version}}-MILESTONE-AUDIT.md:} ⚠ No milestone audit found. Run /gsd:audit-milestone first to verify requirements coverage, cross-phase integration, and E2E flows.
{If audit has gaps:} ⚠ Milestone audit found gaps. Run /gsd:plan-milestone-gaps to create phases that close the gaps, or proceed anyway to accept as tech debt.
{If audit passed:} ✓ Milestone audit passed. Proceeding with completion.
.planning/milestones/v{{version}}-ROADMAP.md.planning/milestones/v{{version}}-REQUIREMENTS.md.planning/REQUIREMENTS.md (fresh one created for next milestone)<details> (if v1.1+)chore: archive v{{version}} milestonegit tag -a v{{version}} -m "[milestone summary]"/gsd:new-milestone — start next milestone (questioning → research → requirements → roadmap)</process>
<output_format> After the milestone is archived and tagged, emit a Milestone Complete continuation block following the pattern in references/continuation-format.md (§ Milestone Complete variant):
## 🎉 Milestone v{{version}} Complete with phase/plan/task summary## ▶ Next Up heading /clear then: before /gsd:new-milestone/gsd:audit-milestone (retrospective audit) or /gsd:review-backlog (deferred items review)Milestone close is the single biggest context-shed point in the workflow. The just-shipped milestone's plan/execute conversation is finished; the next milestone wants a clean slate. Always suggest /clear. </output_format>
<success_criteria>
.planning/milestones/v{{version}}-ROADMAP.md.planning/milestones/v{{version}}-REQUIREMENTS.md.planning/REQUIREMENTS.md deleted (fresh for next milestone)</success_criteria>
<critical_rules>
/gsd:new-milestone which includes requirements definition</critical_rules>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | fail→fail | 9,082 | 4,668 | -49% | 1 | 1 | 0% | 1,736 | 2,279 | +31% | 0 | 0 | — |
case-01 | fail→fail | 1,607 | 5,141 | +220% | 1 | 1 | 0% | 265 | 1,686 | +536% | 0 | 0 | — |
case-02 | fail→fail | 5,558 | 36,295 | +553% | 1 | 1 | 0% | 337 | 1,610 | +378% | 0 | 0 | — |
case-03 | fail→fail | 4,631 | 4,796 | +4% | 1 | 1 | 0% | 294 | 1,634 | +456% | 0 | 0 | — |
case-04 | pass→pass | 15,341 | 11,456 | -25% | 1 | 1 | 0% | 2,610 | 2,960 | +13% | 0 | 0 | — |
case-05 | pass→pass | 11,325 | 3,042 | -73% | 1 | 1 | 0% | 1,834 | 1,909 | +4% | 0 | 0 | — |
case-06 | pass→fail | 12,374 | 2,560 | -79% | 1 | 1 | 0% | 2,214 | 1,666 | -25% | 0 | 0 | — |
case-07 | fail→fail | 6,582 | 4,592 | -30% | 1 | 1 | 0% | 1,135 | 1,717 | +51% | 0 | 0 | — |
case-08 | fail→pass | 3,081 | 4,312 | +40% | 1 | 1 | 0% | 605 | 2,108 | +248% | 0 | 0 | — |
case-09 | fail→fail | 11,230 | 2,897 | -74% | 1 | 1 | 0% | 1,597 | 1,875 | +17% | 0 | 0 | — |
case-10 | fail→fail | 8,594 | 4,768 | -45% | 1 | 1 | 0% | 1,516 | 2,214 | +46% | 0 | 0 | — |
case-11 | fail→pass | 9,738 | 3,853 | -60% | 1 | 1 | 0% | 1,718 | 2,062 | +20% | 0 | 0 | — |
case-12 | fail→fail | 9,316 | 4,960 | -47% | 1 | 1 | 0% | 1,693 | 2,270 | +34% | 0 | 0 | — |
case-13 | fail→fail | 7,186 | 3,257 | -55% | 1 | 1 | 0% | 1,420 | 1,963 | +38% | 0 | 0 | — |
case-15 | fail→pass | 9,194 | 2,786 | -70% | 1 | 1 | 0% | 1,561 | 1,880 | +20% | 0 | 0 | — |
case-16 | fail→pass | 11,268 | 4,945 | -56% | 1 | 1 | 0% | 1,937 | 2,265 | +17% | 0 | 0 | — |
case-17 | fail→fail | 9,936 | 4,858 | -51% | 1 | 1 | 0% | 1,684 | 2,223 | +32% | 0 | 0 | — |
case-18 | fail→fail | 8,889 | 2,885 | -68% | 1 | 1 | 0% | 1,519 | 1,765 | +16% | 0 | 0 | — |
case-19 | fail→fail | 7,496 | 2,740 | -63% | 1 | 1 | 0% | 1,332 | 1,841 | +38% | 0 | 0 | — |
case-20 | fail→fail | 7,759 | 4,538 | -42% | 1 | 1 | 0% | 1,459 | 1,668 | +14% | 0 | 0 | — |
case-21 | fail→pass | 10,216 | 5,773 | -43% | 1 | 1 | 0% | 1,701 | 2,338 | +37% | 0 | 0 | — |
case-22 | fail→pass | 10,814 | 1,842 | -83% | 1 | 1 | 0% | 1,822 | 1,733 | -5% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 17 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.