Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Archive completed milestone and prepare for next version
.claude/skills/coco-research-gsd-complete-milestone/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 23% | 0% |
<objective> Mark milestone $ARGUMENTS complete, archive to milestones/, and update ROADMAP.md and REQUIREMENTS.md.
Purpose: Create historical record of shipped version, archive milestone artifacts (roadmap + requirements), and prepare for next milestone. Output: Milestone archived (roadmap + requirements), PROJECT.md evolved, git tagged. </objective>
<execution_context> Load these files NOW (before proceeding):
</execution_context>
<context> Project files:
.planning/ROADMAP.md.planning/REQUIREMENTS.md.planning/STATE.md.planning/PROJECT.mdUser input:
</context>
<process>
Follow complete-milestone.md workflow:
.planning/v$ARGUMENTS-MILESTONE-AUDIT.md/gsd-audit-milestone firstgaps_found: recommend /gsd-plan-milestone-gaps firstpassed: proceed to step 1markdown ## Pre-flight Check
{If no v$ARGUMENTS-MILESTONE-AUDIT.md:} ⚠ No milestone audit found. Run /gsd-audit-milestone first to verify requirements coverage, cross-phase integration, and E2E flows.
{If audit has gaps:} ⚠ Milestone audit found gaps. Run /gsd-plan-milestone-gaps to create phases that close the gaps, or proceed anyway to accept as tech debt.
{If audit passed:} ✓ Milestone audit passed. Proceeding with completion.
.planning/milestones/v$ARGUMENTS-ROADMAP.md.planning/milestones/v$ARGUMENTS-REQUIREMENTS.md.planning/REQUIREMENTS.md (fresh one created for next milestone)<details> (if v1.1+)chore: archive v$ARGUMENTS milestonegit tag -a v$ARGUMENTS -m "[milestone summary]"/gsd-new-milestone — start next milestone (questioning → research → requirements → roadmap)</process>
<success_criteria>
.planning/milestones/v$ARGUMENTS-ROADMAP.md.planning/milestones/v$ARGUMENTS-REQUIREMENTS.md.planning/REQUIREMENTS.md deleted (fresh for next milestone)</success_criteria>
<critical_rules>
/gsd-new-milestone which includes requirements definition</critical_rules>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | fail→pass | 14,753 | 11,354 | -23% | 1 | 1 | 0% | 1,780 | 2,687 | +51% | 0 | 0 | — |
case-01 | fail→fail | 15,277 | 15,988 | +5% | 1 | 1 | 0% | 1,363 | 1,418 | +4% | 0 | 0 | — |
case-02 | fail→fail | 18,746 | 6,041 | -68% | 1 | 1 | 0% | 2,118 | 1,396 | -34% | 0 | 0 | — |
case-03 | fail→fail | 18,932 | 25,250 | +33% | 1 | 1 | 0% | 1,830 | 1,457 | -20% | 0 | 0 | — |
case-04 | fail→fail | 14,939 | 12,475 | -16% | 1 | 1 | 0% | 1,404 | 2,698 | +92% | 0 | 0 | — |
case-05 | fail→fail | 7,204 | 11,122 | +54% | 1 | 1 | 0% | 1,213 | 1,390 | +15% | 0 | 0 | — |
case-06 | fail→pass | 14,415 | 5,396 | -63% | 1 | 1 | 0% | 1,369 | 2,076 | +52% | 0 | 0 | — |
case-07 | fail→fail | 9,818 | 12,142 | +24% | 1 | 1 | 0% | 750 | 1,547 | +106% | 0 | 0 | — |
case-08 | fail→pass | 13,377 | 8,574 | -36% | 1 | 1 | 0% | 1,407 | 1,785 | +27% | 0 | 0 | — |
case-09 | fail→pass | 10,671 | 10,992 | +3% | 1 | 1 | 0% | 1,721 | 2,236 | +30% | 0 | 0 | — |
case-11 | pass→pass | 13,004 | 8,818 | -32% | 1 | 1 | 0% | 1,148 | 1,641 | +43% | 0 | 0 | — |
case-12 | fail→pass | 7,978 | 7,910 | -1% | 1 | 1 | 0% | 1,217 | 1,494 | +23% | 0 | 0 | — |
case-13 | pass→fail | 16,889 | 6,916 | -59% | 1 | 1 | 0% | 1,667 | 1,424 | -15% | 0 | 0 | — |
case-14 | fail→fail | 10,258 | 9,569 | -7% | 1 | 1 | 0% | 1,587 | 1,881 | +19% | 0 | 0 | — |
case-15 | fail→pass | 14,846 | 7,224 | -51% | 1 | 1 | 0% | 1,640 | 1,501 | -8% | 0 | 0 | — |
case-16 | pass→pass | 14,876 | 1,910 | -87% | 1 | 1 | 0% | 1,255 | 1,401 | +12% | 0 | 0 | — |
case-17 | pass→pass | 13,050 | 3,739 | -71% | 1 | 1 | 0% | 1,375 | 1,593 | +16% | 0 | 0 | — |
case-18 | fail→pass | 17,365 | 8,231 | -53% | 1 | 1 | 0% | 2,106 | 1,645 | -22% | 0 | 0 | — |
case-19 | fail→fail | 12,764 | 8,004 | -37% | 1 | 1 | 0% | 1,307 | 1,681 | +29% | 0 | 0 | — |
case-20 | pass→fail | 18,756 | 10,963 | -42% | 1 | 1 | 0% | 2,325 | 1,469 | -37% | 0 | 0 | — |
case-21 | pass→fail | 22,428 | 37,225 | +66% | 1 | 1 | 0% | 2,831 | 1,620 | -43% | 0 | 0 | — |
case-22 | pass→fail | 17,313 | 17,822 | +3% | 1 | 1 | 0% | 1,960 | 1,589 | -19% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 13 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 13 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.