Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create a detailed, phased implementation plan with documentation discovery. Use when asked to plan a feature, task, or multi-step implementation — especially before executing with do.
.claude/skills/thedotmack-make-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 1% | 0% |
You are an ORCHESTRATOR. Create an LLM-friendly plan in phases that can be executed consecutively in new chat contexts.
Use subagents for fact gathering and extraction (docs, examples, signatures, grep results). Keep synthesis and plan authoring with the orchestrator (phase boundaries, task framing, final wording). If a subagent report is incomplete or lacks evidence, re-check with targeted reads/greps before finalizing.
Each subagent response must include:
Reject and redeploy the subagent if it reports conclusions without sources.
Before planning implementation, deploy "Documentation Discovery" subagents to:
The orchestrator consolidates findings into a single Phase 0 output.
oh-my-issues — the issue-side sibling. When the plan you're being asked to make is rooted in a bug or feature backlog rather than a fresh idea, route through oh-my-issues first to cluster issues by root cause into plan masters and plans/0X-*.md design docs. make-plan then operates on the design doc for one plan slice.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,356 | 3,796 | -75% | 1 | 1 | 0% | 3,101 | 983 | -68% | 0 | 0 | — |
case-02 | fail→fail | 20,824 | 6,153 | -70% | 1 | 1 | 0% | 4,406 | 1,242 | -72% | 0 | 0 | — |
case-03 | fail→fail | 19,544 | 6,692 | -66% | 1 | 1 | 0% | 4,366 | 978 | -78% | 0 | 0 | — |
case-04 | pass→pass | 9,234 | 3,907 | -58% | 1 | 1 | 0% | 1,669 | 1,494 | -10% | 0 | 0 | — |
case-05 | fail→fail | 13,203 | 9,913 | -25% | 1 | 1 | 0% | 2,527 | 2,815 | +11% | 0 | 0 | — |
case-06 | pass→pass | 7,498 | 6,186 | -17% | 1 | 1 | 0% | 1,725 | 1,821 | +6% | 0 | 0 | — |
case-07 | pass→pass | 9,381 | 2,500 | -73% | 1 | 1 | 0% | 1,738 | 1,196 | -31% | 0 | 0 | — |
case-08 | fail→pass | 4,593 | 1,637 | -64% | 1 | 1 | 0% | 835 | 989 | +18% | 0 | 0 | — |
case-09 | fail→pass | 9,829 | 4,862 | -51% | 1 | 1 | 0% | 1,690 | 1,651 | -2% | 0 | 0 | — |
case-10 | pass→fail | 7,864 | 2,424 | -69% | 1 | 1 | 0% | 1,542 | 1,180 | -23% | 0 | 0 | — |
case-11 | fail→pass | 10,071 | 5,554 | -45% | 1 | 1 | 0% | 1,815 | 1,757 | -3% | 0 | 0 | — |
case-12 | fail→pass | 13,086 | 11,901 | -9% | 1 | 1 | 0% | 2,623 | 2,867 | +9% | 0 | 0 | — |
case-13 | fail→pass | 5,248 | 1,613 | -69% | 1 | 1 | 0% | 1,014 | 1,026 | +1% | 0 | 0 | — |
case-14 | fail→pass | 3,919 | 2,056 | -48% | 1 | 1 | 0% | 676 | 1,088 | +61% | 0 | 0 | — |
case-15 | pass→pass | 8,798 | 3,714 | -58% | 1 | 1 | 0% | 1,595 | 1,384 | -13% | 0 | 0 | — |
case-16 | pass→pass | 8,593 | 3,352 | -61% | 1 | 1 | 0% | 1,293 | 1,279 | -1% | 0 | 0 | — |
case-17 | pass→pass | 7,770 | 4,491 | -42% | 1 | 1 | 0% | 1,655 | 1,525 | -8% | 0 | 0 | — |
case-18 | fail→pass | 7,346 | 3,283 | -55% | 1 | 1 | 0% | 1,340 | 1,369 | +2% | 0 | 0 | — |
case-19 | fail→pass | 3,961 | 5,653 | +43% | 1 | 1 | 0% | 621 | 1,959 | +215% | 0 | 0 | — |
case-20 | fail→fail | 7,180 | 4,364 | -39% | 1 | 1 | 0% | 1,559 | 1,088 | -30% | 0 | 0 | — |
case-21 | fail→pass | 9,599 | 5,653 | -41% | 1 | 1 | 0% | 1,961 | 1,748 | -11% | 0 | 0 | — |
case-22 | pass→pass | 8,850 | 5,540 | -37% | 1 | 1 | 0% | 1,679 | 1,706 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.