Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Draft an implementation plan for a feature in .agents/plans. Does not execute the plan. Use when the user asks to plan, design, or scope a feature without building it.
.claude/skills/itsbencodes-create-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-09 | ✓→✗ | ▼ Worse | -54% | 0% |
| case-08 | ✓→✗ | ▼ Worse | -60% | 0% |
| case-12 | ✓→✗ | ▼ Worse | -82% | 0% |
| case-17 | ✓→✗ | ▼ Worse | 50% | 0% |
Write a plan to .agents/plans/. Do not implement.
Ask questions in batches until the feature is unambiguous. Read the code instead of asking whenever the answer is in the repo. Stop at nitpicks.
Present 2–4 high-level approaches. For each: shape, main upside, main cost. Ask which to take. Wait.
Save to .agents/plans/YYYY-MM-DD-<slug>.md (create the dir if missing). Use this template:
The plan can cover a feature, bug fix, refactor, or technical task. Adapt the user story and acceptance criteria accordingly (e.g., for a refactor, "user" may be a developer; for a bug fix, frame the story around the broken behavior).
markdown# <Title> ## User story As a <user>, I want <goal> so that <benefit>. ## Description A few paragraphs explaining what this is, the motivation, the current state, and the desired state. Include enough context that a reader unfamiliar with the work can understand it without external references. ## Acceptance criteria User-perspective only. No technical terms, file names, or APIs. - [ ] ... ## Out of scope - ... ## Implementation details Architecture, key types, dependencies, relevant patterns, trade-offs. ## Testing strategy What to cover and how: unit, integration, e2e (Playwright). Note any new fixtures, mocks, or manual verification steps. ## Tasks - [ ] ... - [ ] Run the `validate` skill (must be the last task) ## Open questions - ...
Tell the user the file path.
Apply requested edits, repeat until approved. Do not implement.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,393 | 3,944 | -74% | 1 | 1 | 0% | 2,221 | 612 | -72% | 0 | 0 | — |
case-02 | fail→fail | 15,362 | 3,871 | -75% | 1 | 1 | 0% | 2,769 | 615 | -78% | 0 | 0 | — |
case-03 | fail→fail | 15,151 | 5,574 | -63% | 1 | 1 | 0% | 2,164 | 585 | -73% | 0 | 0 | — |
case-04 | fail→fail | 4,832 | 4,422 | -8% | 1 | 1 | 0% | 136 | 755 | +455% | 0 | 0 | — |
case-09 | pass→fail | 11,524 | 4,442 | -61% | 1 | 1 | 0% | 1,786 | 818 | -54% | 0 | 0 | — |
case-05 | fail→fail | 59,303 | 4,115 | -93% | 1 | 1 | 0% | 3,990 | 534 | -87% | 0 | 0 | — |
case-06 | fail→fail | 9,393 | 10,990 | +17% | 1 | 1 | 0% | 1,735 | 1,461 | -16% | 0 | 0 | — |
case-07 | fail→fail | 15,206 | 3,748 | -75% | 1 | 1 | 0% | 2,513 | 613 | -76% | 0 | 0 | — |
case-08 | pass→fail | 12,444 | 4,484 | -64% | 1 | 1 | 0% | 1,905 | 756 | -60% | 0 | 0 | — |
case-19 | pass→pass | 13,753 | 11,052 | -20% | 1 | 1 | 0% | 2,212 | 1,467 | -34% | 0 | 0 | — |
case-10 | fail→fail | 2,240 | 3,868 | +73% | 1 | 1 | 0% | 275 | 820 | +198% | 0 | 0 | — |
case-11 | fail→fail | 21,236 | 2,903 | -86% | 1 | 1 | 0% | 3,231 | 597 | -82% | 0 | 0 | — |
case-12 | pass→fail | 21,823 | 3,667 | -83% | 1 | 1 | 0% | 4,048 | 721 | -82% | 0 | 0 | — |
case-13 | fail→fail | 24,511 | 3,839 | -84% | 1 | 1 | 0% | 3,796 | 656 | -83% | 0 | 0 | — |
case-14 | fail→fail | 19,111 | 3,298 | -83% | 1 | 1 | 0% | 3,156 | 518 | -84% | 0 | 0 | — |
case-15 | fail→pass | 23,602 | 28,119 | +19% | 1 | 1 | 0% | 3,788 | 4,595 | +21% | 0 | 0 | — |
case-16 | pass→pass | 15,620 | 9,427 | -40% | 1 | 1 | 0% | 2,701 | 1,997 | -26% | 0 | 0 | — |
case-17 | pass→fail | 3,643 | 25,371 | +596% | 1 | 1 | 0% | 600 | 901 | +50% | 0 | 0 | — |
case-18 | fail→fail | 17,825 | 5,704 | -68% | 1 | 1 | 0% | 3,541 | 747 | -79% | 0 | 0 | — |
case-20 | fail→fail | 22,947 | 5,457 | -76% | 1 | 1 | 0% | 5,568 | 739 | -87% | 0 | 0 | — |
case-21 | fail→fail | 2,865 | 21,756 | +659% | 1 | 1 | 0% | 393 | 3,614 | +820% | 0 | 0 | — |
case-22 | fail→fail | 3,609 | 5,289 | +47% | 1 | 1 | 0% | 507 | 726 | +43% | 0 | 0 | — |
case-23 | pass→fail | 4,843 | 2,946 | -39% | 1 | 1 | 0% | 849 | 579 | -32% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 7 counted toward the lift figure. The other 16 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -33 percentage points is the difference between those two pass rates over the 7 comparable cases. 9 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.