Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Structured task planning with clear breakdowns, dependencies, and verification criteria. Use when implementing features, refactoring, or any multi-step work.
.claude/skills/sickn33-plan-writing/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -47% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -21% | 0% |
> Source: obra/superpowers
This skill provides a framework for breaking down work into clear, actionable tasks with verification criteria.
{task-slug}.md in the PROJECT ROOTauth-feature.md).claude/, docs/, or temp folders> 🔴 NO fixed templates. Each plan is UNIQUE to the task.
| ❌ Wrong | ✅ Right | |----------|----------| | 50 tasks with sub-sub-tasks | 5-10 clear tasks max | | Every micro-step listed | Only actionable items | | Verbose descriptions | One-line per task |
> Rule: If plan is longer than 1 page, it's too long. Simplify.
| ❌ Wrong | ✅ Right | |----------|----------| | "Set up project" | "Run npx create-next-app" | | "Add authentication" | "Install next-auth, create /api/auth/[...nextauth].ts" | | "Style the UI" | "Add Tailwind classes to Header.tsx" |
> Rule: Each task should have a clear, verifiable outcome.
For NEW PROJECT:
For FEATURE ADDITION:
For BUG FIX:
> 🔴 DO NOT copy-paste script commands. Choose based on project type.
| Project Type | Relevant Scripts | |--------------|------------------| | Frontend/React | ux_audit.py, accessibility_checker.py | | Backend/API | api_validator.py, security_scan.py | | Mobile | mobile_audit.py | | Database | schema_validator.py | | Full-stack | Mix of above based on what you touched |
Wrong: Adding all scripts to every plan Right: Only scripts relevant to THIS task
| ❌ Wrong | ✅ Right | |----------|----------| | "Verify the component works correctly" | "Run npm run dev, click button, see toast" | | "Test the API" | "curl localhost:3000/api/users returns 200" | | "Check styles" | "Open browser, verify dark mode toggle works" |
# [Task Name]
## Goal
One sentence: What are we building/fixing?
## Tasks
- [ ] Task 1: [Specific action] → Verify: [How to check]
- [ ] Task 2: [Specific action] → Verify: [How to check]
- [ ] Task 3: [Specific action] → Verify: [How to check]
## Done When
- [ ] [Main success criteria]> That's it. No phases, no sub-sections unless truly needed. > Keep it minimal. Add complexity only when required.
Any important considerations]
---
## Best Practices (Quick Reference)
1. **Start with goal** - What are we building/fixing?
2. **Max 10 tasks** - If more, break into multiple plans
3. **Each task verifiable** - Clear "done" criteria
4. **Project-specific** - No copy-paste templates
5. **Update as you go** - Mark `[x]` when complete
---
## When to Use
- New project from scratch
- Adding a feature
- Fixing a bug (if complex)
- Refactoring multiple files
## Limitations
- Use this skill only when the task clearly matches the scope described above.
- Do not treat the output as a substitute for environment-specific validation, testing, or expert review.
- Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 10,142 | 6,491 | -36% | 1 | 1 | 0% | 2,089 | 2,293 | +10% | 0 | 0 | — |
case-02 | fail→fail | 3,571 | 10,318 | +189% | 1 | 1 | 0% | 261 | 3,009 | +1053% | 0 | 0 | — |
case-03 | fail→fail | 13,207 | 6,414 | -51% | 1 | 1 | 0% | 2,394 | 2,346 | -2% | 0 | 0 | — |
case-04 | pass→pass | 17,632 | 16,866 | -4% | 1 | 1 | 0% | 2,806 | 3,947 | +41% | 0 | 0 | — |
case-05 | pass→pass | 3,724 | 5,669 | +52% | 1 | 1 | 0% | 671 | 2,060 | +207% | 0 | 0 | — |
case-06 | pass→pass | 15,216 | 16,957 | +11% | 1 | 1 | 0% | 3,655 | 5,004 | +37% | 0 | 0 | — |
case-07 | fail→pass | 15,795 | 9,130 | -42% | 1 | 1 | 0% | 3,145 | 2,904 | -8% | 0 | 0 | — |
case-08 | fail→pass | 11,887 | 6,182 | -48% | 1 | 1 | 0% | 2,403 | 2,233 | -7% | 0 | 0 | — |
case-09 | pass→pass | 13,611 | 7,877 | -42% | 1 | 1 | 0% | 2,558 | 2,515 | -2% | 0 | 0 | — |
case-10 | pass→pass | 10,385 | 5,478 | -47% | 1 | 1 | 0% | 1,938 | 2,251 | +16% | 0 | 0 | — |
case-11 | fail→pass | 33,451 | 12,518 | -63% | 1 | 1 | 0% | 6,180 | 3,289 | -47% | 0 | 0 | — |
case-12 | fail→pass | 17,438 | 7,474 | -57% | 1 | 1 | 0% | 3,190 | 2,524 | -21% | 0 | 0 | — |
case-13 | pass→pass | 14,617 | 6,853 | -53% | 1 | 1 | 0% | 2,745 | 2,480 | -10% | 0 | 0 | — |
case-14 | pass→pass | 13,627 | 6,207 | -54% | 1 | 1 | 0% | 2,314 | 2,129 | -8% | 0 | 0 | — |
case-15 | pass→pass | 10,143 | 6,803 | -33% | 1 | 1 | 0% | 1,887 | 2,408 | +28% | 0 | 0 | — |
case-16 | fail→pass | 13,133 | 6,977 | -47% | 1 | 1 | 0% | 2,298 | 2,289 | -0% | 0 | 0 | — |
case-17 | fail→fail | 16,744 | 7,361 | -56% | 1 | 1 | 0% | 3,137 | 2,406 | -23% | 0 | 0 | — |
case-18 | fail→fail | 19,489 | 15,554 | -20% | 1 | 1 | 0% | 3,845 | 4,063 | +6% | 0 | 0 | — |
case-19 | pass→pass | 8,952 | 5,430 | -39% | 1 | 1 | 0% | 1,755 | 2,092 | +19% | 0 | 0 | — |
case-20 | pass→pass | 7,707 | 4,610 | -40% | 1 | 1 | 0% | 1,586 | 1,937 | +22% | 0 | 0 | — |
case-21 | pass→pass | 11,635 | 7,340 | -37% | 1 | 1 | 0% | 2,208 | 2,479 | +12% | 0 | 0 | — |
case-22 | pass→pass | 14,824 | 6,258 | -58% | 1 | 1 | 0% | 2,888 | 2,286 | -21% | 0 | 0 | — |
case-23 | fail→pass | 15,224 | 6,433 | -58% | 1 | 1 | 0% | 2,863 | 2,275 | -21% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +30 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.