Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when user specfically says 'plan harder'.
.claude/skills/aiskillstore-plan-harder/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 276% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 777% | 0% |
| case-22 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -55% | 0% |
Create detailed, phased implementation plans for bugs, features, or tasks. You make phased implementation plans with sprints and atomic tasks.
Use request_user_input to resolve ambiguities. Ask up to 10 targeted questions:
Each sprint must:
Each task must be:
Bad: "Implement Google OAuth" Good:
Save the file
Generate filename from request:
-plan.md suffixExamples:
xyz-bug-plan.mdgoogle-auth-plan.mdAFTER it is saved. Identify potential issues and edge cases in the plan. Address them proactively. Where could something go wrong? What about the plan is ambiguous? Is there a missing step, dependency, or pitfall?
Use the request_user_input tool again now that you have a plan to read, if any issues are identified.
Update the plan if you have improvements.
Provide the plan file location to a subagent for review, and ask it to provide feedback. Provide it useful context so it can make sound decisions. Explicitly tell it not to ask any questions. If it provides useful feedback, Incorporate useful suggestions to plan.
markdown# Plan: [Task Name] **Generated**: [Date] **Estimated Complexity**: [Low/Medium/High] ## Overview [Summary of task and approach] ## Prerequisites - [Dependencies or requirements] - [Tools, libraries, access needed] ## Sprint 1: [Name] **Goal**: [What this accomplishes] **Demo/Validation**: - [How to run/demo] - [What to verify] ### Task 1.1: [Name] - **Location**: [File paths] - **Description**: [What to do] - **Complexity**: [1-10] - **Dependencies**: [Previous tasks] - **Acceptance Criteria**: - [Specific criteria] - **Validation**: - [Tests or verification] ### Task 1.2: [Name] [...] ## Sprint 2: [Name] [...] ## Testing Strategy - [How to test] - [What to verify per sprint] ## Potential Risks & Gotchas - [What could go wrong] - [Mitigation strategies] ## Rollback Plan - [How to undo if needed]
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 18,747 | 14,189 | -24% | 1 | 1 | 0% | 2,236 | 1,199 | -46% | 0 | 0 | — |
case-02 | fail→fail | 18,816 | 31,820 | +69% | 1 | 1 | 0% | 285 | 1,230 | +332% | 0 | 0 | — |
case-03 | fail→fail | 71,977 | 17,426 | -76% | 1 | 1 | 0% | 6,505 | 1,125 | -83% | 0 | 0 | — |
case-04 | pass→fail | 26,297 | 14,200 | -46% | 1 | 1 | 0% | 4,023 | 1,810 | -55% | 0 | 0 | — |
case-05 | pass→fail | 19,998 | 15,486 | -23% | 1 | 1 | 0% | 3,204 | 1,358 | -58% | 0 | 0 | — |
case-06 | pass→fail | 22,217 | 5,094 | -77% | 1 | 1 | 0% | 2,663 | 1,216 | -54% | 0 | 0 | — |
case-07 | fail→fail | 18,659 | 13,594 | -27% | 1 | 1 | 0% | 3,235 | 1,113 | -66% | 0 | 0 | — |
case-08 | fail→pass | 12,580 | 38,728 | +208% | 1 | 1 | 0% | 2,067 | 7,766 | +276% | 0 | 0 | — |
case-09 | pass→fail | 29,729 | 9,403 | -68% | 1 | 1 | 0% | 4,131 | 1,098 | -73% | 0 | 0 | — |
case-10 | pass→pass | 10,536 | 9,013 | -14% | 1 | 1 | 0% | 1,740 | 2,537 | +46% | 0 | 0 | — |
case-11 | pass→pass | 17,452 | 20,527 | +18% | 1 | 1 | 0% | 1,894 | 3,379 | +78% | 0 | 0 | — |
case-12 | pass→fail | 18,475 | 4,771 | -74% | 1 | 1 | 0% | 1,958 | 1,205 | -38% | 0 | 0 | — |
case-13 | pass→fail | 22,510 | 13,683 | -39% | 1 | 1 | 0% | 3,376 | 1,086 | -68% | 0 | 0 | — |
case-14 | pass→fail | 22,115 | 4,858 | -78% | 1 | 1 | 0% | 2,621 | 1,208 | -54% | 0 | 0 | — |
case-15 | pass→pass | 12,251 | 21,529 | +76% | 1 | 1 | 0% | 1,850 | 4,688 | +153% | 0 | 0 | — |
case-16 | fail→pass | 22,806 | 31,145 | +37% | 1 | 1 | 0% | 3,444 | 6,478 | +88% | 0 | 0 | — |
case-17 | pass→fail | 4,932 | 29,748 | +503% | 1 | 1 | 0% | 774 | 2,760 | +257% | 0 | 0 | — |
case-18 | fail→fail | 37,769 | 14,467 | -62% | 1 | 1 | 0% | 8,030 | 1,191 | -85% | 0 | 0 | — |
case-19 | fail→pass | 10,850 | 44,870 | +314% | 1 | 1 | 0% | 966 | 8,474 | +777% | 0 | 0 | — |
case-20 | pass→pass | 16,037 | 14,099 | -12% | 1 | 1 | 0% | 1,734 | 3,249 | +87% | 0 | 0 | — |
case-21 | pass→pass | 16,530 | 27,027 | +64% | 1 | 1 | 0% | 2,041 | 6,339 | +211% | 0 | 0 | — |
case-22 | fail→pass | 16,324 | 8,336 | -49% | 1 | 1 | 0% | 1,722 | 1,418 | -18% | 0 | 0 | — |
case-23 | pass→fail | 21,611 | 9,718 | -55% | 1 | 1 | 0% | 1,869 | 1,200 | -36% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 9 counted toward the lift figure. The other 14 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -22 percentage points is the difference between those two pass rates over the 9 comparable cases. 10 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.