Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage. Use when starting complex tasks, multi-step projects, research tasks, or when the user mentions planning, organizing work, tracking progress, or wants structured output.
.claude/skills/aiskillstore-planning-with-files/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-07 | ✓→✓ | = Same ✓ | 1% | 0% |
| case-08 | ✓→✓ | = Same ✓ | -17% | 0% |
| case-11 | ✓→✓ | = Same ✓ | 114% | 0% |
Work like Manus: Use persistent markdown files as your "working memory on disk."
Before ANY complex task:
task_plan.md in the working directoryFor every non-trivial task, create THREE files:
| File | Purpose | When to Update | | ------------------ | --------------------------- | ---------------- | | task_plan.md | Track phases and progress | After each phase | | notes.md | Store findings and research | During research | | [deliverable].md | Final output | At completion |
Loop 1: Create task_plan.md with goal and phases
Loop 2: Research → save to notes.md → update task_plan.md
Loop 3: Read notes.md → create deliverable → update task_plan.md
Loop 4: Deliver final outputBefore each major action:
bashRead task_plan.md # Refresh goals in attention window
After each phase:
bashEdit task_plan.md # Mark [x], update status
When storing information:
bashWrite notes.md # Don't stuff context, store in file
Create this file FIRST for any complex task:
markdown# Task Plan: [Brief Description] ## Goal [One sentence describing the end state] ## Phases - [ ] Phase 1: Plan and setup - [ ] Phase 2: Research/gather information - [ ] Phase 3: Execute/build - [ ] Phase 4: Review and deliver ## Key Questions 1. [Question to answer] 2. [Question to answer] ## Decisions Made - [Decision]: [Rationale] ## Errors Encountered - [Error]: [Resolution] ## Status **Currently in Phase X** - [What I'm doing now]
For research and findings:
markdown# Notes: [Topic] ## Sources ### Source 1: [Name] - URL: [link] - Key points: - [Finding] - [Finding] ## Synthesized Findings ### [Category] - [Finding] - [Finding]
Never start a complex task without task_plan.md. This is non-negotiable.
Before any major decision, read the plan file. This keeps goals in your attention window.
After completing any phase, immediately update the plan file:
Large outputs go to files, not context. Keep only paths in working memory.
Every error goes in the "Errors Encountered" section. This builds knowledge for future tasks.
Use 3-file pattern for:
Skip for:
| Don't | Do Instead | | ----------------------------- | --------------------------------- | | Use TodoWrite for persistence | Create task_plan.md file | | State goals once and forget | Re-read plan before each decision | | Hide errors and retry | Log errors to plan file | | Stuff everything in context | Store large content in files | | Start executing immediately | Create plan file FIRST |
See reference.md for:
See examples.md for:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,282 | 19,938 | +30% | 1 | 1 | 0% | 509 | 1,872 | +268% | 0 | 0 | — |
case-02 | fail→fail | 14,656 | 11,846 | -19% | 1 | 1 | 0% | 309 | 1,322 | +328% | 0 | 0 | — |
case-03 | fail→fail | 24,097 | 18,136 | -25% | 1 | 1 | 0% | 508 | 1,926 | +279% | 0 | 0 | — |
case-04 | fail→fail | 14,610 | 28,593 | +96% | 1 | 1 | 0% | 2,424 | 2,340 | -3% | 0 | 0 | — |
case-05 | fail→fail | 13,003 | 6,202 | -52% | 1 | 1 | 0% | 1,924 | 1,907 | -1% | 0 | 0 | — |
case-06 | fail→fail | 11,869 | 10,936 | -8% | 1 | 1 | 0% | 1,777 | 2,676 | +51% | 0 | 0 | — |
case-07 | pass→pass | 28,451 | 15,055 | -47% | 1 | 1 | 0% | 3,380 | 3,406 | +1% | 0 | 0 | — |
case-08 | pass→pass | 18,339 | 14,180 | -23% | 1 | 1 | 0% | 2,965 | 2,461 | -17% | 0 | 0 | — |
case-09 | fail→fail | 10,585 | 10,591 | +0% | 1 | 1 | 0% | 1,671 | 1,656 | -1% | 0 | 0 | — |
case-10 | fail→pass | 12,719 | 10,670 | -16% | 1 | 1 | 0% | 1,877 | 2,628 | +40% | 0 | 0 | — |
case-11 | pass→pass | 4,658 | 9,203 | +98% | 1 | 1 | 0% | 749 | 1,605 | +114% | 0 | 0 | — |
case-12 | pass→pass | 18,156 | 31,210 | +72% | 1 | 1 | 0% | 3,328 | 3,877 | +16% | 0 | 0 | — |
case-13 | pass→pass | 24,547 | 27,681 | +13% | 1 | 1 | 0% | 3,082 | 5,696 | +85% | 0 | 0 | — |
case-14 | fail→fail | 10,987 | 2,600 | -76% | 1 | 1 | 0% | 952 | 1,375 | +44% | 0 | 0 | — |
case-15 | fail→fail | 12,593 | 6,574 | -48% | 1 | 1 | 0% | 1,380 | 1,897 | +37% | 0 | 0 | — |
case-16 | fail→fail | 10,310 | 4,437 | -57% | 1 | 1 | 0% | 878 | 1,780 | +103% | 0 | 0 | — |
case-17 | fail→pass | 9,270 | 9,920 | +7% | 1 | 1 | 0% | 1,530 | 1,747 | +14% | 0 | 0 | — |
case-18 | pass→pass | 11,683 | 4,163 | -64% | 1 | 1 | 0% | 1,141 | 1,714 | +50% | 0 | 0 | — |
case-19 | pass→pass | 8,410 | 8,495 | +1% | 1 | 1 | 0% | 1,376 | 1,574 | +14% | 0 | 0 | — |
case-20 | pass→pass | 4,903 | 4,001 | -18% | 1 | 1 | 0% | 820 | 1,595 | +95% | 0 | 0 | — |
case-21 | fail→fail | 2,391 | 10,634 | +345% | 1 | 1 | 0% | 330 | 1,177 | +257% | 0 | 0 | — |
case-22 | pass→pass | 1,757 | 6,326 | +260% | 1 | 1 | 0% | 187 | 1,122 | +500% | 0 | 0 | — |
case-23 | pass→pass | 3,836 | 8,412 | +119% | 1 | 1 | 0% | 679 | 1,559 | +130% | 0 | 0 | — |
case-24 | pass→pass | 13,544 | 9,705 | -28% | 1 | 1 | 0% | 1,427 | 1,810 | +27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 19 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +8 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.