Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Read-only planning and architecture analysis for Penpot — produce a structured implementation plan (Context, Affected modules, Approach, Risks, Testing). Always output to the user; additionally save to .opencode/plans/YYYY-MM-DD-<title>.md.
.claude/skills/penpot-planner/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 69% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 268% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 327% | 0% |
Produce a plan that another engineer or agent can execute without guessing.
names, and test strategy.
breakdown.
Do not use for a small change with obvious scope or an existing executable plan.
Before drafting any plan, work through the project's own guidance:
critical-info (.serena/memories/critical-info.md) — the entry pointthat describes the monorepo structure and module dependency graph.
critical-info, identify which modules your task affects.mem:frontend/core,mem:backend/core, mem:common/core, mem:exporter/core, mem:render-wasm/core. Follow mem: references deeper as needed.
plan can include concrete verification steps.
Skipping this step is the #1 cause of incorrect or incomplete plans.
only file you may write is the plan itself, and only when the command or user explicitly instructs you to save it.
rg, ls,find, cat, bat).
engineer agent or developer.
before their consumers.
details from existing conventions when they do not affect public behavior.
changes, and external dependencies.
over unrelated layer-wide batches. Apply DRY and KISS to the proposed implementation.
Each task follows this structure:
markdown### Task [N]: [Short descriptive title] **Description:** One or two paragraphs explaining what this task accomplishes. Should be clear and concise. **Rationale:** Why this task exists and why this approach over the obvious alternatives — design decisions, trade-offs, constraints discovered during analysis. One or two sentences; skip only if genuinely trivial. **Code sketch (optional):** Signature-, type-, or shape-level example when the intended interface is non-obvious. Keep it short — a skeleton that fixes the contract (function signature, model fields, error shape), never a full implementation. Omit when the task is mechanical. **Acceptance criteria:** - [ ] [Specific, testable condition] - [ ] [Specific, testable condition] **Verification:** - [ ] Relevant tests pass (module-specific test command). - [ ] Lint/formatter passes (module-specific check command), if applicable. - [ ] The core flow works end-to-end, if applicable. **Dependencies:** [Task numbers this depends on, or "None"] **Files likely touched:** - `path/to/file.clj` - `path/to/file_test.clj` **Estimated scope:** [XS: 1 file | S: 1-2 files | M: 3-5 files | L: 5+ files]
Use commands from mem:testing and affected module memories. Never substitute generic text such as "run the tests" when the project documents an exact command.
When possible, design each task with TDD in mind: acceptance criteria double as a test list, and the natural first step of the task is writing those tests before the implementation. Some tasks resist this (config, migrations, pure wiring) — for those, keep the usual verification steps.
| Size | Files | Scope | Example | |------|-------|-------|---------| | XS | 1 | Single function, config change, or schema tweak | Add a validation rule | | S | 1-2 | One handler or component method | Add a new RPC endpoint | | M | 3-5 | One vertical feature slice | Bookmark CRUD with tests | | L | 5-8 | Multi-component feature | Search with filtering and pagination | | XL | 8+ | Too large — break it down further | — |
Split a task when it contains independent outcomes, spans unrelated systems, or cannot be completed and verified in one focused session (if a task is XL, it should be broken into smaller tasks; agents perform best on S and M tasks).
Arrange tasks so that:
Add explicit checkpoints with the relevant module commands:
markdown### Checkpoint: After Tasks 1-3 - [ ] Relevant tests pass (module-specific command). - [ ] The relevant build or compilation passes, if applicable. - [ ] The core flow works end-to-end.
The plan is always delivered in the response so the user sees it regardless of which agent is running the skill. File writes follow Constraints — by default announce the path instead of writing.
Announce the save path .agents/plans/YYYY-MM-DD-<slug>.md (today's date, lowercase hyphen-separated slug, e.g. 2026-09-10-add-batch-get-profiles; an explicit user path wins).
Never invent a fresh slug when the plan derives from an existing one. The derived name is <parent-basename> plus one suffix per level, joined with -- (double hyphen; single hyphens already separate slug words, so -- marks where the derivation starts). The parent name is never edited, and no new date is added — the parent prefix already carries its date, which keeps parent and derivatives adjacent in ls. Record the real creation date inside the plan (Created:).
Valid names match:
^\d{4}-\d{2}-\d{2}-[a-z0-9-]+(--(review-\d{2}|task-\d{2})(-[a-z0-9-]+)?)*\.md$review-NN — a new plan addressing findings of a review-code orreview-plan on already-implemented work. NN counts reviews of that parent from 01. Example: parent 2026-09-14-paste-before-init-crash.md → 2026-09-14-paste-before-init-crash--review-01.md, then --review-02.md.
task-NN-<short-slug> — sub-plan for task NN of a high-levelroadmap plan. NN is the roadmap task number. Example: parent 2026-09-20-upload-pipeline-roadmap.md → 2026-09-20-upload-pipeline-roadmap--task-01-chunk-upload.md. Levels chain: ...--task-02-gc--review-01.md.
Never use v2, final, new, or fix2 as suffixes. Keep the optional short slug to 3-4 lowercase hyphen-separated words.
Every derived plan opens its Context with:
markdownParent: `<parent-basename>.md` Source: review-code over `<commit>` (branch `<branch>`) | task `NN` of roadmap `<parent-basename>.md` Created: YYYY-MM-DD
End the response by suggesting the next steps: /review-plan to get a second opinion on the plan and /implement-plan to execute it.
Use this document shape:
markdown# Plan: Title Status: draft | reviewed | done Review Log: - YYYY-MM-DDTHH:MM:SSZ — <what changed and why, one line per entry> ## Context ## Affected Modules ## Architecture Decisions ## Risks and Considerations ## Approach ## Task List ## Verification and Testing ## Parallelization ## Open Questions
Status and Review Log are write-restricted metadata, not free text:
make-a-plan creates every plan with Status: draft and an emptylog. reviewed never means "a review was emitted" — it means "review feedback was incorporated". A plan approved with no changes goes draft → done without passing through reviewed.
when the user says to apply review-plan findings, make-a-plan applies the changes, flips to reviewed, and appends one UTC ISO 8601 line describing what changed.
implement-plan flips to done on completion and appends one linewith the issue URL when one exists (standalone mode); otherwise just done. Never record commit hashes — they rot on amend/rebase and git already links the commit.
review-plan and review-code never write these fields.Omit empty sections only when they do not apply. Every implementation task still requires acceptance criteria, verification, dependencies, likely files, and scope.
When the plan is purely analytical (e.g. a code review or feasibility study with no implementation), skip the Approach and Task List sections and lead with Findings instead, keeping the rest of the structure.
| Rationalization | Reality | |---|---| | "I'll figure it out as I go" | That's how you end up with a tangled mess and rework. 10 minutes of planning saves hours. | | "The tasks are obvious" | Write them down anyway. Explicit tasks surface hidden dependencies and forgotten edge cases. | | "Planning is overhead" | Planning is the task. Implementation without a plan is just typing. | | "I can hold it all in my head" | Context windows are finite. Written plans survive session boundaries and compaction. |
Before delivering the plan, confirm:
/review-plan and /implement-plan
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 24,095 | 7,082 | -71% | 1 | 1 | 0% | 4,264 | 3,023 | -29% | 0 | 0 | — |
case-02 | fail→fail | 25,740 | 9,144 | -64% | 1 | 1 | 0% | 5,211 | 3,442 | -34% | 0 | 0 | — |
case-03 | fail→fail | 48,276 | 7,073 | -85% | 1 | 1 | 0% | 8,266 | 3,016 | -64% | 0 | 0 | — |
case-04 | fail→fail | 20,203 | 9,000 | -55% | 1 | 1 | 0% | 212 | 3,153 | +1387% | 0 | 0 | — |
case-05 | pass→pass | 4,809 | 12,215 | +154% | 1 | 1 | 0% | 643 | 3,665 | +470% | 0 | 0 | — |
case-06 | pass→fail | 5,498 | 7,241 | +32% | 1 | 1 | 0% | 865 | 2,967 | +243% | 0 | 0 | — |
case-07 | fail→fail | 19,139 | 11,646 | -39% | 1 | 1 | 0% | 3,012 | 3,256 | +8% | 0 | 0 | — |
case-08 | fail→pass | 24,436 | 6,334 | -74% | 1 | 1 | 0% | 4,713 | 3,980 | -16% | 0 | 0 | — |
case-09 | fail→fail | 31,979 | 13,727 | -57% | 1 | 1 | 0% | 2,886 | 3,279 | +14% | 0 | 0 | — |
case-10 | fail→fail | 19,759 | 7,879 | -60% | 1 | 1 | 0% | 2,895 | 3,032 | +5% | 0 | 0 | — |
case-11 | fail→fail | 18,316 | 7,302 | -60% | 1 | 1 | 0% | 2,687 | 2,930 | +9% | 0 | 0 | — |
case-12 | fail→pass | 17,015 | 9,212 | -46% | 1 | 1 | 0% | 2,390 | 4,109 | +72% | 0 | 0 | — |
case-13 | fail→fail | 14,112 | 10,270 | -27% | 1 | 1 | 0% | 2,021 | 3,087 | +53% | 0 | 0 | — |
case-14 | fail→pass | 14,408 | 5,798 | -60% | 1 | 1 | 0% | 2,078 | 3,519 | +69% | 0 | 0 | — |
case-15 | fail→pass | 7,309 | 7,210 | -1% | 1 | 1 | 0% | 1,011 | 3,719 | +268% | 0 | 0 | — |
case-16 | fail→pass | 7,040 | 6,010 | -15% | 1 | 1 | 0% | 827 | 3,532 | +327% | 0 | 0 | — |
case-17 | fail→pass | 11,386 | 5,973 | -48% | 1 | 1 | 0% | 1,564 | 3,734 | +139% | 0 | 0 | — |
case-18 | fail→fail | 24,145 | 9,225 | -62% | 1 | 1 | 0% | 3,569 | 3,028 | -15% | 0 | 0 | — |
case-19 | fail→pass | 6,240 | 4,956 | -21% | 1 | 1 | 0% | 1,069 | 3,383 | +216% | 0 | 0 | — |
case-20 | pass→pass | 15,035 | 6,232 | -59% | 1 | 1 | 0% | 1,965 | 3,456 | +76% | 0 | 0 | — |
case-21 | fail→pass | 9,124 | 4,569 | -50% | 1 | 1 | 0% | 1,249 | 3,461 | +177% | 0 | 0 | — |
case-22 | pass→pass | 16,843 | 14,418 | -14% | 1 | 1 | 0% | 2,700 | 4,649 | +72% | 0 | 0 | — |
case-23 | fail→pass | 8,431 | 124,635 | +1378% | 1 | 1 | 0% | 1,289 | 3,347 | +160% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 12 counted toward the lift figure. The other 11 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +35 percentage points is the difference between those two pass rates over the 12 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.