Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Plan mode: write markdown plan, no execution.
.claude/skills/hezaohezao-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 82% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -38% | 0% |
| case-08 | ✓→✗ | ▼ Worse | 51% | 0% |
| case-09 | ✓→✗ | ▼ Worse | 7% | 0% |
| case-12 | ✓→✗ | ▼ Worse | -33% | 0% |
Use this skill when the user wants a plan instead of execution.
For this turn, you are planning only.
.poirot/plans/.Write a markdown plan that is concrete and actionable.
Include, when relevant:
If the task is code-related, include exact file paths, likely test targets, and verification steps.
Save the plan with write_file under:
.poirot/plans/YYYY-MM-DD_HHMMSS-<slug>.mdTreat that as relative to the active working directory / sandbox workspace. Poirot sandbox file tools are path-aware, so using this relative path keeps the plan with the workspace.
If no specific target path is provided by the runtime, create a sensible timestamped filename yourself under .poirot/plans/.
The rest of this skill is the craft of authoring a good implementation plan — the content that goes inside the markdown file above.
Write comprehensive implementation plans assuming the implementer has zero context for the codebase and questionable taste. Document everything they need: which files to touch, complete code, testing commands, docs to check, how to verify. Give them bite-sized tasks. DRY. YAGNI. TDD. Frequent commits.
Assume the implementer is a skilled developer but knows almost nothing about the toolset or problem domain.
Core principle: A good plan makes implementation obvious. If someone has to guess, the plan is incomplete.
Always use before:
Don't skip when:
Each task = 2-5 minutes of focused work.
Every step is one action:
Too big:
markdown### Task 1: Build authentication system [50 lines of code across 5 files]
Right size:
markdown### Task 1: Create User model with email field [10 lines, 1 file] ### Task 2: Add password hash field to User [8 lines, 1 file] ### Task 3: Create password hashing utility [15 lines, 1 file]
Every plan MUST start with:
markdown# [Feature Name] Implementation Plan **Goal:** [One sentence describing what this builds] **Architecture:** [2-3 sentences about approach] **Tech Stack:** [Key technologies/libraries] ---
Each task follows this format:
`markdown### Task N: [Descriptive Name] **Objective:** What this task accomplishes (one sentence) **Files:** - Create: `exact/path/to/new_file.py` - Modify: `exact/path/to/existing.py:45-67` (line numbers if known) - Test: `tests/path/to/test_file.py` **Step 1: Write failing test**
def test_specific_behavior(): result = function(input) assert result == expected
**Step 2: Run test to verify failure**
Run: `pytest tests/path/test.py::test_specific_behavior -v`
Expected: FAIL — "function not defined"
**Step 3: Write minimal implementation**
def function(input): return expected
**Step 4: Run test to verify pass**
Run: `pytest tests/path/test.py::test_specific_behavior -v`
Expected: PASS
**Step 5: Commit**
git add tests/path/test.py src/path/file.py git commit -m "feat: add specific feature"
Read and understand:
Use Poirot tools to understand the project:
python# Understand project structure list_dir("src/") # Look at similar features bash("grep -rl 'similar_pattern' src/ --include='*.py'") # Check existing tests list_dir("tests/") # Read key files read_file("src/app.py")
Decide:
Create tasks in order:
For each task, include:
src/config/settings.py)Check:
Bad: Copy-paste validation in 3 places Good: Extract validation function, use everywhere
Bad: Add "flexibility" for future requirements Good: Implement only what's needed now
python# Bad — YAGNI violation class User: def __init__(self, name, email): self.name = name self.email = email self.preferences = {} # Not needed yet! self.metadata = {} # Not needed yet! # Good — YAGNI class User: def __init__(self, name, email): self.name = name self.email = email
Every task that produces code should include the full TDD cycle:
See test-driven-development skill for details.
Commit after every task:
bashgit add [files] git commit -m "type: description"
Bad: "Add authentication" Good: "Create User model with email and password_hash fields"
Bad: "Step 1: Add validation function" Good: "Step 1: Add validation function" followed by the complete function code
Bad: "Step 3: Test it works" Good: "Step 3: Run pytest tests/test_auth.py -v, expected: 3 passed"
Bad: "Create the model file" Good: "Create: src/models/user.py"
After saving the plan, offer the execution approach:
"Plan complete and saved. Ready to execute — I'll work through tasks sequentially with TDD. Shall I proceed?"
Note: Poirot currently has no subagent delegation. Execute tasks manually in order, committing after each. If subagent support is added later, this skill will reference it for parallel task execution.
Bite-sized tasks (2-5 min each)
Exact file paths
Complete code (copy-pasteable)
Exact commands with expected output
Verification steps
DRY, YAGNI, TDD
Frequent commitsA good plan makes implementation obvious.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,156 | 5,099 | -17% | 1 | 1 | 0% | 341 | 2,407 | +606% | 0 | 0 | — |
case-02 | fail→fail | 6,803 | 10,065 | +48% | 1 | 1 | 0% | 283 | 2,408 | +751% | 0 | 0 | — |
case-03 | fail→fail | 13,396 | 4,288 | -68% | 1 | 1 | 0% | 277 | 2,379 | +759% | 0 | 0 | — |
case-04 | fail→fail | 34,654 | 4,583 | -87% | 1 | 1 | 0% | 2,635 | 2,368 | -10% | 0 | 0 | — |
case-05 | pass→fail | 20,860 | 4,022 | -81% | 1 | 1 | 0% | 3,737 | 2,323 | -38% | 0 | 0 | — |
case-06 | fail→fail | 15,800 | 15,460 | -2% | 1 | 1 | 0% | 2,334 | 2,479 | +6% | 0 | 0 | — |
case-07 | fail→fail | 20,731 | 18,365 | -11% | 1 | 1 | 0% | 1,609 | 2,351 | +46% | 0 | 0 | — |
case-08 | pass→fail | 15,294 | 5,922 | -61% | 1 | 1 | 0% | 1,615 | 2,432 | +51% | 0 | 0 | — |
case-09 | pass→fail | 16,334 | 11,804 | -28% | 1 | 1 | 0% | 2,320 | 2,473 | +7% | 0 | 0 | — |
case-10 | pass→pass | 11,329 | 5,665 | -50% | 1 | 1 | 0% | 1,641 | 2,792 | +70% | 0 | 0 | — |
case-11 | fail→fail | 12,959 | 3,996 | -69% | 1 | 1 | 0% | 1,877 | 2,212 | +18% | 0 | 0 | — |
case-12 | pass→fail | 37,610 | 6,723 | -82% | 1 | 1 | 0% | 3,329 | 2,223 | -33% | 0 | 0 | — |
case-13 | pass→fail | 6,851 | 7,783 | +14% | 1 | 1 | 0% | 1,008 | 2,382 | +136% | 0 | 0 | — |
case-14 | pass→pass | 15,447 | 23,818 | +54% | 1 | 1 | 0% | 1,878 | 2,831 | +51% | 0 | 0 | — |
case-15 | fail→pass | 12,159 | 10,810 | -11% | 1 | 1 | 0% | 1,682 | 3,068 | +82% | 0 | 0 | — |
case-16 | pass→pass | 12,145 | 48,702 | +301% | 1 | 1 | 0% | 1,787 | 3,564 | +99% | 0 | 0 | — |
case-17 | pass→pass | 6,125 | 4,067 | -34% | 1 | 1 | 0% | 472 | 2,623 | +456% | 0 | 0 | — |
case-18 | pass→pass | 10,748 | 4,540 | -58% | 1 | 1 | 0% | 1,501 | 2,635 | +76% | 0 | 0 | — |
case-19 | pass→fail | 12,523 | 20,104 | +61% | 1 | 1 | 0% | 2,021 | 2,644 | +31% | 0 | 0 | — |
case-20 | pass→fail | 11,081 | 10,607 | -4% | 1 | 1 | 0% | 1,993 | 2,392 | +20% | 0 | 0 | — |
case-21 | fail→fail | 3,694 | 9,150 | +148% | 1 | 1 | 0% | 456 | 2,644 | +480% | 0 | 0 | — |
case-22 | fail→fail | 12,246 | 6,512 | -47% | 1 | 1 | 0% | 1,967 | 2,295 | +17% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 6 counted toward the lift figure. The other 16 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -27 percentage points is the difference between those two pass rates over the 6 comparable cases. 11 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.