Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Write a markdown plan to .hermes/plans/; no execution.
.claude/skills/nousresearch-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-20 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-05 | ✓→✗ | ▼ Worse | 18% | 0% |
| case-09 | ✓→✗ | ▼ Worse | 27% | 0% |
| case-10 | ✓→✗ | ▼ Worse | 4% | 0% |
| case-15 | ✓→✗ | ▼ Worse | 25% | 0% |
Use this skill when the user wants a plan instead of execution.
For this turn, you are planning only.
.hermes/plans/.Write a markdown plan that is concrete and actionable.
Include, when relevant:
If the task is code-related, include exact file paths, likely test targets, and verification steps.
Save the plan with write_file under:
.hermes/plans/YYYY-MM-DD_HHMMSS-<slug>.mdTreat that as relative to the active working directory / backend workspace. Hermes file tools are backend-aware, so using this relative path keeps the plan with the workspace on local, docker, ssh, modal, and daytona backends.
If the runtime provides a specific target path, use that exact path. If not, create a sensible timestamped filename yourself under .hermes/plans/.
/plan, infer the task from the current conversation context.The rest of this skill is the craft of authoring a good implementation plan — the content that goes inside the markdown file above.
Write comprehensive implementation plans assuming the implementer has zero context for the codebase and questionable taste. Document everything they need: which files to touch, complete code, testing commands, docs to check, how to verify. Give them bite-sized tasks. DRY. YAGNI. TDD. Frequent commits.
Assume the implementer is a skilled developer but knows almost nothing about the toolset or problem domain. Assume they don't know good test design very well.
Core principle: A good plan makes implementation obvious. If someone has to guess, the plan is incomplete.
Always use before:
Don't skip when:
Each task = 2-5 minutes of focused work.
Every step is one action:
Too big:
markdown### Task 1: Build authentication system [50 lines of code across 5 files]
Right size:
markdown### Task 1: Create User model with email field [10 lines, 1 file] ### Task 2: Add password hash field to User [8 lines, 1 file] ### Task 3: Create password hashing utility [15 lines, 1 file]
Every plan MUST start with:
markdown# [Feature Name] Implementation Plan > **For Hermes:** Use subagent-driven-development skill to implement this plan task-by-task. **Goal:** [One sentence describing what this builds] **Architecture:** [2-3 sentences about approach] **Tech Stack:** [Key technologies/libraries] ---
Each task follows this format:
`markdown### Task N: [Descriptive Name] **Objective:** What this task accomplishes (one sentence) **Files:** - Create: `exact/path/to/new_file.py` - Modify: `exact/path/to/existing.py:45-67` (line numbers if known) - Test: `tests/path/to/test_file.py` **Step 1: Write failing test**
def test_specific_behavior(): result = function(input) assert result == expected
**Step 2: Run test to verify failure**
Run: `pytest tests/path/test.py::test_specific_behavior -v`
Expected: FAIL — "function not defined"
**Step 3: Write minimal implementation**
def function(input): return expected
**Step 4: Run test to verify pass**
Run: `pytest tests/path/test.py::test_specific_behavior -v`
Expected: PASS
**Step 5: Commit**
git add tests/path/test.py src/path/file.py git commit -m "feat: add specific feature"
Read and understand:
Use Hermes tools to understand the project:
python# Understand project structure search_files("*.py", target="files", path="src/") # Look at similar features search_files("similar_pattern", path="src/", file_glob="*.py") # Check existing tests search_files("*.py", target="files", path="tests/") # Read key files read_file("src/app.py")
Decide:
Create tasks in order:
For each task, include:
src/config/settings.py)Check:
Bad: Copy-paste validation in 3 places Good: Extract validation function, use everywhere
Bad: Add "flexibility" for future requirements Good: Implement only what's needed now
python# Bad — YAGNI violation class User: def __init__(self, name, email): self.name = name self.email = email self.preferences = {} # Not needed yet! self.metadata = {} # Not needed yet! # Good — YAGNI class User: def __init__(self, name, email): self.name = name self.email = email
Every task that produces code should include the full TDD cycle:
See test-driven-development skill for details.
Commit after every task:
bashgit add [files] git commit -m "type: description"
Bad: "Add authentication" Good: "Create User model with email and password_hash fields"
Bad: "Step 1: Add validation function" Good: "Step 1: Add validation function" followed by the complete function code
Bad: "Step 3: Test it works" Good: "Step 3: Run pytest tests/test_auth.py -v, expected: 3 passed"
Bad: "Create the model file" Good: "Create: src/models/user.py"
After saving the plan, offer the execution approach:
"Plan complete and saved. Ready to execute using subagent-driven-development — I'll dispatch a fresh subagent per task with two-stage review (spec compliance then code quality). Shall I proceed?"
When executing, use the subagent-driven-development skill:
delegate_task per task with full contextBite-sized tasks (2-5 min each)
Exact file paths
Complete code (copy-pasteable)
Exact commands with expected output
Verification steps
DRY, YAGNI, TDD
Frequent commitsA good plan makes implementation obvious.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 3,683 | 2,935 | -20% | 1 | 1 | 0% | 289 | 2,432 | +742% | 0 | 0 | — |
case-01 | fail→fail | 5,562 | 4,944 | -11% | 1 | 1 | 0% | 326 | 2,535 | +678% | 0 | 0 | — |
case-02 | fail→fail | 4,023 | 3,601 | -10% | 1 | 1 | 0% | 309 | 2,421 | +683% | 0 | 0 | — |
case-04 | fail→fail | 2,901 | 4,348 | +50% | 1 | 1 | 0% | 153 | 2,523 | +1549% | 0 | 0 | — |
case-05 | pass→fail | 10,828 | 3,740 | -65% | 1 | 1 | 0% | 2,129 | 2,507 | +18% | 0 | 0 | — |
case-06 | fail→fail | 2,489 | 5,247 | +111% | 1 | 1 | 0% | 273 | 2,479 | +808% | 0 | 0 | — |
case-07 | fail→fail | 2,640 | 2,313 | -12% | 1 | 1 | 0% | 141 | 2,346 | +1564% | 0 | 0 | — |
case-08 | fail→fail | 4,613 | 4,111 | -11% | 1 | 1 | 0% | 293 | 2,382 | +713% | 0 | 0 | — |
case-09 | pass→fail | 10,850 | 3,574 | -67% | 1 | 1 | 0% | 1,941 | 2,464 | +27% | 0 | 0 | — |
case-10 | pass→fail | 14,393 | 3,502 | -76% | 1 | 1 | 0% | 2,368 | 2,457 | +4% | 0 | 0 | — |
case-11 | fail→fail | 11,931 | 5,003 | -58% | 1 | 1 | 0% | 2,025 | 2,539 | +25% | 0 | 0 | — |
case-12 | pass→pass | 6,566 | 4,359 | -34% | 1 | 1 | 0% | 1,040 | 2,940 | +183% | 0 | 0 | — |
case-13 | fail→fail | 7,537 | 3,070 | -59% | 1 | 1 | 0% | 1,264 | 2,340 | +85% | 0 | 0 | — |
case-14 | pass→pass | 7,227 | 4,416 | -39% | 1 | 1 | 0% | 1,225 | 3,080 | +151% | 0 | 0 | — |
case-15 | pass→fail | 10,010 | 4,324 | -57% | 1 | 1 | 0% | 1,983 | 2,470 | +25% | 0 | 0 | — |
case-16 | pass→fail | 11,689 | 5,275 | -55% | 1 | 1 | 0% | 2,152 | 2,617 | +22% | 0 | 0 | — |
case-17 | pass→pass | 8,372 | 16,479 | +97% | 1 | 1 | 0% | 1,356 | 3,713 | +174% | 0 | 0 | — |
case-18 | pass→fail | 10,299 | 12,743 | +24% | 1 | 1 | 0% | 1,753 | 3,316 | +89% | 0 | 0 | — |
case-19 | fail→fail | 10,688 | 6,997 | -35% | 1 | 1 | 0% | 1,759 | 2,601 | +48% | 0 | 0 | — |
case-20 | fail→pass | 7,497 | 2,248 | -70% | 1 | 1 | 0% | 1,223 | 2,530 | +107% | 0 | 0 | — |
case-21 | pass→pass | 8,190 | 5,050 | -38% | 1 | 1 | 0% | 1,510 | 3,097 | +105% | 0 | 0 | — |
case-22 | pass→fail | 5,920 | 5,320 | -10% | 1 | 1 | 0% | 1,138 | 2,619 | +130% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 5 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -27 percentage points is the difference between those two pass rates over the 5 comparable cases. 7 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.