Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Before ANY significant development task (new feature, refactor, integration, migration), run a complete planning ritual - clarify scope, assess risks, estimate time, get confirmation. Do not start coding until the battle plan is approved. Use when the user is about to start non-trivial work.
.claude/skills/adityapeshave-sophos-battle-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 590% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 33% | 0% |
A basic plan mode is just "read-only mode + write a markdown file." It doesn't assess risks, estimate time, define done, or check scope. This skill orchestrates a complete planning ritual using composable skills before any significant work begins.
Trigger before:
Do NOT trigger for:
Before anything, clarify what we're actually building:
Scope Check
What are we building?
-> [Clear description of the goal]
What's the definition of done?
-> [Specific, verifiable completion criteria]
What's explicitly OUT of scope?
-> [Things we're NOT doing, even if related]
Any ambiguity I should clarify before planning?If scope is vague, ask clarifying questions. Do not proceed with vague scope.
Map the terrain before planning the attack:
Recon
Files likely affected:
- [file1.py] - [why]
- [file2.py] - [why]
Existing patterns to follow:
- [Pattern observed in codebase]
Dependencies/callers to check:
- [What depends on code we're changing]
Tests that exist:
- [Relevant test files]Use search and file reading to actually explore. Don't guess.
Imagine it's 2 weeks later and this failed. Why?
Pre-mortem
If this fails, it's probably because:
1. HIGH: [Specific risk]
Mitigation: [How to prevent]
2. MED: [Specific risk]
Mitigation: [How to prevent]
3. LOW: [Specific risk]
Mitigation: [How to prevent]Reorder plan to address HIGH risks first.
Provide realistic time estimate:
Estimate
Scope: [X files, Y estimated lines changed]
Complexity: [Simple / Medium / Complex]
Breakdown:
- Recon & setup: ~X min
- Implementation: ~Y min
- Testing: ~Z min
- Verification: ~W min
Total: ~[range] minutes
Checkpoint: [If >15 min, where's the halfway check-in?]Now write the actual plan - structured, not prose:
Battle Plan: [Task Name]
SCOPE
Goal: [One sentence]
Done when: [Verifiable criteria]
Not doing: [Explicit exclusions]
STEPS
1. [ ] [Step] -> [Output/artifact]
Confidence: HIGH/MED/LOW
2. [ ] [Step] -> [Output/artifact]
Confidence: HIGH/MED/LOW
Depends on: Step 1
3. [ ] [Step] -> [Output/artifact]
Confidence: HIGH/MED/LOW
RISKS & MITIGATIONS
[Risk 1] -> [Mitigation baked into steps]
[Risk 2] -> [Mitigation baked into steps]
ROLLBACK
If this goes wrong: [How to undo safely]
ESTIMATE
~[X-Y] minutes
Checkpoint at: [Step N / halfway point]
Ready to execute?Do not proceed without explicit approval:
Type 'go' to execute this plan, or tell me what to adjust.Accept only:
For each step, assess confidence:
HIGH - Done this before, clear path, low ambiguity MED - Reasonable approach, some unknowns LOW - Best guess, might need to pivot
If multiple steps are LOW confidence, flag it:
Several steps have LOW confidence. Consider spiking/prototyping first?This skill orchestrates patterns from:
If those skills are installed, their full capabilities enhance this ritual. If not, this skill includes the core patterns inline.
User: "Add Stripe payment integration"
Scope Check
What are we building?
-> Accept one-time payments via Stripe Checkout
Definition of done:
-> User can click "Pay", complete Stripe Checkout, and see success confirmation
-> Payment recorded in our database
-> Works in test mode with test cards
Explicitly OUT of scope:
-> Subscriptions (future task)
-> Refunds (future task)
-> Multiple payment methods (just card for now)
Recon
Files likely affected:
- routes/payments.py - new endpoints
- models/order.py - add payment_status field
- templates/checkout.html - payment button
- tests/test_payments.py - new file
Existing patterns:
- Routes use Flask blueprints
- Models use SQLAlchemy with migrations
- Tests use pytest + fixtures
Pre-mortem
If this fails in 2 weeks, it's probably because:
1. HIGH: Webhook signature validation wrong
Mitigation: Set up Stripe CLI, test webhooks FIRST
2. HIGH: Test vs live key confusion
Mitigation: Env validation on startup
3. MED: No idempotency on payment creation
Mitigation: Add idempotency key from start
Estimate
Scope: 4 files, ~200 lines
Complexity: Medium
Total: ~35-45 minutes
Checkpoint: After webhook setup works (~15 min mark)
Battle Plan: Stripe Payment Integration
SCOPE
Goal: One-time payments via Stripe Checkout
Done when: Test payment completes end-to-end
Not doing: Subscriptions, refunds, multi-method
STEPS
1. [ ] Set up Stripe CLI + webhook endpoint -> Working signature validation
Confidence: HIGH
2. [ ] Add payment routes (create session, success, webhook) -> Endpoints responding
Confidence: HIGH
Depends on: Step 1
3. [ ] Database migration for payment_status -> Migration applied
Confidence: HIGH
4. [ ] Frontend checkout button -> Redirects to Stripe
Confidence: MED (need to check existing JS patterns)
Depends on: Step 2
5. [ ] Integration tests with Stripe test mode -> Tests passing
Confidence: HIGH
Depends on: Steps 1-4
RISKS & MITIGATIONS
Webhook validation -> Addressed in Step 1 (before business logic)
Key confusion -> Env check added in Step 2
Idempotency -> Built into Step 2
ROLLBACK
If broken: Revert migration, remove routes, disable checkout button
ESTIMATE
~35-45 minutes
Checkpoint at: Step 1 complete (webhook working)
Ready to execute?| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 18,448 | 3,804 | -79% | 1 | 1 | 0% | 3,102 | 2,046 | -34% | 0 | 0 | — |
case-02 | fail→pass | 19,389 | 22,455 | +16% | 1 | 1 | 0% | 3,335 | 5,222 | +57% | 0 | 0 | — |
case-03 | fail→pass | 27,197 | 20,108 | -26% | 1 | 1 | 0% | 4,277 | 4,631 | +8% | 0 | 0 | — |
case-04 | pass→fail | 3,567 | 1,771 | -50% | 1 | 1 | 0% | 517 | 1,987 | +284% | 0 | 0 | — |
case-05 | pass→pass | 24,353 | 26,395 | +8% | 1 | 1 | 0% | 3,156 | 4,442 | +41% | 0 | 0 | — |
case-06 | fail→pass | 9,365 | 3,100 | -67% | 1 | 1 | 0% | 319 | 2,201 | +590% | 0 | 0 | — |
case-07 | fail→pass | 17,582 | 11,881 | -32% | 1 | 1 | 0% | 3,098 | 3,639 | +17% | 0 | 0 | — |
case-08 | fail→pass | 18,930 | 22,920 | +21% | 1 | 1 | 0% | 3,330 | 4,444 | +33% | 0 | 0 | — |
case-09 | fail→pass | 17,713 | 17,601 | -1% | 1 | 1 | 0% | 3,118 | 4,771 | +53% | 0 | 0 | — |
case-10 | fail→pass | 20,125 | 16,430 | -18% | 1 | 1 | 0% | 3,896 | 4,503 | +16% | 0 | 0 | — |
case-11 | fail→fail | 31,136 | 12,519 | -60% | 1 | 1 | 0% | 2,708 | 3,651 | +35% | 0 | 0 | — |
case-12 | fail→fail | 26,533 | 17,256 | -35% | 1 | 1 | 0% | 3,305 | 4,562 | +38% | 0 | 0 | — |
case-13 | fail→pass | 14,210 | 13,478 | -5% | 1 | 1 | 0% | 2,477 | 4,192 | +69% | 0 | 0 | — |
case-14 | fail→pass | 16,671 | 13,598 | -18% | 1 | 1 | 0% | 2,929 | 4,041 | +38% | 0 | 0 | — |
case-15 | fail→pass | 19,690 | 21,819 | +11% | 1 | 1 | 0% | 3,520 | 5,591 | +59% | 0 | 0 | — |
case-16 | fail→pass | 15,112 | 11,161 | -26% | 1 | 1 | 0% | 2,722 | 3,664 | +35% | 0 | 0 | — |
case-17 | fail→pass | 18,529 | 14,169 | -24% | 1 | 1 | 0% | 2,653 | 3,732 | +41% | 0 | 0 | — |
case-18 | fail→pass | 39,497 | 14,749 | -63% | 1 | 1 | 0% | 4,228 | 4,237 | +0% | 0 | 0 | — |
case-19 | fail→pass | 22,555 | 13,609 | -40% | 1 | 1 | 0% | 3,614 | 4,219 | +17% | 0 | 0 | — |
case-20 | fail→pass | 20,970 | 14,741 | -30% | 1 | 1 | 0% | 3,503 | 4,080 | +16% | 0 | 0 | — |
case-21 | fail→pass | 16,200 | 12,193 | -25% | 1 | 1 | 0% | 2,819 | 3,538 | +26% | 0 | 0 | — |
case-22 | pass→fail | 13,509 | 4,007 | -70% | 1 | 1 | 0% | 2,271 | 1,972 | -13% | 0 | 0 | — |
case-23 | fail→pass | 20,243 | 12,543 | -38% | 1 | 1 | 0% | 3,601 | 3,988 | +11% | 0 | 0 | — |
case-24 | fail→pass | 14,877 | 10,762 | -28% | 1 | 1 | 0% | 2,919 | 3,613 | +24% | 0 | 0 | — |
case-25 | fail→pass | 16,742 | 16,509 | -1% | 1 | 1 | 0% | 2,898 | 3,482 | +20% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 23 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +68 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.