Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review a proposed plan with a single-pass structured critique and a clear verdict. Use when the user asks to review a plan, stress-test a plan, critique a workflow, identify risks, or improve an execution sequence.
.claude/skills/chrisblattman-review-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 277% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 132% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 179% | 0% |
Read the plan, critique it hard, deliver a verdict. For plans of any substance the default is a small council of independent critics with a separate synthesis — three reviewers with different lenses catch more than one. quick or single-pass opts down to a single fresh-context review.
In priority order:
file:path argument. If the path doesn't exist: "File not found: path]. Check the path and try again."~/.claude/plans/ (Glob ~/.claude/plans/*.md), when the user just finished plan mode.If no plan is found anywhere, reply: "No plan found. Run /review-plan right after plan mode, or /review-plan file:path/to/plan.md." and stop.
If the plan is under ~50 words, warn — "This plan is very brief; the review may be limited." — then proceed.
quick or single-pass, or the plan is short and low-stakes (a simple to-do sequence).Assign an expert role for the critique — infer from the plan's content; the user's role:"..." always wins:
| The plan is about | Reviewer role | |-------------------|---------------| | Software, automation, AI tools | AI engineering specialist | | Grants, proposals, funders | Research funding strategist | | Papers, studies, data analysis | Research methodology specialist | | Project management, workflows, operations | Operations specialist | | Anything else | Strategic planning specialist |
Announce: "Reviewing as: meticulous role]. (Say role:\"...\" to override.)"
quick)Privacy check before any search: web search sends search terms outside Claude. If the plan includes confidential records, personal contact details, research-participant or human-subjects material, unpublished sensitive work, private financial/legal/medical/school details, or anything the user would not paste into a public website, skip web search and say: "I am skipping web search because the plan appears sensitive; reviewing from the local plan and general domain knowledge." If the plan is safe, run two web searches: "[the plan's main approach] best practices" and "[the plan's domain] common pitfalls". Distill 3–5 principles that bear on THIS plan — they feed the critique. If web search is unavailable, continue and note: "Web research unavailable; reviewing from domain knowledge."
Follow the bundled council skill — ${CLAUDE_PLUGIN_ROOT}/skills/council/SKILL.md — with:
--peer codex|gemini: pass it through if the user set it. The council skill checks whether that engine is actually available, asks for explicit send confirmation, and falls back to all-Claude with a friendly note if not. Privacy check first: a peer review sends the full plan to another AI service — if the plan contains confidential, personal, or research-participant material, don't pass --peer; tell the user why and run all-Claude.The council skill owns the persona-file check, the parallel dispatch, and the mandatory separate synthesis. Don't re-implement those here.
quick / single-pass / trivial plan only)Dispatch ONE Task call, subagent_type: general-purpose — a fresh context avoids grading your own homework — with this prompt:
You are a meticulous [role]. Find what's missing, what will break, and what's
wishful thinking. Do not rationalize or hedge.
PLAN TO REVIEW: [full plan text]
BEST PRACTICES CONTEXT: [principles from Step 3, or "none — quick mode"]
Review against 6 dimensions:
1. PRE-MORTEM — it's 3 months later and this failed. Top 3 causes?
2. COMPLETENESS — what's missing that a domain expert would expect?
3. FEASIBILITY — which steps depend on unconfirmed resources, approvals, or data?
4. BEST-PRACTICE FIT — where does the plan deviate, and is the deviation justified?
5. SEQUENCING — hidden blockers? Would reordering reduce risk?
6. SPECIFICITY — could someone unfamiliar execute each step? Flag hand-waves
("figure out", "as needed") and missing success criteria.
Classify each finding: Red (will likely cause failure or major rework), Yellow
(risky but survivable), Green (minor). End with
VERDICT: APPROVE | REVISE — [one-line reason]════════════════════════════════════════
PLAN REVIEW — [plan title]
════════════════════════════════════════
Reviewing as: meticulous [role] · Source: [file / plan mode / conversation]
STRENGTHS
1. [what the plan gets right — be specific]
WEAKNESSES & GAPS
🔴 [issue] → Fix: [specific change]
🟡 [issue] → Fix: [specific change]
🟢 [issue] → Fix: [specific change]
VERDICT
✅ APPROVE — [reason] or 🔄 REVISE — [reason] + revised plan belowWhen the council ran: take the verdict, blockers, and patches from the synthesis output; include the per-critic one-line verdicts; put the raw critiques in a <details> block at the bottom. If the verdict is REVISE, draft the revised plan yourself from the blockers and patches, marking changed sections [CHANGED] and additions [NEW].
Ask: "Apply these revisions, or give me feedback to refine further?"
/review-plan ← right after plan mode
/review-plan file:docs/migration-plan.md
/review-plan quick
/review-plan single-pass role:"clinical trial design specialist"
/review-plan --peer codex| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | fail→fail | 15,088 | 24,933 | +65% | 1 | 1 | 0% | 2,412 | 4,714 | +95% | 0 | 0 | — |
case-01 | fail→fail | 11,823 | 7,197 | -39% | 1 | 1 | 0% | 780 | 1,901 | +144% | 0 | 0 | — |
case-02 | pass→pass | 4,759 | 7,172 | +51% | 1 | 1 | 0% | 724 | 2,681 | +270% | 0 | 0 | — |
case-03 | fail→fail | 21,632 | 4,659 | -78% | 1 | 1 | 0% | 3,106 | 1,830 | -41% | 0 | 0 | — |
case-04 | pass→fail | 24,372 | 41,828 | +72% | 1 | 1 | 0% | 3,916 | 7,774 | +99% | 0 | 0 | — |
case-05 | fail→fail | 3,017 | 6,477 | +115% | 1 | 1 | 0% | 441 | 2,656 | +502% | 0 | 0 | — |
case-06 | pass→pass | 41,399 | 29,595 | -29% | 1 | 1 | 0% | 6,182 | 6,180 | -0% | 0 | 0 | — |
case-08 | pass→pass | 18,639 | 29,340 | +57% | 1 | 1 | 0% | 2,764 | 5,588 | +102% | 0 | 0 | — |
case-09 | fail→pass | 9,570 | 20,666 | +116% | 1 | 1 | 0% | 1,305 | 4,920 | +277% | 0 | 0 | — |
case-10 | fail→pass | 19,094 | 24,887 | +30% | 1 | 1 | 0% | 2,977 | 4,464 | +50% | 0 | 0 | — |
case-11 | fail→fail | 8,007 | 8,328 | +4% | 1 | 1 | 0% | 1,258 | 2,875 | +129% | 0 | 0 | — |
case-12 | pass→fail | 16,671 | 7,295 | -56% | 1 | 1 | 0% | 2,585 | 1,919 | -26% | 0 | 0 | — |
case-13 | fail→fail | 6,346 | 8,527 | +34% | 1 | 1 | 0% | 979 | 2,608 | +166% | 0 | 0 | — |
case-14 | fail→fail | 11,854 | 11,296 | -5% | 1 | 1 | 0% | 1,796 | 3,627 | +102% | 0 | 0 | — |
case-15 | fail→fail | 6,864 | 8,262 | +20% | 1 | 1 | 0% | 873 | 3,073 | +252% | 0 | 0 | — |
case-16 | fail→pass | 10,437 | 14,466 | +39% | 1 | 1 | 0% | 1,713 | 3,971 | +132% | 0 | 0 | — |
case-17 | fail→pass | 14,121 | 15,474 | +10% | 1 | 1 | 0% | 2,202 | 4,145 | +88% | 0 | 0 | — |
case-18 | pass→fail | 12,350 | 6,788 | -45% | 1 | 1 | 0% | 1,858 | 2,372 | +28% | 0 | 0 | — |
case-19 | fail→pass | 7,817 | 9,495 | +21% | 1 | 1 | 0% | 1,181 | 3,294 | +179% | 0 | 0 | — |
case-20 | fail→fail | 6,594 | 11,457 | +74% | 1 | 1 | 0% | 1,016 | 3,006 | +196% | 0 | 0 | — |
case-21 | pass→pass | 12,179 | 15,371 | +26% | 1 | 1 | 0% | 1,962 | 3,276 | +67% | 0 | 0 | — |
case-22 | pass→fail | 11,164 | 5,491 | -51% | 1 | 1 | 0% | 1,734 | 2,444 | +41% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +5 percentage points is the difference between those two pass rates over the 19 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.