Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Parallel 4-agent cleanup of recent code changes.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | 228% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 130% | 0% |
| case-05 | ✓→✗ | ▼ Worse | 147% | 0% |
| case-11 | ✓→✗ | ▼ Worse | 94% | 0% |
| case-12 | ✓→✓ | = Same ✓ | 172% | 0% |
Review your recent code changes with four focused reviewers running in parallel, aggregate their findings, and apply the fixes worth applying.
This is a cleanup pass, not a bug hunt. You are improving the quality of code that already works — removing duplication, flattening needless complexity, cutting waste, and deepening band-aid fixes. Do not go hunting for correctness bugs here; that's what requesting-code-review is for.
Core principle: Four narrow reviewers beat one broad reviewer. Each one deeply searches the codebase for a single class of problem — reuse, quality, efficiency, altitude — without diluting its attention across all four. They run concurrently, so you pay the latency of one review, not four.
Trigger this skill when the user says any of:
Optional modifiers the user may add — honor them:
(or weight the aggregation toward it). Recognized focuses: reuse, quality (also accepts simplification), efficiency, altitude.
four reviewers, present findings, apply NOTHING. Ask before applying.
src/foo.py" → narrow the diff source accordingly (see Phase 1).
Do NOT auto-run this after every edit or tack it onto the end of unrelated tasks. It costs four subagents' worth of tokens — invoke it only when the user explicitly asks.
Capture the diff to review. Pick the source by what the user asked for, in this default order:
bash# 1. Default: uncommitted working-tree changes (tracked files) git diff # 2. If that's empty, include staged changes git diff HEAD # 3. Scoped variants the user may request: git diff --staged # "staged changes" git diff HEAD~1 # "the last commit" git diff main...HEAD # "this branch" / "my PR" git diff -- src/foo.py # specific file(s)
If git diff and git diff HEAD are both empty and there's no git repo or no changes, fall back to the files the user explicitly named or that were recently created/edited in this session. If you genuinely can't find any changed code, say so and stop — there's nothing to simplify.
Capture the full diff text. Note its size: if it's very large (say >2000 changed lines), warn the user that four subagents each carrying the full diff will be token-heavy, and offer to scope it down (per-directory, per-commit) before proceeding.
Use delegate_task batch mode — pass all four tasks in one tasks array so they run concurrently. Four is the right fan-out for this pattern; it's within the delegation.max_concurrent_children budget on any default install.
No delegation available? If you can't call delegate_task in this context (you're a leaf subagent, delegation is disabled, or the budget is exhausted), do NOT skip the review or drop angles. Work through all four reviewer angles yourself, sequentially, in this context — same search standards, same finding format. Then say clearly in your final summary that this was a single-pass inline review, not the parallel fan-out, so the user knows what actually ran.
Give every reviewer the complete diff (not fragments — cross-file issues hide in the gaps) plus the absolute repo path so they can search the wider codebase. Each reviewer gets terminal, file, and search toolsets (so they can git, read_file, and search_files/grep).
Tell each reviewer to:
git blame on the line to understand why it exists. If you can't determine the original purpose, mark it confidence: low — don't guess.
and risk: file:line → problem → cost (what's duplicated/wasted/harder to maintain) → suggested fix | confidence: high/medium/low | risk: SAFE/CAREFUL/RISKY The cost field forces each finding to justify itself — a finding that can't articulate what the problem actually costs is probably a nit.
code, pass-through wrappers). Auto-apply these.
flatten nested ternary, extract helper). Apply with test verification.
restructuring, public API rename, memory lifecycle change). Flag for human review — do NOT auto-apply.
the code.
Pass these four goals (drop any the user's focus excludes):
Reviewer 1 — Code Reuse > Review this diff for code that duplicates functionality already in the > codebase. Search utility modules, shared helpers, and adjacent files > (use search_files / grep) for existing functions, constants, or patterns > the new code could call instead of reimplementing. Flag: new functions > that duplicate existing ones; hand-rolled logic that an existing utility > already does (manual string/path manipulation, custom env checks, ad-hoc > type guards, re-implemented parsing). For each, name the existing thing to > use and where it lives.
Reviewer 2 — Code Quality > Review this diff for quality problems. Look for: redundant state (values > that duplicate or could be derived from existing state; caches that don't > need to exist); parameter sprawl (new params bolted on where the function > should have been restructured); copy-paste-with-variation (near-duplicate > blocks that should share an abstraction); leaky abstractions (exposing > internals, breaking an existing encapsulation boundary); stringly-typed > code (raw strings where a constant/enum/registry already exists — check the > canonical registries before flagging); deeply nested conditionals (ternary > chains, 3+-level if/else pyramids — flatten with guard clauses, early > returns, or a lookup table); AI-generated slop patterns (extra > comments restating obvious code like // increment counter above count++; > unnecessary defensive null-checks on already-validated inputs; as any > casts that bypass the type system; patterns inconsistent with the rest of > the file). For each, give the concrete refactor.
Reviewer 3 — Efficiency > Review this diff for efficiency problems. Look for: unnecessary work > (redundant computation, repeated file reads, duplicate API calls, N+1 > access patterns); missed concurrency (independent ops run sequentially); > hot-path bloat (heavy/blocking work on startup or per-request paths); > TOCTOU anti-patterns (existence pre-checks before an op instead of doing > the op and handling the error); memory issues (unbounded growth, missing > cleanup, listener/handle leaks; long-lived callbacks or objects built as > closures that capture the whole enclosing scope — everything captured > stays alive as long as the object does, so prefer a small class or > explicit-fields struct that copies only what it needs); overly broad reads > (loading whole files when a slice would do); silent failures (empty catch > blocks, ignored error returns, except: pass, .catch(() => {}) with no > handling, error propagation gaps — these hide bugs and should at minimum > log before swallowing). For each, give the concrete fix and why it's > faster or safer.
Reviewer 4 — Altitude > Review this diff for changes implemented at the wrong depth — band-aids > layered on top of shared infrastructure instead of fixes to the > infrastructure itself. Signs of a too-shallow fix: a special case added to > a generic code path to handle one caller (an if (caller == X) branch, a > type check, a magic-value escape hatch); a symptom patched at the call > site while sibling call sites keep the same flaw; a workaround stacked on > an earlier workaround; a wrapper added to avoid touching the thing that > actually needs changing; configuration or flags introduced to route around > a broken default instead of fixing the default. For each, identify the > underlying mechanism the change is dodging and describe the deeper fix — > generalize the shared path, fix the root default, or fix the whole bug > class — and honestly note when the deeper fix is large enough that it > should be its own task rather than part of this cleanup. Read the > surrounding code and git blame first: what looks like a band-aid is > sometimes a deliberate boundary (compat shims, staged migrations, > vendored-code isolation). Don't flag those.
Wait for all four to return (batch mode returns them together).
when two findings target the same line or the same underlying mechanism, collapse them into one.
argue with a reviewer, just drop weak or wrong suggestions silently.
util X"; Reviewer 3: "X is slow, inline it"). Default resolution order: correctness > the user's stated focus > readability/reuse > micro-perf. Don't apply a perf "fix" that hurts clarity unless the path is genuinely hot. When two suggestions are mutually exclusive and both defensible, pick the one that touches less code and note the alternative.
pass-through wrappers, redundant type assertions. Run tests after.
locals, flatten ternaries, extract helpers, consolidate dupes. Run tests after each file. Revert any that break.
public API changes, concurrency fixes, error-handling changes. Present each with risk description and test coverage status. Altitude findings usually land here — deepening a fix means touching shared infrastructure, so present the deeper fix and let the user decide whether to do it now or as a follow-up. If the user opted for a dry run, present all three tiers and apply nothing.
the touched files (not the full suite), and re-run any linter/type check the repo uses. If a fix breaks a test, revert that one fix and report it.
reviewer category and risk tier, plus any findings you deliberately skipped and why. If you ran inline (no delegation), say so here.
conflicting suggestions to reconcile, not better coverage. The four categories cover the space.
defeats the design — cross-file duplication and N+1s only show up with the full picture.
the existing utility ("there's probably a helper for this") is noise. Require file:line evidence; drop findings that lack it.
license to refactor the whole module. Keep edits scoped to what the diff touched plus the minimal surrounding change a fix requires. Altitude findings are the exception that proves the rule: when the right fix is deeper than the diff, FLAG it — don't unilaterally rebuild the shared mechanism inside a cleanup pass.
correctness bug, report it prominently — but as a separate "found a bug" note, not folded into cleanup fixes. Correctness review is a different pass with different verification standards.
HERMES.md or a linter config, fold those rules into the reviewer prompts so suggestions match house style instead of fighting it.
delegating — four subagents each carrying a 5000-line diff is expensive and may truncate.
knip, ts-prune, and depcheck flagexports that ARE used dynamically (string-based imports, reflection). Always grep for the symbol name before removing — a clean tool report is not proof.
paths, DB column names, and config keys are contracts — even if the name is bad, renaming breaks consumers. Tag public-contract changes as RISKY; never auto-rename them.
error might be intentional — the error is expected and benign in that context. Flag it, don't remove it; let the human decide.
and isolation layers around vendored code look like altitude violations but are deliberate design. Check git blame and surrounding comments before flagging; when the intent is unclear, mark confidence: low.
If your install has the subagent-driven-development skill (optional), it covers the complementary case: parallel review during implementation, per task. This skill is the standalone after-the-fact cleanup pass. Use requesting-code-review for the pre-commit security/quality gate — that's the bug hunt; this is the cleanup.
Other measured skills in the registry, with their headline benchmark lift.