Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Conduct evidence-backed Happier code, plan-completeness, session, worktree, feature, commit, branch, PR, codebase, and release-readiness reviews with affected-corridor analysis, high-confidence finding triage, meaningful parallel lanes, proportionate but comprehensive QA, optional root-cause fixes, and independent closeout. Use for deep review, audit, QA, pre-merge assessment, plan-vs-implementation verification, review-and-fix loops, or when asked to inspect all related code rather than only ch
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 159% | 0% |
| case-04 | ✓→✗ | ▼ Worse | 162% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 83% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 302% | 0% |
| case-19 | ✓→✗ | ▼ Worse | -4% | 0% |
Use this as the only general review orchestrator for Happier. The superseded generic review-protocol and code-reviewer skills are archived and must not be invoked here; this skill absorbs their useful review categories and routes to the narrower repository skills that own testing, compatibility, diagnosis, and verification.
Determine five independent values before reviewing:
report, comment, qa, staged, or fix.advisory, boundary, or ship.Infer these from clear wording; ask only when the wrong target/base/mode would materially change the work or authorize writes the user did not request. Read targets.md for target inference and mixed dirty-worktree attribution. For plan targets or plan-backed reviews, identify whether the plan is a draft under pre-approval review or an approved execution contract, then read plan-completeness.md. Do not turn an implementation-completeness audit into a second plan-design review.
Mode rules:
report: review and QA only; recommendations, no source/config/test edits.comment: same as report, with concise PR-style comments backed by the full evidence record.qa: exercise the requested behavior and investigate failures without source edits or a general code-review mandate.staged: finish and present the reviewed finding/fix set, then wait for approval before source edits.fix: review, triage, implement authorized root-cause fixes, validate, and re-review.If the request contains several example prompts with different modes, treat them as examples. Follow the user's actual requested deliverable; do not merge contradictory example clauses into one execution mode.
Review-class rules:
advisory: inspect moving, dirty, partial, or completed work and report evidence-backed findings without a completeness or ship verdict;boundary: review a substantial integrated batch or plan gate and decide whether that boundary is ready to close;ship: issue a final security/data/schema/user-visible/release verdict using current deciding evidence and a different reviewer where required.An explicit user review request is sufficient reason to review current work. Do not wait for unrelated work to stop moving; reconcile materially changed observations before a boundary or ship verdict.
Never equate the diff with the whole review scope.
Classify observations as introduced, exposed/activated, pre-existing corridor debt required for coherence, or unrelated observation. Unrelated pre-existing issues do not enter the merge verdict. Read review-standard.md.
For architecture or cross-module relationship reviews, follow the repository Graphify instructions before raw exploration when its corpus is relevant.
Create a compact inventory before findings or QA:
Discovery is complete when those facts are sufficient to decide correctness and coverage. Do not stop early to save tokens, and do not keep retrieving optional confirmation after the material gaps are closed.
Run available deterministic inventory, schema, generated-output, formatting, type, and contract checks before spending independent semantic-review effort on the same facts. Their success is supporting evidence only; it never substitutes for behavior, architecture, security, compatibility, UX, or completeness judgment.
Always review:
Activate conditional scopes only when reachable:
Review categories are prompts to investigate, not automatic findings. Severity follows demonstrated impact, not category. Read review-standard.md for the integrated correctness, security, architecture, test, and maintainability standard.
Do not weaken QA in the name of proportionality. First enumerate the affected behavior and state space, then execute every material reachable scenario or give an evidence-backed disposition.
Across each independently observable user flow or canonical-owner contract, inventory these dimensions and exercise every materially reachable one:
For a materially user-visible corridor audit, exercise the primary user-facing flow end to end when runnable or record its evidence-backed disposition. One scenario may cover several dimensions; do not duplicate equivalent rows for every minor behavior or implementation step.
Every material row ends PASS, FAIL, BLOCKED, UNREACHABLE, OUT_OF_SCOPE with rationale, or an explicitly authorized DEFERRED. Avoid irrelevant Cartesian multiplication, not relevant edge cases. Add a bounded exploratory charter after scripted flows for user-visible changes. Read qa-coverage.md before any substantive QA.
Use subagents when independent, lane-sized work improves depth or throughput—not to perform parallelism theatrically.
fork_turns="none" in Codex).Read orchestration.md whenever delegation is used. Use skills/decompose-gates for hard lane boundaries. Check routine lane outputs with deciding scripts, tests, diffs, and focused source inspection; use skills/verify-claims for decision-material delegated claims consolidated at the applicable boundary.
Review changed behavior first, while reading as much unchanged corridor code as correctness requires. When a durable workspace exists, store bulky logs and inventories there; otherwise keep tool output bounded to summaries plus decisive excerpts without creating ad-hoc artifacts.
Review may begin against moving, dirty, partial, or completed work. Record a concise observed basis: applicable plan revision, current HEAD and dirty-state acknowledgement, relevant paths/symbols/flows, checks run, and runtime artifact when relevant. Advisory review can report defects, emerging split-brains, incomplete wiring, unsafe direction, or plan drift without claiming completeness. Boundary and ship verdicts reconcile only materially affected observations that changed during the review; never freeze, hash, lease, manifest, or globally snapshot the worktree for review orchestration.
Authors perform compact in-place self-review during implementation without creating a formal review program. Formal independent review is normally batched at substantial integrated boundaries, explicit user-requested review points, and high-risk triggers that cannot safely wait—not after every lane, commit, gate, or microchange.
Scale review depth with reachability, user impact, reversibility, and silence of failure—not with the number of mechanisms or files. Dormant code receives one architecture/activation assessment; do not exhaustively harden it as if it were shipping. Live behavior, schema/data, security, compatibility, release, and user-visible gates retain independent risk-appropriate review.
For every candidate finding:
Only high-confidence, objective, actionable issues become findings. Material uncertainty stays an investigation question with its falsifying check; immaterial uncertainty is discarded. Use review-standard.md for the finding and comment contracts.
Any decision-material finding, refutation, or "already fixed" claim inherited from a prior report, round, or external review is re-verified against the current relevant implementation before entering a verdict. Carried claims not re-verified remain labeled assumptions; use numeric verified-N-of-M accounting only when that census itself changes the decision. A green test asserting the defective behavior is wrong-contract evidence, not acceptance; rewrite it with an authorized fix rather than citing it as protection.
Reviewers may investigate and propose any evidence-backed correction, simplification, addition, removal, mechanism, process change, or plan amendment relevant to the target. Findings are candidate claims, not implementation orders or plan authority. The orchestrator adjudicates four dimensions separately: claim (CONFIRMED, REFUTED, INVESTIGATE), impact (MATERIAL, IMMATERIAL, UNRELATED), proposed response (ACCEPT, REPLACE_WITH_SIMPLER_FIX, DEFER, REJECT), and authority (WITHIN_PLAN, AMENDMENT_REQUIRED, OUT_OF_SCOPE). A mechanism-sized response still requires a reproduced failure, reachable risk, or named live consumer.
After accepted fixes, review the changed finding delta and affected corridor. Restart a full round only when the approved contract, architecture, scope, review boundary, or risk materially changed. If repeated rounds keep finding hazards created by the proposed mechanism, stop hardening it and run a deletion/simplification or redesign test.
In fix mode, or after approval in staged mode:
skills/happier-implement, including its bug-fix loop when the cause is not already established;skills/happier-compatibility for released/predecessor seams;skills/happier-diagnose only when a runtime/session/provider/auth incident requires its support evidence workflow;For approved plan-backed work, a finding does not authorize deviation from the plan. If the root-cause fix would materially change an approved requirement, mark it AMENDMENT_REQUIRED, document the evidence and smallest proposed amendment, and wait for user approval before implementing that deviation.
Never edit implementation files in report or comment mode.
Before a non-trivial boundary or ship verdict:
skills/attack-conclusion against alternative causes, neighboring cases, blast radius, environment gap, hypothesis lock, split-brains, and compatibility provenance;skills/verify-claims for decision-material delegated claims consolidated at this boundary;Route release-candidate sign-off to skills/happier-release-validation-review; do not recreate its release evidence protocol here.
Use no workspace unless the user requested durable tracking or the review is long-lived, multi-lane, plan-wide, substantively QA-heavy, or has evidence too bulky to manage safely in the handoff. When one of those conditions holds, bootstrap one isolated ignored workspace:
bashnode skills/happier-review/scripts/bootstrap-review.mjs --slug <short-slug> --target <target> --mode <mode> --class <advisory|boundary|ship> --path <relevant-path>
Read tracking-and-evidence.md and output-modes.md before creating or updating artifacts. Never put credentials in review artifacts; supplied development access stays runtime-only. Do not assume managed services hot-reload changes or restart/stop them without authorization—attest the actual loaded bundle/binary/revision when it matters.
A review is complete only when:
Use skills/handoff-report: outcome and blockers first, evidence-pointed findings next, coverage and rejected findings after, residual risk and exact next action last. Never claim “all flows,” “the full codebase,” or “the plan is complete” without an auditable coverage basis. A review may be final for its explicit target and observed basis while still naming unexamined surfaces and residual risk.
Other measured skills in the registry, with their headline benchmark lift.