Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes — Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/PRD asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to "review since X".
.claude/skills/fradser-code-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 75% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 101% | 0% |
Two-axis review of the diff between HEAD and a fixed point the user supplies:
/mattpocock:bdd), those are the authoritative acceptance criteria.Both axes run as parallel sub-agents so they don't pollute each other's context, then this skill aggregates their findings.
The issue tracker should have been provided to you. If docs/agents/issue-tracker.md is missing, tell the user to run /mattpocock:setup-matt-pocock-skills.
Whatever the user said is the fixed point (a commit SHA, branch name, tag, main, HEAD~5, etc.). If they didn't specify one, ask for it.
Capture the diff command once: git diff <fixed-point>...HEAD (three-dot, so the comparison is against the merge-base). Also note the list of commits via git log <fixed-point>..HEAD --oneline.
Before going further, confirm the fixed point resolves (git rev-parse <fixed-point>) and the diff is non-empty. A bad ref or empty diff should fail here, not inside two parallel sub-agents.
Look for the originating spec, in this order:
#123, Closes #45, GitLab !67, etc.), fetched via the workflow in docs/agents/issue-tracker.md.docs/, specs/, or .scratch/ matching the branch name or feature.Anything in the repo that documents how code should be written, such as CODING_STANDARDS.md or CONTRIBUTING.md.
On top of whatever the repo documents, the Standards axis always carries the smell baseline below: a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:
Each smell reads what it is → how to fix; match it against the diff:
switch/if-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.a.b().c().d() navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.Send a single message with two Agent tool calls.
Standards sub-agent prompt should include:
Spec sub-agent prompt should include:
If the spec is missing, skip the Spec sub-agent and note this in the final report.
Present the two reports under ## Standards and ## Spec headings, verbatim or lightly cleaned. Do not merge or rerank findings, because the two axes are deliberately separate (see _Why two axes_).
End with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes: that's the reranking the separation exists to prevent.
Before scoring, load the project's known traps: Glob docs/memory/*.md, then read every file whose frontmatter category is pitfall or convention and whose summary is topically related to this review (the diff's files, patterns, or stack). These are evaluator-verified failure modes this repo already paid for.
In step 5.5's Standards scoring, treat a read memory file as a red flag: if the diff exhibits a known pitfall, record it as a FAIL in the Standards report with evidence citing the memory file (docs/memory/pitfall_<slug>.md). If docs/memory/ does not exist, skip this step.
After presenting the two reports, emit a binary verdict per axis:
Refute-before-PASS (red-team protocol): Before assigning PASS to either axis, attempt to refute your own PASS:
A PASS without an attempted refutation is invalid — downgrade to REWORK and list the un-refuted risk.
Final verdict: REWORK if either axis is REWORK; PASS only if both axes PASS (post-refutation).
A PASS without an attempted refutation is invalid. Before assigning PASS to either axis, state the strongest reason the work might still fail, cite the evidence (diff hunk or missing scenario) that would confirm it, and hold PASS only when concrete contrary evidence refutes it. Never merge or rerank the two axes — the separation exists to stop one axis masking the other.
A change can pass one axis and fail the other:
Reporting them separately stops one axis from masking the other.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 17,510 | 17,906 | +2% | 1 | 1 | 0% | 1,807 | 2,372 | +31% | 0 | 0 | — |
case-02 | fail→fail | 19,810 | 21,439 | +8% | 1 | 1 | 0% | 2,561 | 2,532 | -1% | 0 | 0 | — |
case-03 | fail→fail | 12,630 | 13,905 | +10% | 1 | 1 | 0% | 2,074 | 2,650 | +28% | 0 | 0 | — |
case-04 | pass→pass | 13,499 | 7,786 | -42% | 1 | 1 | 0% | 2,014 | 3,372 | +67% | 0 | 0 | — |
case-05 | pass→pass | 11,255 | 4,909 | -56% | 1 | 1 | 0% | 1,532 | 2,832 | +85% | 0 | 0 | — |
case-06 | fail→pass | 9,804 | 4,581 | -53% | 1 | 1 | 0% | 1,532 | 2,684 | +75% | 0 | 0 | — |
case-07 | fail→fail | 10,717 | 4,232 | -61% | 1 | 1 | 0% | 1,408 | 2,674 | +90% | 0 | 0 | — |
case-08 | pass→pass | 27,019 | 6,302 | -77% | 1 | 1 | 0% | 2,158 | 2,856 | +32% | 0 | 0 | — |
case-09 | pass→pass | 7,356 | 4,698 | -36% | 1 | 1 | 0% | 1,202 | 2,754 | +129% | 0 | 0 | — |
case-10 | pass→pass | 9,036 | 4,285 | -53% | 1 | 1 | 0% | 1,330 | 2,727 | +105% | 0 | 0 | — |
case-11 | pass→pass | 5,746 | 6,418 | +12% | 1 | 1 | 0% | 838 | 2,764 | +230% | 0 | 0 | — |
case-12 | pass→pass | 12,585 | 5,015 | -60% | 1 | 1 | 0% | 1,747 | 2,669 | +53% | 0 | 0 | — |
case-13 | fail→fail | 12,179 | 4,107 | -66% | 1 | 1 | 0% | 1,650 | 2,552 | +55% | 0 | 0 | — |
case-14 | fail→pass | 13,823 | 2,739 | -80% | 1 | 1 | 0% | 1,856 | 2,329 | +25% | 0 | 0 | — |
case-15 | fail→pass | 9,465 | 3,879 | -59% | 1 | 1 | 0% | 1,539 | 2,613 | +70% | 0 | 0 | — |
case-16 | fail→fail | 9,438 | 4,625 | -51% | 1 | 1 | 0% | 1,607 | 2,707 | +68% | 0 | 0 | — |
case-17 | fail→fail | 12,062 | 5,024 | -58% | 1 | 1 | 0% | 1,665 | 2,798 | +68% | 0 | 0 | — |
case-18 | pass→fail | 14,778 | 4,515 | -69% | 1 | 1 | 0% | 1,967 | 2,529 | +29% | 0 | 0 | — |
case-19 | fail→pass | 13,939 | 7,178 | -49% | 1 | 1 | 0% | 1,857 | 2,635 | +42% | 0 | 0 | — |
case-20 | fail→pass | 8,950 | 3,550 | -60% | 1 | 1 | 0% | 1,274 | 2,556 | +101% | 0 | 0 | — |
case-21 | pass→pass | 6,611 | 9,750 | +47% | 1 | 1 | 0% | 699 | 2,443 | +249% | 0 | 0 | — |
case-22 | fail→pass | 11,799 | 4,676 | -60% | 1 | 1 | 0% | 2,027 | 2,693 | +33% | 0 | 0 | — |
case-23 | pass→fail | 13,620 | 20,268 | +49% | 1 | 1 | 0% | 2,039 | 2,372 | +16% | 0 | 0 | — |
case-24 | pass→fail | 22,560 | 12,574 | -44% | 1 | 1 | 0% | 2,319 | 3,049 | +31% | 0 | 0 | — |
case-25 | pass→fail | 15,270 | 44,563 | +192% | 1 | 1 | 0% | 2,832 | 2,666 | -6% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 19 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +8 percentage points is the difference between those two pass rates over the 19 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/29/2026 | +5% |
Other measured skills in the registry, with their headline benchmark lift.