Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Pre-commit cross-check — spawn an isolated, fresh-context critic to review the working diff for correctness, grounded in a real build/test run. The fast per-commit gate that feeds fewer issues into /code-review and /ship.
.claude/skills/nudgebee-judge/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 407% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-10 | ✓→✗ | ▼ Worse | -5% | 0% |
<!-- Single-source-of-truth skill: this file is canonical under .claude/skills/judge/. .gemini/skills/judge is a directory symlink to this directory (create it with: ln -s ../../.claude/skills/judge .gemini/skills/judge). Both agents parse the YAML frontmatter — Claude reads user-invocable + allowed-tools, Gemini reads name. Do NOT copy this skill into .gemini/skills/ — use the symlink. -->
Run an independent correctness check on the current working diff before committing. The critic runs in a fresh context — it did not write the code, so it evaluates the artifact, not the reasoning that produced it. This is self-review's blind spot removed.
The failure mode we defend against: committing code that looks right to the author (who is attached to it) but is wrong.
Relationship to your other gates: /judge is the fast, per-commit gate (1 critic, correctness only, in Phase 1). /code-review stays the deep, per-PR gate (all lenses). /challenge is the pre-plan gate. Don't collapse them — different granularity, same philosophy.
Use before committing any non-trivial code change — especially:
Skip for:
If the diff is trivial by the above, say so in one line and stop — do not spawn a critic for a typo.
Run git diff --stat (working tree; unstaged + staged). Decide:
Do not paste the whole diff here — the critic will read it itself in its own context. Keeping the diff out of this context is the point (that's the token saving).
Use the Task/Agent tool with subagent_type: correctness-critic. Give it a one-line scope prompt, e.g.:
> Review the current working diff for correctness. Build the affected module(s), run the changed packages' tests, read the changed lines for logic defects, and report grounded findings only. Working tree diff (not the branch). The task this change claims to do: <one line — what the user asked for>.
Pass the claimed intent so the critic can check "does it actually do that," not just "does it compile." Let the critic run its own build — do not run it here and hand results in; its independence is the value.
Fallback (Gemini / no subagent support): if the Task/Agent tool is unavailable, perform the correctness-critic procedure inline (build the affected module, run targeted tests, read the changed lines, ground every finding in evidence). Weaker — same context that may have written the code — but better than nothing. Note in the output that it ran in-loop, not isolated.
Present the critic's report to the user as-is (it's already formatted). Then:
/judge./judge reports; the human commits. It is a gate, not a pipeline step./code-review), not here.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | pass→pass | 18,091 | 11,317 | -37% | 1 | 1 | 0% | 1,660 | 1,866 | +12% | 0 | 0 | — |
case-01 | fail→fail | 11,952 | 13,769 | +15% | 1 | 1 | 0% | 206 | 1,428 | +593% | 0 | 0 | — |
case-02 | fail→fail | 17,971 | 12,203 | -32% | 1 | 1 | 0% | 320 | 1,537 | +380% | 0 | 0 | — |
case-03 | fail→fail | 13,188 | 11,168 | -15% | 1 | 1 | 0% | 337 | 1,289 | +282% | 0 | 0 | — |
case-04 | pass→pass | 24,688 | 6,092 | -75% | 1 | 1 | 0% | 2,960 | 2,033 | -31% | 0 | 0 | — |
case-21 | pass→pass | 13,655 | 4,953 | -64% | 1 | 1 | 0% | 1,920 | 1,823 | -5% | 0 | 0 | — |
case-05 | fail→pass | 14,900 | 8,851 | -41% | 1 | 1 | 0% | 1,932 | 1,611 | -17% | 0 | 0 | — |
case-06 | fail→fail | 10,017 | 12,261 | +22% | 1 | 1 | 0% | 620 | 1,290 | +108% | 0 | 0 | — |
case-07 | fail→pass | 10,130 | 11,694 | +15% | 1 | 1 | 0% | 413 | 2,095 | +407% | 0 | 0 | — |
case-08 | fail→fail | 17,216 | 22,455 | +30% | 1 | 1 | 0% | 212 | 1,393 | +557% | 0 | 0 | — |
case-09 | fail→fail | 15,227 | 20,110 | +32% | 1 | 1 | 0% | 490 | 1,718 | +251% | 0 | 0 | — |
case-22 | pass→pass | 19,919 | 12,487 | -37% | 1 | 1 | 0% | 2,035 | 2,070 | +2% | 0 | 0 | — |
case-10 | pass→fail | 12,869 | 9,960 | -23% | 1 | 1 | 0% | 1,814 | 1,715 | -5% | 0 | 0 | — |
case-11 | fail→pass | 18,912 | 10,239 | -46% | 1 | 1 | 0% | 1,994 | 1,719 | -14% | 0 | 0 | — |
case-12 | fail→fail | 17,004 | 13,433 | -21% | 1 | 1 | 0% | 1,897 | 2,340 | +23% | 0 | 0 | — |
case-13 | pass→pass | 13,434 | 3,639 | -73% | 1 | 1 | 0% | 1,198 | 1,507 | +26% | 0 | 0 | — |
case-14 | pass→pass | 16,084 | 13,829 | -14% | 1 | 1 | 0% | 1,340 | 1,540 | +15% | 0 | 0 | — |
case-16 | fail→fail | 6,176 | 7,749 | +25% | 1 | 1 | 0% | 740 | 1,383 | +87% | 0 | 0 | — |
case-17 | pass→pass | 15,152 | 10,198 | -33% | 1 | 1 | 0% | 1,368 | 1,756 | +28% | 0 | 0 | — |
case-18 | pass→pass | 14,716 | 4,666 | -68% | 1 | 1 | 0% | 1,229 | 1,650 | +34% | 0 | 0 | — |
case-19 | fail→pass | 19,856 | 9,443 | -52% | 1 | 1 | 0% | 2,256 | 2,357 | +4% | 0 | 0 | — |
case-20 | pass→pass | 15,247 | 10,635 | -30% | 1 | 1 | 0% | 1,225 | 1,814 | +48% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 15 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.