Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Ultra-compressed code review comments. Cuts noise from PR feedback while preserving the actionable signal. Each comment is one line: location, problem, fix. Use when user says "review this PR", "code review", "review the diff", "/review", or invokes /caveman-review. Auto-triggers when reviewing pull requests.
.claude/skills/hoangnguyen0403-caveman-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 22% | 0% |
Write code review comments terse and actionable. One line per finding. Location, problem, fix. No throat-clearing.
Format: L<line>: <problem>. <fix>. — or <file>:L<line>: ... when reviewing multi-file diffs.
Severity prefix (optional, when mixed):
🔴 bug: — broken behavior, will cause incident🟡 risk: — works but fragile (race, missing null check, swallowed error)🔵 nit: — style, naming, micro-optim. Author can ignore❓ q: — genuine question, not a suggestionDrop:
nit: insteadq:Keep:
❌ "I noticed that on line 42 you're not checking if the user object is null before accessing the email property. This could potentially cause a crash if the user is not found in the database. You might want to add a null check here."
✅ L42: 🔴 bug: user can be null after .find(). Add guard before .email.
❌ "It looks like this function is doing a lot of things and might benefit from being broken up into smaller functions for readability."
✅ L88-140: 🔵 nit: 50-line fn does 4 things. Extract validate/normalize/persist.
❌ "Have you considered what happens if the API returns a 429? I think we should probably handle that case."
✅ L23: 🟡 risk: no retry on 429. Wrap in withBackoff(3).
Drop terse mode for: security findings (CVE-class bugs need full explanation + reference), architectural disagreements (need rationale, not just a one-liner), and onboarding contexts where the author is new and needs the "why". In those cases write a normal paragraph, then resume terse for the rest.
Reviews only — does not write the code fix, does not approve/request-changes, does not run linters. Output the comment(s) ready to paste into the PR. "stop caveman-review" or "normal mode": revert to verbose review style.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 5,315 | 2,749 | -48% | 1 | 1 | 0% | 919 | 1,034 | +13% | 0 | 0 | — |
case-11 | pass→pass | 5,686 | 5,301 | -7% | 1 | 1 | 0% | 992 | 1,122 | +13% | 0 | 0 | — |
case-01 | fail→pass | 4,670 | 2,256 | -52% | 1 | 1 | 0% | 783 | 948 | +21% | 0 | 0 | — |
case-03 | fail→pass | 5,025 | 2,516 | -50% | 1 | 1 | 0% | 963 | 1,088 | +13% | 0 | 0 | — |
case-04 | pass→pass | 4,567 | 2,530 | -45% | 1 | 1 | 0% | 873 | 1,007 | +15% | 0 | 0 | — |
case-05 | fail→pass | 2,839 | 2,697 | -5% | 1 | 1 | 0% | 510 | 1,108 | +117% | 0 | 0 | — |
case-06 | fail→fail | 6,320 | 3,012 | -52% | 1 | 1 | 0% | 947 | 1,082 | +14% | 0 | 0 | — |
case-07 | pass→pass | 3,335 | 4,081 | +22% | 1 | 1 | 0% | 659 | 1,371 | +108% | 0 | 0 | — |
case-08 | pass→pass | 5,078 | 3,571 | -30% | 1 | 1 | 0% | 880 | 1,204 | +37% | 0 | 0 | — |
case-09 | fail→pass | 5,348 | 3,651 | -32% | 1 | 1 | 0% | 1,085 | 1,322 | +22% | 0 | 0 | — |
case-10 | pass→pass | 13,170 | 10,608 | -19% | 1 | 1 | 0% | 696 | 1,636 | +135% | 0 | 0 | — |
case-12 | pass→pass | 5,889 | 4,243 | -28% | 1 | 1 | 0% | 1,092 | 1,292 | +18% | 0 | 0 | — |
case-13 | pass→pass | 5,362 | 3,802 | -29% | 1 | 1 | 0% | 985 | 1,247 | +27% | 0 | 0 | — |
case-14 | pass→pass | 5,517 | 2,846 | -48% | 1 | 1 | 0% | 946 | 1,145 | +21% | 0 | 0 | — |
case-15 | fail→pass | 6,215 | 2,586 | -58% | 1 | 1 | 0% | 1,078 | 1,136 | +5% | 0 | 0 | — |
case-20 | pass→fail | 3,499 | 2,174 | -38% | 1 | 1 | 0% | 566 | 917 | +62% | 0 | 0 | — |
case-16 | pass→pass | 5,777 | 3,701 | -36% | 1 | 1 | 0% | 923 | 1,223 | +33% | 0 | 0 | — |
case-17 | pass→pass | 4,485 | 3,649 | -19% | 1 | 1 | 0% | 702 | 1,216 | +73% | 0 | 0 | — |
case-18 | pass→pass | 6,561 | 2,827 | -57% | 1 | 1 | 0% | 1,167 | 1,042 | -11% | 0 | 0 | — |
case-19 | pass→fail | 6,245 | 3,330 | -47% | 1 | 1 | 0% | 1,135 | 1,245 | +10% | 0 | 0 | — |
case-21 | fail→fail | 2,236 | 2,516 | +13% | 1 | 1 | 0% | 318 | 1,029 | +224% | 0 | 0 | — |
case-22 | pass→pass | 7,336 | 2,660 | -64% | 1 | 1 | 0% | 1,329 | 1,085 | -18% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.