Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when receiving code review feedback, before implementing suggestions, especially if feedback seems unclear or technically questionable - requires technical rigor and verification, not performative agreement or blind implementation
.claude/skills/mkurman-receiving-code-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-02 | ✓→✓ | = Same ✓ | -10% | 0% |
| case-03 | ✓→✓ | = Same ✓ | -31% | 0% |
| case-04 | ✓→✓ | = Same ✓ | -25% | 0% |
---|-----| | Performative agreement | State requirement or just act | | Blind implementation | Verify against codebase first | | Batch without testing | One at a time, test each | | Assuming reviewer is right | Check if breaks things | | Avoiding pushback | Technical correctness > comfort | | Partial implementation | Clarify all items first | | Can't verify, proceed anyway | State limitation, ask for direction |
Performative Agreement (Bad):
Reviewer: "Remove legacy code"
❌ "You're absolutely right! Let me remove that..."Technical Verification (Good):
Reviewer: "Remove legacy code"
✅ "Checking... build target is 10.15+, this API needs 13+. Need legacy for backward compat. Current impl has wrong bundle ID - fix it or drop pre-13 support?"YAGNI (Good):
Reviewer: "Implement proper metrics tracking with database, date filters, CSV export"
✅ "Grepped codebase - nothing calls this endpoint. Remove it (YAGNI)? Or is there usage I'm missing?"Unclear Item (Good):
your human partner: "Fix items 1-6"
You understand 1,2,3,6. Unclear on 4,5.
✅ "Understand 1,2,3,6. Need clarification on 4 and 5 before implementing."When replying to inline review comments on GitHub, reply in the comment thread (gh api repos/{owner}/{repo}/pulls/{pr}/comments/{id}/replies), not as a top-level PR comment.
External feedback = suggestions to evaluate, not orders to follow.
Verify. Question. Then implement.
No performative agreement. Technical rigor always.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 8,158 | 3,627 | -56% | 1 | 1 | 0% | 1,162 | 956 | -18% | 0 | 0 | — |
case-02 | pass→pass | 5,022 | 2,580 | -49% | 1 | 1 | 0% | 934 | 844 | -10% | 0 | 0 | — |
case-03 | pass→pass | 9,022 | 3,545 | -61% | 1 | 1 | 0% | 1,352 | 928 | -31% | 0 | 0 | — |
case-04 | pass→pass | 7,134 | 3,207 | -55% | 1 | 1 | 0% | 1,206 | 909 | -25% | 0 | 0 | — |
case-05 | pass→pass | 10,109 | 5,571 | -45% | 1 | 1 | 0% | 1,821 | 1,336 | -27% | 0 | 0 | — |
case-06 | fail→pass | 13,281 | 4,903 | -63% | 1 | 1 | 0% | 1,923 | 1,044 | -46% | 0 | 0 | — |
case-07 | pass→pass | 12,092 | 7,053 | -42% | 1 | 1 | 0% | 1,886 | 1,424 | -24% | 0 | 0 | — |
case-08 | pass→pass | 9,757 | 3,272 | -66% | 1 | 1 | 0% | 1,454 | 931 | -36% | 0 | 0 | — |
case-09 | pass→pass | 9,852 | 4,777 | -52% | 1 | 1 | 0% | 1,183 | 1,111 | -6% | 0 | 0 | — |
case-10 | pass→pass | 8,850 | 5,624 | -36% | 1 | 1 | 0% | 1,463 | 1,353 | -8% | 0 | 0 | — |
case-11 | pass→pass | 3,243 | 4,049 | +25% | 1 | 1 | 0% | 527 | 979 | +86% | 0 | 0 | — |
case-12 | pass→pass | 13,035 | 6,588 | -49% | 1 | 1 | 0% | 1,930 | 1,332 | -31% | 0 | 0 | — |
case-13 | pass→pass | 9,524 | 5,478 | -42% | 1 | 1 | 0% | 1,364 | 1,147 | -16% | 0 | 0 | — |
case-14 | fail→fail | 6,310 | 3,487 | -45% | 1 | 1 | 0% | 894 | 788 | -12% | 0 | 0 | — |
case-15 | pass→pass | 3,445 | 2,969 | -14% | 1 | 1 | 0% | 549 | 787 | +43% | 0 | 0 | — |
case-16 | pass→pass | 24,138 | 8,807 | -64% | 1 | 1 | 0% | 1,953 | 1,615 | -17% | 0 | 0 | — |
case-17 | pass→pass | 13,859 | 4,458 | -68% | 1 | 1 | 0% | 1,976 | 1,022 | -48% | 0 | 0 | — |
case-18 | pass→pass | 24,128 | 6,694 | -72% | 1 | 1 | 0% | 1,855 | 1,305 | -30% | 0 | 0 | — |
case-19 | pass→pass | 6,954 | 9,763 | +40% | 1 | 1 | 0% | 1,234 | 867 | -30% | 0 | 0 | — |
case-20 | pass→pass | 12,818 | 8,942 | -30% | 1 | 1 | 0% | 1,925 | 1,716 | -11% | 0 | 0 | — |
case-21 | pass→pass | 8,578 | 4,486 | -48% | 1 | 1 | 0% | 1,442 | 1,111 | -23% | 0 | 0 | — |
case-22 | pass→pass | 6,455 | 2,813 | -56% | 1 | 1 | 0% | 1,217 | 889 | -27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.