Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the superpowers-evaluator returned REWORK on a batch, when fixing rework items from an evaluation report, or when receiving any code review feedback on superpowers output. Requires technical rigor and verification instead of performative agreement or blind implementation.
.claude/skills/fradser-receiving-code-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 77% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 6% | 0% |
The superpowers-evaluator returns REWORK with file:line evidence. Your job is to fix the right thing — not to please the evaluator.
Core principle: Verify before implementing. Ask before assuming. Technical correctness over social comfort. The evaluator is a red-team reviewer, not an order-giver.
Run this for each rework item before changing code:
NEVER output, even when the evaluator is right:
INSTEAD:
If ANY rework item is unclear, do NOT implement the clear ones and defer the unclear ones. Items may be related — partial understanding produces wrong implementation.
Evaluator: 4 rework items. You understand items 1,2,4. Unclear on item 3.
WRONG: implement 1,2,4 now, ask about 3 later
RIGHT: re-read the cited checklist item for 3; if still unclear, note the
ambiguity in your return and implement 1,2,4 only after 3 is resolved
(or explicitly mark 3 as [AUTO-RESOLVED] per the sprint contract's
Autonomous Resolution Protocol with the most concrete interpretation)Push back (with technical reasoning, not defensiveness) when a rework item:
## Global Constraints blockHow to push back in your return:
If the evaluator was right and you were wrong: state the correction factually and move on. No long apology, no defending why you pushed back.
verification-before-completion)Evaluator REWORK = findings to evaluate, not orders to follow. Verify each item against the codebase. Question the wrong ones. Implement the right ones one at a time with verification. No performative agreement, no blind batch implementation.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | pass→pass | 10,663 | 6,564 | -38% | 1 | 1 | 0% | 1,920 | 1,850 | -4% | 0 | 0 | — |
case-09 | pass→pass | 10,642 | 5,102 | -52% | 1 | 1 | 0% | 1,703 | 1,677 | -2% | 0 | 0 | — |
case-10 | pass→pass | 9,641 | 5,679 | -41% | 1 | 1 | 0% | 1,632 | 1,824 | +12% | 0 | 0 | — |
case-01 | fail→pass | 12,480 | 12,876 | +3% | 1 | 1 | 0% | 1,976 | 2,830 | +43% | 0 | 0 | — |
case-02 | fail→pass | 12,270 | 11,631 | -5% | 1 | 1 | 0% | 2,096 | 2,715 | +30% | 0 | 0 | — |
case-03 | fail→fail | 12,690 | 4,075 | -68% | 1 | 1 | 0% | 2,680 | 1,253 | -53% | 0 | 0 | — |
case-04 | pass→fail | 34,448 | 11,534 | -67% | 1 | 1 | 0% | 6,131 | 2,401 | -61% | 0 | 0 | — |
case-05 | pass→pass | 17,290 | 11,792 | -32% | 1 | 1 | 0% | 2,791 | 2,801 | +0% | 0 | 0 | — |
case-06 | pass→fail | 11,156 | 3,311 | -70% | 1 | 1 | 0% | 1,792 | 1,099 | -39% | 0 | 0 | — |
case-07 | fail→pass | 5,649 | 3,585 | -37% | 1 | 1 | 0% | 806 | 1,427 | +77% | 0 | 0 | — |
case-11 | pass→pass | 7,469 | 3,984 | -47% | 1 | 1 | 0% | 1,330 | 1,691 | +27% | 0 | 0 | — |
case-12 | fail→pass | 9,004 | 6,239 | -31% | 1 | 1 | 0% | 1,393 | 1,863 | +34% | 0 | 0 | — |
case-13 | fail→pass | 8,521 | 5,599 | -34% | 1 | 1 | 0% | 1,661 | 1,755 | +6% | 0 | 0 | — |
case-14 | pass→pass | 6,216 | 4,495 | -28% | 1 | 1 | 0% | 1,037 | 1,531 | +48% | 0 | 0 | — |
case-20 | fail→pass | 3,719 | 3,155 | -15% | 1 | 1 | 0% | 641 | 1,356 | +112% | 0 | 0 | — |
case-15 | pass→pass | 12,986 | 5,958 | -54% | 1 | 1 | 0% | 2,380 | 2,012 | -15% | 0 | 0 | — |
case-16 | fail→pass | 9,678 | 6,248 | -35% | 1 | 1 | 0% | 1,424 | 1,850 | +30% | 0 | 0 | — |
case-17 | fail→pass | 10,611 | 3,989 | -62% | 1 | 1 | 0% | 1,658 | 1,523 | -8% | 0 | 0 | — |
case-18 | pass→fail | 5,488 | 4,913 | -10% | 1 | 1 | 0% | 866 | 1,735 | +100% | 0 | 0 | — |
case-19 | pass→pass | 10,557 | 7,434 | -30% | 1 | 1 | 0% | 1,700 | 2,052 | +21% | 0 | 0 | — |
case-21 | fail→fail | 11,280 | 5,243 | -54% | 1 | 1 | 0% | 1,729 | 1,716 | -1% | 0 | 0 | — |
case-22 | fail→pass | 4,127 | 5,790 | +40% | 1 | 1 | 0% | 658 | 1,919 | +192% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 20 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.