Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when receiving critical feedback on an analysis or manuscript, before implementing suggestions, especially if feedback seems unclear or methodologically questionable - requires verification, not performative agreement or blind changes
.claude/skills/k-dense-ai-receiving-critical-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 42% | 0% |
Critique of an analysis requires technical evaluation, not emotional performance.
Core principle: Verify before changing. Ask before assuming. Methodological correctness over social comfort.
A reviewer pointing at a confound is doing you a favor only if you check whether the confound is real. Reflexively agreeing — or reflexively defending — both skip the verification that makes review worthwhile.
WHEN receiving review feedback:
1. READ: The complete feedback without reacting
2. UNDERSTAND: Restate the methodological concern in your own words (or ask)
3. VERIFY: Check it against the data, code, and pre-registration
4. EVALUATE: Is it correct FOR THIS analysis?
5. RESPOND: Technical acknowledgment or reasoned pushback
6. ACT: One item at a time, re-running and re-verifying eachNEVER:
INSTEAD:
IF any item is unclear:
STOP - do not change anything yet
ASK for clarification
WHY: Methodological items interact. Adjusting for the wrong confound
because you misread the concern can bias the result further.BEFORE changing anything:
1. Is it methodologically correct for THIS data and design?
2. Would the change itself introduce bias (e.g., adding a collider as a covariate)?
3. Is there a pre-registered reason the analysis is the way it is?
4. Does the reviewer have the full context (the pre-registration, the data structure)?
IF the suggestion seems wrong:
Push back with technical reasoning and evidence
IF you can't verify:
Say so: "I can't verify this without [X]. Should I [investigate/ask]?"
IF it conflicts with the pre-registration:
A change to the registered analysis is a deviation. It renders that
analysis exploratory. Flag this explicitly before making the change.A reviewer may suggest a "better" analysis. Before adopting it:
How: technical reasoning, not defensiveness. Show the diagnostic, the DAG, the pre-registration line, or the reproduced number.
When feedback IS correct:
✅ "Verified — site does confound this. Added it; estimate drops to 0.09 [−0.01, 0.19]. Now inconclusive."
✅ "Confirmed leakage: the scaler was fit before the split. Refit on train only; AUC falls to 0.71."
✅ [Just fix it, re-run, and show the new number]
❌ "You're absolutely right!"
❌ "Great catch!"
❌ "Thanks for spotting that!"
❌ ANY gratitude or praise performanceWhy no thanks: the corrected, re-verified result shows you heard the feedback. State the fix and the new evidence.
If you pushed back and were wrong:
✅ "You were right — I checked and the assumption is violated. Re-running with the robust estimator."
❌ Long apology or defense of why you pushed back| Mistake | Fix | |---------|-----| | Performative agreement | State the concern or verify | | Blind change | Check against data/code/prereg first | | Batch edits without re-running | One at a time, re-verify each | | Assuming the reviewer is right | Check whether the change biases the result | | Silently changing the registered analysis | Flag it as a deviation → exploratory | | Avoiding pushback | Methodological correctness > comfort |
External feedback = hypotheses to verify, not orders to follow.
Verify against the data and the pre-registration. Question. Then act — and re-verify with science-superpowers:verifying-results-before-claiming.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-16 | fail→pass | 13,701 | 7,640 | -44% | 1 | 1 | 0% | 2,035 | 2,302 | +13% | 0 | 0 | — |
case-01 | fail→pass | 15,206 | 16,103 | +6% | 1 | 1 | 0% | 2,678 | 3,630 | +36% | 0 | 0 | — |
case-02 | fail→pass | 9,931 | 5,027 | -49% | 1 | 1 | 0% | 1,685 | 1,897 | +13% | 0 | 0 | — |
case-03 | fail→pass | 9,488 | 9,158 | -3% | 1 | 1 | 0% | 1,570 | 2,665 | +70% | 0 | 0 | — |
case-04 | pass→fail | 13,672 | 5,269 | -61% | 1 | 1 | 0% | 2,318 | 1,910 | -18% | 0 | 0 | — |
case-05 | pass→pass | 22,980 | 17,370 | -24% | 1 | 1 | 0% | 3,832 | 3,877 | +1% | 0 | 0 | — |
case-06 | pass→pass | 10,073 | 8,264 | -18% | 1 | 1 | 0% | 1,875 | 2,667 | +42% | 0 | 0 | — |
case-07 | fail→fail | 13,173 | 8,602 | -35% | 1 | 1 | 0% | 2,050 | 2,507 | +22% | 0 | 0 | — |
case-08 | fail→pass | 9,998 | 7,500 | -25% | 1 | 1 | 0% | 1,531 | 2,167 | +42% | 0 | 0 | — |
case-09 | pass→pass | 12,865 | 7,162 | -44% | 1 | 1 | 0% | 1,921 | 2,121 | +10% | 0 | 0 | — |
case-10 | fail→pass | 10,600 | 6,300 | -41% | 1 | 1 | 0% | 1,595 | 2,040 | +28% | 0 | 0 | — |
case-11 | fail→pass | 15,641 | 10,806 | -31% | 1 | 1 | 0% | 2,118 | 2,652 | +25% | 0 | 0 | — |
case-12 | fail→pass | 12,945 | 6,731 | -48% | 1 | 1 | 0% | 1,838 | 2,078 | +13% | 0 | 0 | — |
case-13 | pass→pass | 17,694 | 7,968 | -55% | 1 | 1 | 0% | 2,638 | 2,218 | -16% | 0 | 0 | — |
case-14 | pass→pass | 12,706 | 7,590 | -40% | 1 | 1 | 0% | 1,919 | 2,159 | +13% | 0 | 0 | — |
case-15 | fail→fail | 16,416 | 13,806 | -16% | 1 | 1 | 0% | 2,312 | 3,037 | +31% | 0 | 0 | — |
case-17 | pass→pass | 13,709 | 7,217 | -47% | 1 | 1 | 0% | 1,954 | 2,229 | +14% | 0 | 0 | — |
case-18 | fail→pass | 21,395 | 11,896 | -44% | 1 | 1 | 0% | 3,358 | 2,844 | -15% | 0 | 0 | — |
case-19 | fail→fail | 15,641 | 11,533 | -26% | 1 | 1 | 0% | 2,070 | 2,697 | +30% | 0 | 0 | — |
case-20 | pass→pass | 13,015 | 7,370 | -43% | 1 | 1 | 0% | 1,960 | 2,178 | +11% | 0 | 0 | — |
case-21 | pass→pass | 17,562 | 7,163 | -59% | 1 | 1 | 0% | 2,523 | 2,047 | -19% | 0 | 0 | — |
case-22 | pass→pass | 14,247 | 7,846 | -45% | 1 | 1 | 0% | 2,124 | 2,183 | +3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.