Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Interpret third-party feedback by running parallel internal and peer interpretations to surface intent, correctness concerns, and ambiguities. Use when the user asks to "interpret feedback", "interpret comments", "what does this feedback mean", "clarify reviewer intent", "understand this review", or "interpret these suggestions".
.claude/skills/tobihagemann-interpret-feedback/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -29% | 0% |
Run two independent interpretations of third-party feedback in parallel (internal + codex peer), then reconcile into enriched items with clear intent summaries. Designed for feedback where the author's intent is ambiguous or the correctness of suggestions is uncertain.
Determine the feedback to interpret:
For each item, collect whatever context is available: code snippets, diffs, surrounding discussion, file paths, line numbers. More context produces better interpretation.
Emit both Agent tool calls below in one assistant message. Each Agent call uses model: "opus" and no name. Wait for every agent to report before continuing. Do not begin the next step on a partial set, and do not relaunch an agent that has not yet reported. That is two Agent tool calls total. Both agents' prompts must direct them to treat the shared working tree and its git index as read-only and to interpret by reading and reasoning. HEAD stays where it is: read other refs with git show <ref>:<path> rather than git checkout or git switch.
Spawn a subagent with the feedback items and all available context. Instruct it to:
/peer-review SkillLaunch an Agent tool call whose prompt instructs the subagent to invoke /peer-review via the Skill tool. Describe the request in natural language:
The prompt must also state explicitly that the subagent's final assistant message must contain the verbatim findings text /peer-review produced.
Merge the two interpretations for each feedback item:
| Agreement | Action | |-----------|--------| | Both agree on intent and correctness | High confidence. Use the shared interpretation. | | Intent agrees, correctness differs | Flag the correctness concern with both perspectives. | | Intent disagrees | Flag as ambiguous. Present both readings and note which has stronger evidence. |
For each feedback item, output the original feedback followed by the interpretation:
### Item <N>: <short label>
**Original:** <feedback text, truncated if long>
**File:** <path:line if applicable>
**Intent:** <reconciled interpretation of what the author wants>
**Correctness:** <sound | concern: <explanation>>
**Confidence:** <high | medium | low>
**Ambiguity:** <none | <description of unclear aspects>>
<If interpreters disagreed, show both perspectives>After all items, add a summary:
## Interpretation Summary
- Total items: <N>
- High confidence: <N>
- Correctness concerns: <N>
- Ambiguous intent: <N>Then use the TaskList tool and proceed to any remaining task.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,169 | 35,327 | +332% | 1 | 1 | 0% | 1,282 | 1,460 | +14% | 0 | 0 | — |
case-02 | fail→fail | 4,621 | 6,308 | +37% | 1 | 1 | 0% | 224 | 1,610 | +619% | 0 | 0 | — |
case-03 | fail→fail | 29,085 | 6,697 | -77% | 1 | 1 | 0% | 5,221 | 1,382 | -74% | 0 | 0 | — |
case-04 | fail→pass | 9,411 | 13,981 | +49% | 1 | 1 | 0% | 1,983 | 2,758 | +39% | 0 | 0 | — |
case-05 | pass→pass | 8,174 | 3,321 | -59% | 1 | 1 | 0% | 1,527 | 1,632 | +7% | 0 | 0 | — |
case-06 | fail→pass | 11,270 | 1,672 | -85% | 1 | 1 | 0% | 1,865 | 1,301 | -30% | 0 | 0 | — |
case-07 | pass→pass | 5,285 | 2,361 | -55% | 1 | 1 | 0% | 801 | 1,433 | +79% | 0 | 0 | — |
case-08 | pass→pass | 11,385 | 4,594 | -60% | 1 | 1 | 0% | 1,880 | 1,818 | -3% | 0 | 0 | — |
case-09 | pass→pass | 8,747 | 4,981 | -43% | 1 | 1 | 0% | 1,570 | 1,837 | +17% | 0 | 0 | — |
case-10 | fail→pass | 10,486 | 5,187 | -51% | 1 | 1 | 0% | 1,619 | 1,389 | -14% | 0 | 0 | — |
case-11 | pass→pass | 31,260 | 3,042 | -90% | 1 | 1 | 0% | 2,783 | 1,601 | -42% | 0 | 0 | — |
case-12 | fail→pass | 9,173 | 1,880 | -80% | 1 | 1 | 0% | 1,718 | 1,323 | -23% | 0 | 0 | — |
case-13 | fail→pass | 14,077 | 3,541 | -75% | 1 | 1 | 0% | 2,443 | 1,723 | -29% | 0 | 0 | — |
case-14 | fail→fail | 9,751 | 1,696 | -83% | 1 | 1 | 0% | 1,687 | 1,295 | -23% | 0 | 0 | — |
case-15 | fail→fail | 9,728 | 2,133 | -78% | 1 | 1 | 0% | 1,686 | 1,316 | -22% | 0 | 0 | — |
case-16 | fail→fail | 19,117 | 3,475 | -82% | 1 | 1 | 0% | 3,311 | 1,399 | -58% | 0 | 0 | — |
case-17 | fail→pass | 7,162 | 1,806 | -75% | 1 | 1 | 0% | 1,235 | 1,340 | +9% | 0 | 0 | — |
case-18 | fail→pass | 8,185 | 2,219 | -73% | 1 | 1 | 0% | 1,356 | 1,416 | +4% | 0 | 0 | — |
case-19 | fail→pass | 5,177 | 2,264 | -56% | 1 | 1 | 0% | 921 | 1,385 | +50% | 0 | 0 | — |
case-20 | pass→fail | 8,964 | 20,655 | +130% | 1 | 1 | 0% | 1,526 | 4,274 | +180% | 0 | 0 | — |
case-21 | fail→fail | 3,072 | 8,007 | +161% | 1 | 1 | 0% | 489 | 1,819 | +272% | 0 | 0 | — |
case-22 | fail→fail | 3,346 | 7,064 | +111% | 1 | 1 | 0% | 359 | 1,376 | +283% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 17 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/21/2026 | +18% |
Other measured skills in the registry, with their headline benchmark lift.