Install any skill in seconds. Free to start, no credit card required.
Get Started Free →When the user wants to research customer pain points, complaints, or sentiment using review platforms like Trustpilot, G2, Capterra, or app stores. Also use when the user mentions "what are users saying", "competitor reviews", "pain points", or "voice of customer research".
.claude/skills/mkurman-review-mining/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 35% | 0% |
---|-----------|-----------|-------| | ... | ... | ... | ... |
Words users use for the problem: list of exact phrases] Words users use for the desired outcome: list of exact phrases] Emotional language: frustration words, relief words]
## Frameworks & Best Practices
**Where to mine by product type:**
| Product Type | Best Sources |
|-------------|-------------|
| B2B SaaS | G2, Capterra, TrustRadius |
| B2C / Consumer | Trustpilot, App Store, Play Store |
| Developer Tools | Reddit, Hacker News, GitHub Issues |
| E-commerce / DTC | Trustpilot, Amazon reviews |
| Any | Twitter/X complaints, Reddit threads |
**Review analysis principles:**
- **1-2 star reviews** reveal deal-breakers and switching triggers
- **3 star reviews** reveal "good enough but frustrated" — the most persuadable users
- **4-5 star reviews** reveal what users truly value (defend these in your product)
- **Recent reviews** (last 6-12 months) matter more than old ones
- **Verified purchase/user** reviews carry more weight
**Verbatim language is the output.** The exact words users use to describe their pain are more valuable than your summary. These become headlines, email subject lines, ad copy, and landing page copy.
**Common mistakes:**
- Only reading negative reviews (you miss what users actually value)
- Summarizing instead of quoting (you lose the authentic language)
- Treating all complaints equally (frequency x severity matters)
- Ignoring the context of who's reviewing (enterprise vs SMB, power user vs casual)
- Mining once and never returning (do this quarterly)
## Related Skills
- `competitive-analysis` — for broader competitor research beyond reviews
- `user-research-synthesis` — for synthesizing your own customer interviews
- `feedback-synthesis` — for analyzing feedback from your own users
- `cold-outreach` — use voice-of-customer language in prospecting emails
## Examples
**Prompt:** "I'm building a project management tool. What are the biggest pain points people have with Asana and Monday.com?"
**Good output includes:** Mining Trustpilot, G2, and Capterra for Asana and Monday.com, extracting the top 5-7 pain points with verbatim quotes, identifying switching triggers, and mapping them to positioning opportunities.
**Prompt:** "We're a Trustpilot alternative. Help me understand what businesses hate about Trustpilot."
**Good output includes:** Mining Trustpilot's own reviews (meta!), G2, and Reddit for complaints about Trustpilot, extracting themes like review gating, pricing, fake review handling, and producing a voice-of-customer swipe file the founder can use in outreach.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-16 | pass→pass | 12,596 | 10,894 | -14% | 1 | 1 | 0% | 1,942 | 2,354 | +21% | 0 | 0 | — |
case-01 | pass→pass | 23,284 | 19,436 | -17% | 1 | 1 | 0% | 3,659 | 3,495 | -4% | 0 | 0 | — |
case-02 | pass→pass | 24,991 | 24,282 | -3% | 1 | 1 | 0% | 3,641 | 4,420 | +21% | 0 | 0 | — |
case-03 | pass→pass | 19,650 | 17,533 | -11% | 1 | 1 | 0% | 2,771 | 3,425 | +24% | 0 | 0 | — |
case-04 | fail→fail | 6,752 | 4,032 | -40% | 1 | 1 | 0% | 995 | 1,217 | +22% | 0 | 0 | — |
case-05 | pass→pass | 14,838 | 13,568 | -9% | 1 | 1 | 0% | 2,110 | 2,623 | +24% | 0 | 0 | — |
case-06 | pass→fail | 20,133 | 20,137 | +0% | 1 | 1 | 0% | 3,193 | 3,977 | +25% | 0 | 0 | — |
case-07 | fail→pass | 16,242 | 17,880 | +10% | 1 | 1 | 0% | 2,399 | 3,398 | +42% | 0 | 0 | — |
case-08 | pass→pass | 13,544 | 14,885 | +10% | 1 | 1 | 0% | 2,083 | 2,885 | +39% | 0 | 0 | — |
case-09 | fail→pass | 16,401 | 15,864 | -3% | 1 | 1 | 0% | 2,451 | 3,069 | +25% | 0 | 0 | — |
case-10 | fail→pass | 14,698 | 16,967 | +15% | 1 | 1 | 0% | 2,125 | 3,096 | +46% | 0 | 0 | — |
case-11 | fail→pass | 17,700 | 17,004 | -4% | 1 | 1 | 0% | 2,571 | 3,036 | +18% | 0 | 0 | — |
case-12 | pass→pass | 12,990 | 13,712 | +6% | 1 | 1 | 0% | 1,948 | 2,701 | +39% | 0 | 0 | — |
case-13 | fail→fail | 12,468 | 12,723 | +2% | 1 | 1 | 0% | 2,016 | 2,732 | +36% | 0 | 0 | — |
case-14 | pass→pass | 14,486 | 10,531 | -27% | 1 | 1 | 0% | 2,007 | 2,205 | +10% | 0 | 0 | — |
case-15 | fail→fail | 16,543 | 14,091 | -15% | 1 | 1 | 0% | 2,527 | 2,747 | +9% | 0 | 0 | — |
case-17 | pass→pass | 16,239 | 13,502 | -17% | 1 | 1 | 0% | 2,429 | 2,744 | +13% | 0 | 0 | — |
case-18 | pass→pass | 19,983 | 19,398 | -3% | 1 | 1 | 0% | 2,888 | 3,478 | +20% | 0 | 0 | — |
case-19 | pass→pass | 14,407 | 14,007 | -3% | 1 | 1 | 0% | 2,173 | 2,712 | +25% | 0 | 0 | — |
case-20 | pass→pass | 13,109 | 13,418 | +2% | 1 | 1 | 0% | 1,994 | 2,726 | +37% | 0 | 0 | — |
case-21 | fail→pass | 11,218 | 11,024 | -2% | 1 | 1 | 0% | 1,694 | 2,286 | +35% | 0 | 0 | — |
case-22 | fail→pass | 17,242 | 17,897 | +4% | 1 | 1 | 0% | 2,634 | 3,324 | +26% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.