Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Turn public 1-3-star Shopify App Store review rows into a P0-P3 triage brief: incident risk, repeated friction, pricing confusion, feature requests, and an explicit needs-human-read bucket.
.claude/skills/sickn33-shopify-review-triage/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 542% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 173% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 235% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 184% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 234% | 0% |
Takes rows of public Shopify App Store review text and produces one prioritized brief a product or support owner can act on: what kind of problem each review describes, how badly it can hurt, what to do first, and where the original wording came from.
It is built for independent Shopify app teams and the agencies that run their support — the case where low-star reviews arrive scattered across several listings plus a few watched competitors, and the failure mode is treating them all as equally urgent.
The rubric below is not invented here. It is the published rule set behind a free review triage worksheet and manual triage guide (links under Additional Resources), reproduced so a manual pass, the worksheet, and this skill sort the same row the same way.
This skill needs no network access, no scripts, and no system packages. The person you are helping supplies the review text.
prioritized, or clustered — even when they never say "triage", "severity", or "P0".
whether it is an incident, a UX problem, a pricing copy problem, or a feature request.
watched competitors.
rules below.
These are not style preferences. Breaking one makes the output worse than nothing.
order data, personal contact details, internal telemetry, or anything else not already public on a listing page. If such data appears in the input, stop, say which rows are affected, and ask for them to be removed before continuing.
URL that was not supplied. A row with no link gets source: not captured — never a guessed one.
labeled first pass — not human-checked. Only a person who read the review and checked it against their own systems may relabel an item human-checked.
showed a blank screen", never "the editor is broken". The distinction survives into the brief.
of exhaustive coverage of a listing, a period, or an app.
advice. Suggest actions; do not predict results.
support ticket, message a reviewer, or publish anything. Hand the draft back to the team and let a person decide what to send.
First ask which app names the team owns. Before any row is classified, ask for two lists of app names, spelled exactly as they appear in the rows:
textowned: Example Popup App, Example Currency App competitors: Rival Popup App, Rival Currency App
This is the only thing that makes tie-break 4 (a competitor's incident never becomes your P0) applicable, so collect it first. It stays public data: app names as published on their listings, nothing about accounts, merchants, org structure, or internal identifiers. Do not ask for more than the names, and do not infer ownership from the review text, the first-person voice in a review, or which app appears most often.
If an app name in a row appears in neither list, its ownership is unknown. Classify the row's content normally, then file it under needs human read with ownership: not supplied instead of placing it in a priority bucket or in competitor watch — a guessed owner is exactly the kind of invented evidence hard rule 2 forbids.
Then ask for one review per line. The full form keeps the source link, which the brief needs:
textrating | app name | review date | public reviews URL | review text
The shorter form used by the free worksheet is also fine — treat field 1 as the rating when it is a bare 1–5 (optionally followed by star/stars/★), otherwise as the app name:
textrating | app name | review text
Rules for this step:
# are comments. Blank lines are skipped.source: not captured through to the brief. Do not dropthe row and do not fabricate a link.
helping pastes the public rows they already opened.
classify correctly (a 5★ review often lands in feature requests or needs-human-read), so keep them if they were supplied, but never present them as low-star signal.
Lower-case the review text and normalize curly apostrophes (’ → ') before matching, so a pasted "won’t load" still matches won't load.
Five buckets. Each row gets exactly one primary bucket — the first dimension below, in this order, with any matching keyword. Further matches are recorded as secondary, never as a second brief item.
The purchase path, app activation, or merchant data may be at stake right now. Left alone it costs the merchant money and the team installs.
Suggested action. Try to reproduce on a test store today. If confirmed, treat it as an incident: fix or mitigate first, then reply to the reviewer with what changed.
Signal keywords. won't load, wont load, won't open, wont open, can't close, cannot close, won't close, blank screen, broken, crash, stopped working, not working, doesn't work, does not work, checkout, losing sales, lost sales, error
The product works, but the same struggle keeps showing up across reviews or against an open support theme. Repetition is the signal, not volume of adjectives.
Suggested action. Log it against the matching support theme. If the same complaint repeats across rows, schedule a UX fix ahead of new feature work.
Signal keywords. confusing, unclear, hard to, difficult, complicated, clunky, slow, couldn't figure, could not figure, annoying, had to contact support, setup took, too many steps
What the merchant expected to pay and what happened diverged. Usually a copy problem in the listing, the plan limits, or the upgrade prompts — not a code problem.
Suggested action. Compare what the reviewer expected with the listing's pricing section and in-app upgrade prompts; clarify the copy where they diverge.
Signal keywords. pricing, price, charged, charge, billing, billed, expensive, free plan, trial, refund, hidden fee, hidden cost, paywall
The merchant wants something the app does not do, or could not find. Valuable as a log entry, rarely urgent on its own.
Suggested action. Add it to the feature-request log with a link to the review. If the capability already exists, reply to the reviewer with where to find it.
Signal keywords. wish, would be great, would love, please add, feature request, missing, if only, would like, no option to, needs an option, hope you add, add support for
No keyword matched. Vague frustration, sarcasm, mixed praise, or a story that needs context.
Suggested action. No keyword matched. Read the full review yourself and file it manually — the heuristic makes no guess here.
Priority. The worksheet labels this bucket P2 and sorts it last. Treat that label as provisional placement in the queue, not as a severity judgment — nothing has been judged yet.
P0 with pricing noted as secondary. Never split one review across two brief items.
reviews within about 60 days, move it up one level and say how many rows drove the change.
problem, unless a recent row corroborates it. Cite it as context, never as the headline.
ownership lists from step 1: owned keeps its rubric bucket, competitors moves to the competitor watch section whatever its keywords matched, and a name in neither list goes to needs human read with ownership: not supplied. A competitor's incident is roadmap, positioning, or copy input — never your P0.
uncertainty into a priority label.
The first pass is where this skill stops being able to help on its own. Before any item is presented as more than a keyword match, a person on the team has to:
support inbox for matching signals from the same period;
Ask for these outcomes rather than assuming them. Until you have them, every item stays labeled first pass — not human-checked, including in the summary line. An unverified P0 is a candidate, not an incident.
Known limits to state plainly when they apply: keyword matching is English-only, misses sarcasm and context, can misfile a review that mentions "checkout" in passing, and sees only the rows supplied.
One document per portfolio, sections in rubric order, every item carrying an owner, a next action, and a source link. An item without an owner is a note, not a brief entry.
<!-- brief-template -->
markdown# Low-star review brief — {portfolio or team name} — week of {YYYY-MM-DD} Scope: {apps monitored} · {competitors watched} · {N} rows supplied, {date range}. Covers only the rows supplied — no claim of exhaustive coverage. Reviews are customer reports, not verified defects. Items marked "first pass" are unverified keyword matches; "human-checked" means a person read the review and checked it. ## P0 — Incident risk - **{App} — {signal in a few words}** ({rating}★, {review date}, source: {public reviews URL or not captured}) - Reviewer reports: {one sentence, in their words where possible} - Status: first pass — not human-checked / human-checked - Reproduced: {yes / no / attempted — notes} - Next action: {action} — owner {name}, due {date} ## P1 — Repeated friction - **{App} — {theme}** ({rating}★, {date}, source: {public reviews URL or not captured}; also seen: {where}) - Status: first pass — not human-checked / human-checked - Next action: {UX or docs change} — owner {name}, due {date} ## P2 — Pricing confusion - **{App} — {signal}** ({rating}★, {date}, source: {public reviews URL or not captured}) - Expected vs. actual: {one line} - Status: first pass — not human-checked / human-checked - Next action: {copy or prompt change} — owner {name}, due {date} ## P3 — Feature requests - **{App} — {request}** ({rating}★, {date}, source: {public reviews URL or not captured}) — {log it / already exists → reply with where to find it} ## Needs human read - **{App}** ({rating}★, {date}, source: {public reviews URL or not captured}) — {no keyword matched; what a human should look for}{, or: ownership: not supplied — app name on neither list} ## Competitor watch - **{Competitor} — {signal}**: {what it implies for our roadmap, copy, or positioning} ## Decisions this week - {one decision or experiment, with the row(s) that motivated it}
Open the summary line with the counts, e.g. "Triaged 8 rows supplied: 3 incident risk, 2 repeated friction, 1 pricing confusion, 1 feature request, 1 needs human read — first pass, not human-checked."
Refuse to deliver until every line is true:
source: not captured.owned list; every competitor row sits in competitorwatch; every unlisted app name says ownership: not supplied under needs human read.
that did not happen.
These eight fictional rows are the worksheet's own example set, so the two tools can be compared directly. Two of them are deliberately 4★ and 5★, to exercise the feature-request and needs-human-read buckets.
Ownership context, collected before any of it is classified:
textowned: Example Popup App, Example Currency App, Example Reviews App competitors: (none supplied)
text1 | Example Popup App | The editor shows a blank screen and the popup won't load. We are losing sales every day. 2 | Example Popup App | The overlay can't close on mobile and it blocks the checkout button. 1 | Example Currency App | Conversion is broken at checkout and we were still billed for the month. 3 | Example Currency App | Setup took hours and the settings screen is confusing. Support was slow to reply. 3 | Example Reviews App | The widget looks fine but the template editor is confusing and hard to use on a tablet. 2 | Example Currency App | We kept getting charged after uninstalling, and the pricing page never mentioned this. 4 | Example Reviews App | Great app, but I wish it could export reviews to CSV. Please add filtering by country. 5 | Example Reviews App | Does what it promises and support replied the same day.
First pass over those rows:
textrow 1 → P0 incident risk row 2 → P0 incident risk row 3 → P0 incident risk (secondary: pricing confusion) row 4 → P1 repeated friction row 5 → P1 repeated friction row 6 → P2 pricing confusion row 7 → P3 feature request row 8 → needs human read
Explanation: Rows 4 and 5 both matched confusing, so they are flagged as a repeated theme — two rows, which is a cluster to watch, not yet the three that trigger escalation. Row 3 is a single P0 item with pricing recorded as secondary, never two items. Row 8 matched nothing and stays unjudged. All three app names are on the owned list, so every bucket above is the team's own queue and competitor watch is empty; had Example Reviews App been listed as a competitor instead, rows 5, 7, and 8 would move there and none of them could become a P0. None of these rows carried a source URL, so each item would read source: not captured until the team supplies the listing links.
text1 | Example Popup App | 2026-07-28 | https://apps.shopify.com/example-popup-app/reviews?ratings%5B%5D=1 | The editor shows a blank screen and the popup won't load. We are losing sales every day.
Rendered into the brief:
markdown## P0 — Incident risk - **Example Popup App — editor reported blank, popup reported not loading** (1★, 2026-07-28, [source](https://apps.shopify.com/example-popup-app/reviews?ratings%5B%5D=1)) - Reviewer reports: the editor shows a blank screen, the popup does not load, and they are losing sales daily. - Status: first pass — not human-checked - Reproduced: not yet attempted - Next action: attempt reproduction on a development store today — owner {name}, due {date}
Explanation: It files as a P0 only because Example Popup App is on the owned list; the same row from a competitor listing would render under competitor watch instead. The wording stays a report ("the reviewer reports"), the status stays first pass — not human-checked until a person verifies it, and the source link is the listing's public reviews page with the rating filter kept — the App Store has no per-review permalink.
source: not captured forward when a row has no link, so the gap is visible.needs-human-read; that is the correct outcome, not a bug to work around by translating first.
"checkout" or "missing" in passing.
team's error tracker, or their support inbox.
until a person reproduces it.
for clarification if required inputs, permissions, or safety boundaries are missing.
fetch listings, call APIs, or read files outside what the person supplies.
personal contact details, or internal telemetry appear in the input, stop, name the affected rows, and ask for them to be removed before continuing.
posting a public developer reply, opening a ticket, or contacting a reviewer is out of scope under every circumstance (hard rule 7).
"the reviewer".
Solution: Cite the listing's public reviews page, keep the rating filter if one was used (…/reviews?ratings%5B%5D=1), and pin the item with the review date plus the reviewer's first few words so a human can find it again.
checkout is the noisiest keyword in the set — it fires on "we love the checkoutupsell". Solution: A P0 whose only evidence is the word checkout is a needs-human-read row wearing a P0 badge. Say so instead of promoting it.
missing and error cross buckets — "missing a dark mode" is P3, "settings pageerrors out" is P0. Solution: Primary-bucket order resolves the collision mechanically; the human pass fixes the ones where it guessed wrong.
Solution: It still goes to competitor watch. A competitor's P0 is never yours.
inflating every count in the summary line. Solution: One review, one item. Secondary matches are annotations.
second | into the review text, so a five-field row displays its date and URL inside the quote. Solution: Paste the short form into the worksheet and keep the long form here.
@customer-research — when the goal is broader voice-of-customer synthesis rather thanprioritizing a specific set of low-star review rows.
@shopify-apps — when the next step is actually building or fixing the Shopify app behavior atriaged P0 points at.
@before-you-build — when a P3 feature request needs product-risk review before it becomesroadmap work.
This skill packages the public rubric behind Shopify App Review Brief, an independent open-source project that is not affiliated with or endorsed by Shopify Inc. or any app developer. The same four dimensions, priorities, keyword lists, and suggested actions are published in three places:
Upstream source repository: alfredtech2026/shopify-app-review-brief (MIT).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | fail→pass | 7,425 | 8,606 | +16% | 1 | 1 | 0% | 1,128 | 7,241 | +542% | 0 | 0 | — |
case-01 | fail→pass | 47,067 | 42,862 | -9% | 1 | 1 | 0% | 3,123 | 8,519 | +173% | 0 | 0 | — |
case-02 | fail→pass | 15,203 | 24,807 | +63% | 1 | 1 | 0% | 3,167 | 10,603 | +235% | 0 | 0 | — |
case-03 | fail→pass | 16,056 | 42,096 | +162% | 1 | 1 | 0% | 2,850 | 8,092 | +184% | 0 | 0 | — |
case-04 | fail→fail | 12,177 | 12,409 | +2% | 1 | 1 | 0% | 1,878 | 7,601 | +305% | 0 | 0 | — |
case-05 | fail→pass | 10,657 | 5,049 | -53% | 1 | 1 | 0% | 1,928 | 6,437 | +234% | 0 | 0 | — |
case-06 | fail→pass | 8,150 | 5,889 | -28% | 1 | 1 | 0% | 1,342 | 6,648 | +395% | 0 | 0 | — |
case-08 | fail→pass | 9,855 | 13,748 | +40% | 1 | 1 | 0% | 1,715 | 8,467 | +394% | 0 | 0 | — |
case-09 | pass→pass | 8,473 | 11,995 | +42% | 1 | 1 | 0% | 1,343 | 7,965 | +493% | 0 | 0 | — |
case-10 | fail→pass | 7,711 | 11,403 | +48% | 1 | 1 | 0% | 1,236 | 7,831 | +534% | 0 | 0 | — |
case-11 | fail→pass | 7,859 | 10,288 | +31% | 1 | 1 | 0% | 1,413 | 7,662 | +442% | 0 | 0 | — |
case-12 | fail→pass | 7,482 | 6,937 | -7% | 1 | 1 | 0% | 1,166 | 6,901 | +492% | 0 | 0 | — |
case-13 | fail→pass | 7,304 | 5,857 | -20% | 1 | 1 | 0% | 1,099 | 6,721 | +512% | 0 | 0 | — |
case-14 | fail→pass | 7,840 | 10,126 | +29% | 1 | 1 | 0% | 1,334 | 7,560 | +467% | 0 | 0 | — |
case-15 | fail→pass | 8,465 | 7,236 | -15% | 1 | 1 | 0% | 1,261 | 7,002 | +455% | 0 | 0 | — |
case-16 | fail→pass | 5,886 | 7,440 | +26% | 1 | 1 | 0% | 1,104 | 7,234 | +555% | 0 | 0 | — |
case-17 | pass→pass | 8,972 | 6,648 | -26% | 1 | 1 | 0% | 1,530 | 6,869 | +349% | 0 | 0 | — |
case-18 | pass→pass | 9,446 | 8,033 | -15% | 1 | 1 | 0% | 1,422 | 7,195 | +406% | 0 | 0 | — |
case-19 | pass→pass | 5,756 | 9,142 | +59% | 1 | 1 | 0% | 869 | 7,474 | +760% | 0 | 0 | — |
case-20 | fail→pass | 5,791 | 7,696 | +33% | 1 | 1 | 0% | 851 | 7,087 | +733% | 0 | 0 | — |
case-21 | fail→pass | 7,799 | 7,800 | +0% | 1 | 1 | 0% | 1,175 | 7,117 | +506% | 0 | 0 | — |
case-22 | fail→pass | 7,417 | 9,133 | +23% | 1 | 1 | 0% | 1,127 | 7,476 | +563% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +77 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.