Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Produce honest product reviews and buying guides without fabricated first-hand experience, using disclosed evidence tiers: verified specs, owner-experience synthesis at scale, expert triangulation, and hands-on only when true. Use this skill whenever the user wants to write product reviews or buying guides without hands-on access, set up a review methodology, add evidence disclosure to review content, align reviews with Google's reviews system or FTC affiliate disclosure expectations, or decide
.claude/skills/rampstackco-evidence-based-reviews/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 106% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 144% | 0% |
Produce product reviews and buying guides whose claims match their evidence. The method is evidence synthesis with a disclosed basis, and one rule anchors everything: never claim hands-on time that did not happen.
Most review sites fail this rule quietly. "We tested" appears above content assembled from spec sheets, and readers have learned to smell it. The honest alternative is stronger, not weaker: synthesis of owner experience at scale, verified specifications, and triangulated expert sources is demonstrable original analysis, and disclosing it builds the trust that fabricated testing destroys.
content-and-copy; this skill governs the evidence and claims layer of review content specifically)brand-voice)seo-onpage)content-brief-authoring)If genuine hands-on testing at scale exists, with instrumentation and protocols, this skill still applies: the tiers do not change, the methodology block simply states tier 4 and the testing protocol becomes the disclosed basis.
The anchor rule: never claim hands-on time that did not happen. Every other part of the method exists to make honest content strong enough that the lie is unnecessary.
Specifications cited to the maker's own page, not to aggregator copies that drift. The verification step is the work: cross-check the spec against the manufacturer's current listing, note the date, and flag discrepancies between sources instead of picking one silently. A spec table built this way is a fact layer; a spec table pasted from another review site is a rumor layer.
When it suffices: factual comparisons, fit and compatibility answers, price-tier groupings. It never supports a quality verdict on its own.
Retailer review corpora, forum threads, warranty and return patterns, read across sources and summarized honestly. Honest synthesis means:
When it suffices: durability and reliability verdicts, real-world quirks specs never show, satisfaction patterns by use case. This tier is the workhorse of honest no-hands-on reviewing, and it is genuine original analysis when done at scale.
Named expert sources compared against each other. Triangulation means surfacing agreement and disagreement, not averaging verdicts into mush. When two credible testers reach opposite conclusions, the honest move is to say so and explain the conditions that might account for it. Anonymous "experts agree" is not a tier; it is decoration.
When it suffices: performance claims that require instrumentation the site lacks, technique-dependent judgments, category context.
Stated only when it happened, flagged as such, with the extent quantified: what was done, how much, under what conditions. When hands-on experience arrives later for a piece published on tiers 1 to 3, the piece is upgraded and the update is marked with what changed. Never backfilled to look like it was always hands-on; the upgrade trail is itself a trust signal.
Every review and buying guide carries a short disclosure block naming its evidence basis, placed with the criteria, written for readers rather than lawyers. The fillable template and a worked example live in references/methodology-block-template.md.
The block names: the criteria in the order they were weighted, the tiers actually used (specifically, with sources), what was done hands-on or the words "none claimed," and the update line. A block that claims a tier the piece did not use is the same lie the anchor rule bans, in smaller type.
The guidance asks for demonstrated first-hand experience OR demonstrable original research and analysis. The second branch is this skill's lane, and "original analysis" means concrete work product:
A site that does tier 2 and 3 work honestly is doing original analysis by the guidance's own definition. The methodology block is how the work shows.
Affiliate relationships are disclosed in plain language, placed where the recommendation is, not in a footer. The disclosure travels with the monetized content: near the picks, in body-size text, before or beside the first affiliate link a reader can act on. "Plain language" means a reader who has never heard the word affiliate understands that the site earns a commission and that the price they pay does not change. Euphemisms ("partner links," "support the site") fail the plain-language test.
Markup is a claim. Apply the same honesty to structured data as to prose:
For a methodology setup: a short methodology standard the site adopts (the tiers, the block template filled with the site's sources, the schema policy), suitable for docs or an editorial guide.
For a piece-level engagement: the methodology block for that piece, the criteria list, the evidence file (sources gathered per tier with sample sizes), and the claims audit if the piece existed before this skill did.
This skill's output depends on data, measurements, or tool results it cannot generate on its own. When a required input, tool, or data source is unavailable or unverifiable, the sanctioned output is the deliverable with the gap stated: what was needed, what was actually obtained or verified, and which parts of the output are affected. Fabricating, estimating, or interpolating a required number to complete the deliverable is never sanctioned. A stated gap is a complete answer.
references/evidence-tiers.md - The four tiers in depth: verification steps, synthesis honesty, triangulation practice, and the upgrade-and-mark pattern.references/methodology-block-template.md - The fillable per-piece disclosure block, a worked example, and placement rules.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→pass | 18,763 | 18,175 | -3% | 1 | 1 | 0% | 2,958 | 5,079 | +72% | 0 | 0 | — |
case-04 | pass→pass | 15,559 | 13,024 | -16% | 1 | 1 | 0% | 2,536 | 4,578 | +81% | 0 | 0 | — |
case-05 | pass→pass | 11,282 | 8,940 | -21% | 1 | 1 | 0% | 1,744 | 3,623 | +108% | 0 | 0 | — |
case-01 | fail→pass | 33,943 | 21,214 | -38% | 1 | 1 | 0% | 5,350 | 5,843 | +9% | 0 | 0 | — |
case-02 | fail→fail | 20,006 | 22,067 | +10% | 1 | 1 | 0% | 3,203 | 5,982 | +87% | 0 | 0 | — |
case-06 | pass→pass | 14,426 | 9,644 | -33% | 1 | 1 | 0% | 2,025 | 3,663 | +81% | 0 | 0 | — |
case-07 | pass→pass | 13,003 | 10,125 | -22% | 1 | 1 | 0% | 1,831 | 3,819 | +109% | 0 | 0 | — |
case-08 | fail→pass | 14,859 | 12,852 | -14% | 1 | 1 | 0% | 2,066 | 4,256 | +106% | 0 | 0 | — |
case-09 | pass→pass | 18,918 | 15,028 | -21% | 1 | 1 | 0% | 3,032 | 4,745 | +56% | 0 | 0 | — |
case-10 | fail→pass | 13,514 | 8,326 | -38% | 1 | 1 | 0% | 2,027 | 3,526 | +74% | 0 | 0 | — |
case-11 | pass→pass | 10,120 | 4,912 | -51% | 1 | 1 | 0% | 1,551 | 3,083 | +99% | 0 | 0 | — |
case-12 | pass→pass | 13,296 | 5,967 | -55% | 1 | 1 | 0% | 1,956 | 3,285 | +68% | 0 | 0 | — |
case-13 | pass→pass | 18,283 | 13,170 | -28% | 1 | 1 | 0% | 2,729 | 4,206 | +54% | 0 | 0 | — |
case-14 | pass→pass | 6,815 | 8,509 | +25% | 1 | 1 | 0% | 1,198 | 3,719 | +210% | 0 | 0 | — |
case-15 | pass→fail | 16,892 | 20,086 | +19% | 1 | 1 | 0% | 2,552 | 5,348 | +110% | 0 | 0 | — |
case-16 | pass→fail | 10,542 | 9,544 | -9% | 1 | 1 | 0% | 1,952 | 3,991 | +104% | 0 | 0 | — |
case-17 | pass→fail | 21,476 | 22,715 | +6% | 1 | 1 | 0% | 3,688 | 6,037 | +64% | 0 | 0 | — |
case-18 | pass→fail | 16,545 | 17,653 | +7% | 1 | 1 | 0% | 2,654 | 4,978 | +88% | 0 | 0 | — |
case-19 | pass→pass | 12,913 | 9,097 | -30% | 1 | 1 | 0% | 1,887 | 3,696 | +96% | 0 | 0 | — |
case-20 | pass→pass | 22,043 | 9,024 | -59% | 1 | 1 | 0% | 3,951 | 4,002 | +1% | 0 | 0 | — |
case-21 | fail→pass | 9,937 | 7,965 | -20% | 1 | 1 | 0% | 1,488 | 3,630 | +144% | 0 | 0 | — |
case-22 | pass→pass | 8,290 | 8,951 | +8% | 1 | 1 | 0% | 1,341 | 3,852 | +187% | 0 | 0 | — |
case-23 | pass→pass | 12,267 | 8,617 | -30% | 1 | 1 | 0% | 1,946 | 3,677 | +89% | 0 | 0 | — |
case-24 | pass→pass | 14,958 | 14,451 | -3% | 1 | 1 | 0% | 2,101 | 4,375 | +108% | 0 | 0 | — |
case-25 | pass→pass | 16,337 | 10,976 | -33% | 1 | 1 | 0% | 2,259 | 4,009 | +77% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +4 percentage points is the difference between those two pass rates over the 25 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.