Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Before showing the user any substantive GTM deliverable (positioning, value prop, homepage, launch post, pricing, sales script, the brief or roadmap), stress-test it against the standard as an independent critic, because the agent that wrote it is the worst judge of whether it is good. Use as a gate right before presenting work, or when the user asks whether something is actually strong.
.claude/skills/aidevgtm-review-the-work/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 77% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 96% | 0% |
> The person who wrote it is the worst judge of whether it is good. Before this reaches the founder, stop being the author and become the skeptic.
Use this when: you are about to present any substantive deliverable, or the founder asks "is this actually good?" It is a gate, not a stage. It has no place in the linear sequence. Run it around any piece of work, every time.
An agent that just wrote something is biased to ship it. That bias is how generic, plausible, quietly wrong work reaches a founder who does not yet know enough to catch it. So separate the two jobs: the author drafts, a different lens judges. You do not present work because you made it. You present it because it survived a skeptic.
Steal the one rule that makes this work: the author may submit the work, the author may not issue the verdict. Switch roles on purpose. Become the developer who is skeptical, the buyer who is busy, the reviewer who has seen a hundred of these, and try to break it before the market does.
[assumption] is flagged out loud, not smuggled in as truth.Run the general standard, then the relevant skill's own checklist:
positioning-and-story): problem-first, not solution-first. Has a villain. Survives the "every competitor could say this" test.value-prop-that-converts): no banned words, built on a real job-to-be-done, and a developer would repeat it in their own words.the-homepage): written for the developer who visits, not the buyer who never does. Time to first value is obvious.launch-it): honest, not salesy. Opens on the developer's problem, not the product. Would not get flamed for overclaiming.pricing): anchored to a value metric, not a guessed number. The buyer could justify it to whoever holds the budget.founder-led-sales): passes the four-part deal test. Does not mistake a friendly free user for a buyer.start-here): strongest asset at the top, [validated] / [assumption] tags intact, in the founder's own words.strategy-and-roadmap): opens with a one-line Diagnosis, then Now / Next / Later. It is a plan, not a re-saved copy of the brief.[assumption] it depends on most, and decide whether it is safe to ship on or belongs in talk-to-users first.Built from real dev-tool GTM experience, with frameworks from Adam Frankl (The Developer-Facing Startup) and Jakub Czakon (markepear.dev). When a framework can't make the call, that's what a human is for: The DevTool GTM Company.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 9,817 | 9,729 | -1% | 1 | 1 | 0% | 1,413 | 2,435 | +72% | 0 | 0 | — |
case-02 | pass→pass | 12,351 | 9,419 | -24% | 1 | 1 | 0% | 1,698 | 2,566 | +51% | 0 | 0 | — |
case-03 | fail→pass | 11,239 | 12,812 | +14% | 1 | 1 | 0% | 1,611 | 2,853 | +77% | 0 | 0 | — |
case-04 | fail→pass | 13,193 | 14,314 | +8% | 1 | 1 | 0% | 1,980 | 3,060 | +55% | 0 | 0 | — |
case-05 | pass→fail | 16,600 | 16,235 | -2% | 1 | 1 | 0% | 2,639 | 3,466 | +31% | 0 | 0 | — |
case-06 | fail→pass | 14,389 | 10,647 | -26% | 1 | 1 | 0% | 2,158 | 2,706 | +25% | 0 | 0 | — |
case-07 | fail→pass | 15,634 | 14,700 | -6% | 1 | 1 | 0% | 2,313 | 3,335 | +44% | 0 | 0 | — |
case-08 | fail→fail | 10,608 | 9,699 | -9% | 1 | 1 | 0% | 1,593 | 2,511 | +58% | 0 | 0 | — |
case-09 | fail→pass | 10,175 | 11,245 | +11% | 1 | 1 | 0% | 1,459 | 2,861 | +96% | 0 | 0 | — |
case-10 | fail→pass | 14,571 | 12,682 | -13% | 1 | 1 | 0% | 2,069 | 2,840 | +37% | 0 | 0 | — |
case-11 | pass→pass | 13,557 | 11,317 | -17% | 1 | 1 | 0% | 1,845 | 2,944 | +60% | 0 | 0 | — |
case-12 | fail→fail | 11,016 | 12,973 | +18% | 1 | 1 | 0% | 1,658 | 2,793 | +68% | 0 | 0 | — |
case-13 | pass→pass | 15,489 | 9,711 | -37% | 1 | 1 | 0% | 2,035 | 2,594 | +27% | 0 | 0 | — |
case-14 | fail→pass | 12,267 | 9,302 | -24% | 1 | 1 | 0% | 1,855 | 2,371 | +28% | 0 | 0 | — |
case-15 | fail→fail | 6,260 | 12,803 | +105% | 1 | 1 | 0% | 885 | 2,928 | +231% | 0 | 0 | — |
case-16 | fail→fail | 15,026 | 11,133 | -26% | 1 | 1 | 0% | 1,924 | 2,779 | +44% | 0 | 0 | — |
case-17 | pass→pass | 12,588 | 13,272 | +5% | 1 | 1 | 0% | 1,827 | 3,032 | +66% | 0 | 0 | — |
case-18 | fail→pass | 11,373 | 10,069 | -11% | 1 | 1 | 0% | 1,838 | 2,532 | +38% | 0 | 0 | — |
case-19 | fail→pass | 11,119 | 12,469 | +12% | 1 | 1 | 0% | 1,576 | 2,802 | +78% | 0 | 0 | — |
case-20 | pass→pass | 14,006 | 12,133 | -13% | 1 | 1 | 0% | 2,332 | 2,815 | +21% | 0 | 0 | — |
case-21 | pass→pass | 12,419 | 9,579 | -23% | 1 | 1 | 0% | 1,538 | 2,412 | +57% | 0 | 0 | — |
case-22 | pass→pass | 12,220 | 7,900 | -35% | 1 | 1 | 0% | 1,687 | 2,322 | +38% | 0 | 0 | — |
case-23 | pass→pass | 13,194 | 22,099 | +67% | 1 | 1 | 0% | 2,086 | 4,637 | +122% | 0 | 0 | — |
case-24 | pass→pass | 15,697 | 16,373 | +4% | 1 | 1 | 0% | 2,223 | 3,515 | +58% | 0 | 0 | — |
case-25 | pass→pass | 3,271 | 5,790 | +77% | 1 | 1 | 0% | 497 | 1,921 | +287% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 25 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.