Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Extract and risk-rate hidden assumptions in a product brief or PRD. Use when asked to review a product brief for assumptions, audit a PRD for risks, find hidden assumptions, validate product plans, or run an assumption analysis. Produces a prioritised assumption map with confidence and impact scores, recommended validation methods, and critical assumption flags.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 77% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 104% | 0% |
Surface and prioritize the untested assumptions embedded in any product plan before development begins.
Ask the user for these if not provided:
This is the front of the product-decision spine: assumption-mapper → /prd-template → /rice-prioritisation → /roadmap-narrative. It takes a raw idea or brief and hands the next skill one thing: the riskiest assumption, and whether it survived a cheap test. Shared terms (assumption, load-bearing, confidence, provenance) are defined once in docs/craft/product-decisions.md — consult it rather than re-deriving them. Writing a PRD on top of an untested load-bearing assumption is the failure this skill exists to prevent, so run it before /prd-template, not after.
Four phases. Phase 3 is the skill; the rest feed it. Each ends on a completion criterion — don't advance until it's met.
want it?), Feasibility (can we build it?), Viability (will the business sustain it?), Usability (can users actually use it?). The dangerous assumptions are the ones so obvious no one wrote them down. Done when: at least one assumption per lens, and re-reading the brief for the emptiest lens surfaces nothing new.
if it's false?) and confidence (1–5, how sure are we it's true?). Priority = load-bearing − confidence. Tag each fact's provenance (data]/hunch]). Done when: every assumption has both scores and a provenance tag, and the highest-priority one is genuinely the scariest — not the easiest to test.
low-confidence) is the one that can sink the whole plan. Name the cheapest test that could disprove it before a line of code is written (see the disclosed cheap-tests reference for the menu). Done when: the single riskiest assumption is named, with a test that could run this week and a clear "what a fail looks like."
/prd-template must treat as validated-or-open. An unresolved riskiest assumption becomes an Open Question in the PRD, not a silent bet. Done when: the downstream skill could start from this output without re-asking what the risky bet is.
| Assumption | Category | Confidence | Impact | Priority | Validation Method | |------------|----------|------------|--------|----------|-------------------| | assumption] | type] | 1-5] | 1-5] | score] | method] |
Flagged items with detailed validation recommendations]
Detailed recommendations including specific research method, estimated effort, and what the result would change]
Input: "We're building a self-serve onboarding flow to reduce time-to-value for SMB customers."
| Assumption | Category | Confidence | Impact | Priority | Validation Method | |------------|----------|------------|--------|----------|-------------------| | SMB users can complete onboarding without human help | Usability | 2 | 5 | 3 | Unmoderated usability test (n=8) | | Faster onboarding correlates with higher retention | Viability | 3 | 4 | 1 | Cohort analysis of current onboarding times vs. 90-day retention | | The current onboarding is the primary reason for slow time-to-value | Desirability | 2 | 4 | 2 | User interviews with recent churned SMB accounts |
This skill ships with support files — use them when they are available:
references/cheap-tests.md — The Cheap-Test Catalog: Right-Sizing Validation. Apply it while producing the output; it carries the calibration and judgment calls the method summary above compresses.templates/assumption-board.md — a fill-in version of the deliverable with the quality gates inline. Offer it when the user wants to work the document themselves rather than have it generated.Score any output of this skill before handing it over; 32+ is ship-quality.
| Dimension | 0 | 5 | 10 | |---|---|---|---| | Category coverage | Desirability-only — the feasibility and viability assumptions most likely to kill the plan are absent | Three categories populated, but the empty one wasn't re-mined from the brief; coverage is token (one throwaway row) | All four categories populated with substantive rows, with visible digging into whichever category the brief itself neglected | | Scoring discipline | Confidence/impact numbers arbitrary or missing; priority arithmetic inconsistent; no critical flags | Scores present and Priority = Impact − Confidence holds, but confidence is inflated for unchallenged assumptions and critical flags applied selectively | Scores defensible (unchallenged ≠ high confidence), arithmetic consistent including negative priorities left visible, and the CRITICAL flag applied mechanically at Impact 4+ / Confidence ≤2 — even to assumptions the team likes | | Validation method fit | "User interviews" (or "do research") pasted into every row | Methods vary but several are mismatched to the assumption type, missing sample sizes, or unpriced | Each method matched to the assumption (data audit, backtest, fake door, desk check, spike…) with sample size and effort; untestable assumptions flagged unknowable and converted to owned risks, not given fake tests | | Decision leverage | Top-3 list missing, or tests whose outcome would change nothing | Top 3 named with effort, but "what the result changes" is vague or the tests validate comfortable assumptions over dangerous ones | Top 3 are the highest-priority testable assumptions, each with effort, a pre-committed threshold where relevant, and a concrete decision the result would change |
Other measured skills in the registry, with their headline benchmark lift.