Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when a founder has a rough product idea and wants autonomous deep validation, market and competitor research, and an evidence-based MVP decision with minimal back-and-forth.
.claude/skills/bilal140202-idea-validation-autopilot/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 63% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 83% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 105% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 161% | 0% |
Turn a rough idea into an evidence-backed build decision in one run.
This skill is a single orchestrator for:
Default behavior favors action over over-analysis:
Use this skill when:
Do not use this skill when:
If user context is missing, proceed with defaults instead of blocking:
speed-to-learning > polishnear-zero external spendsolo builder or very small teamone focused discovery cycleOnly ask questions when missing data would invalidate the result (for example: unclear target user or regulated domain).
Copy this checklist and track progress:
mdProgress - [ ] Step 1: Normalize idea into problem hypothesis - [ ] Step 2: Run 4 parallel research tracks - [ ] Step 3: Grade evidence quality and resolve contradictions - [ ] Step 4: Produce decision scorecard and verdict - [ ] Step 5: Define MVP scope and exclusions - [ ] Step 6: Define first experiments and stop rules - [ ] Step 7: Deliver final report using template
Convert raw idea into this structure:
If unclear, propose your best assumption and mark it explicitly.
Dispatch four independent subagents (or equivalent parallel workers).
Use evidence tiers:
Tier A: behavioral or monetary signal (payment, waitlist intent with commitment, repeated real usage)Tier B: strong secondary evidence (credible reports, robust competitor/user data)Tier C: weak signal (opinions, generic trend articles, unsupported claims)Rules:
Score 0-100 using weighted dimensions:
| Dimension | Weight | | --- | ---: | | Problem severity and frequency | 25 | | Distribution reachability | 20 | | Willingness-to-pay potential | 20 | | MVP speed/feasibility | 20 | | Strategic differentiation | 15 |
Scoring rules (fixed):
0..100weighted_i = score_i * weight_i / 100total_score = round(sum(weighted_i), 1)total_score using the bands belowVerdict bands:
80-100: Build now60-79: Validate-first (run targeted tests before building)40-59: Pivot<40: DropUse strict scope slicing:
Must: smallest set proving core valueShould: useful but deferrableWon't (now): explicitly excluded featuresOutput a 2-week implementation target:
For top risks, define:
Keep experiments cheap and fast. Favor reversible steps.
Use assets/final-report-template.md.
Output path rules:
reports/ does not exist, create it first (mkdir -p reports)reports/YYYY-MM-DD-<idea-slug>-idea-validation.mdRequired output qualities:
Adapt to available tools:
reports/If one tool is unavailable, continue with the best fallback and document the limitation in assumptions.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 26,674 | 35,069 | +31% | 1 | 1 | 0% | 4,286 | 6,987 | +63% | 0 | 0 | — |
case-02 | fail→pass | 26,771 | 31,924 | +19% | 1 | 1 | 0% | 4,460 | 6,460 | +45% | 0 | 0 | — |
case-03 | fail→fail | 24,709 | 24,371 | -1% | 1 | 1 | 0% | 4,161 | 5,554 | +33% | 0 | 0 | — |
case-04 | pass→pass | 21,334 | 17,239 | -19% | 1 | 1 | 0% | 3,667 | 4,236 | +16% | 0 | 0 | — |
case-05 | pass→pass | 19,573 | 17,919 | -8% | 1 | 1 | 0% | 3,959 | 5,385 | +36% | 0 | 0 | — |
case-06 | pass→pass | 28,226 | 19,688 | -30% | 1 | 1 | 0% | 4,992 | 5,151 | +3% | 0 | 0 | — |
case-07 | fail→pass | 19,365 | 29,904 | +54% | 1 | 1 | 0% | 2,930 | 5,373 | +83% | 0 | 0 | — |
case-08 | fail→pass | 21,336 | 33,851 | +59% | 1 | 1 | 0% | 3,512 | 7,198 | +105% | 0 | 0 | — |
case-09 | fail→pass | 16,755 | 5,003 | -70% | 1 | 1 | 0% | 831 | 2,168 | +161% | 0 | 0 | — |
case-10 | fail→pass | 10,557 | 3,972 | -62% | 1 | 1 | 0% | 1,799 | 1,937 | +8% | 0 | 0 | — |
case-11 | fail→pass | 6,113 | 3,276 | -46% | 1 | 1 | 0% | 1,043 | 1,983 | +90% | 0 | 0 | — |
case-12 | pass→pass | 7,637 | 4,933 | -35% | 1 | 1 | 0% | 1,283 | 2,166 | +69% | 0 | 0 | — |
case-13 | fail→pass | 16,177 | 16,220 | +0% | 1 | 1 | 0% | 2,730 | 4,285 | +57% | 0 | 0 | — |
case-14 | pass→pass | 9,934 | 6,309 | -36% | 1 | 1 | 0% | 1,656 | 2,337 | +41% | 0 | 0 | — |
case-15 | fail→pass | 19,867 | 32,539 | +64% | 1 | 1 | 0% | 3,314 | 6,655 | +101% | 0 | 0 | — |
case-16 | fail→pass | 9,799 | 3,122 | -68% | 1 | 1 | 0% | 1,738 | 1,960 | +13% | 0 | 0 | — |
case-17 | fail→pass | 10,721 | 17,034 | +59% | 1 | 1 | 0% | 1,652 | 4,058 | +146% | 0 | 0 | — |
case-18 | pass→pass | 15,707 | 14,736 | -6% | 1 | 1 | 0% | 2,863 | 3,693 | +29% | 0 | 0 | — |
case-19 | pass→pass | 13,810 | 6,471 | -53% | 1 | 1 | 0% | 1,964 | 2,350 | +20% | 0 | 0 | — |
case-20 | pass→pass | 10,647 | 5,796 | -46% | 1 | 1 | 0% | 1,607 | 2,154 | +34% | 0 | 0 | — |
case-21 | fail→fail | 14,654 | 8,895 | -39% | 1 | 1 | 0% | 2,241 | 2,836 | +27% | 0 | 0 | — |
case-22 | pass→pass | 11,344 | 4,567 | -60% | 1 | 1 | 0% | 1,672 | 2,076 | +24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.