Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when a user shares a raw product idea or problem statement and wants a structured pipeline from clarifying questions through deep research, a PRD, and a phased execution plan — written as files they can take forward. This is an end-to-end idea-to-build-ready workflow, not a quick-feedback skill (use `idea-refine` for that) and not a PRD editor (use `product-management` for that). Trigger phrases include "I have an idea for…", "help me build X", "validate and plan this concept", "what should
.claude/skills/bilal140202-idea-os/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 96% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 177% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-20 | ✓→✗ | ▼ Worse | -35% | 0% |
An operating system for turning a raw idea into a build-ready plan. Takes a rough problem statement and produces four files: clarifying questions, deep research, a PRD, and a phased execution plan with platform/stack picks, a user-journey diagram, and kill criteria.
Four files in the current working directory (use pwd). Filenames are canonical.
questions.md — sharp clarifying questions (skip if the brief is already dense or run is autonomous)research.md — market, competitors, SWOT, JTBD, distribution, risks, insightsPRD.md — problem, users, scope, non-goals, metrics, solution shapeplan.md — user journey + mermaid, platform + stack, phased build, kill criteria, next stepsDo not skip or reorder. Weak inputs produce weak outputs; each phase feeds the next.
Classify the idea on two axes. Depth of research, PRD, plan, and question-count scale with tier; vocabulary scales with sophistication.
Idea tier (T1/T2/T3) — complexity:
Sophistication (S1/S2/S3) — builder vocabulary:
State your classification in one line (e.g. "T2 · S2 — moderate SaaS, builder has shipped before") before proceeding. If ambiguous, ask at most one question.
See references/clarifying-questions.md for the question bank tiered to these levels.
Write questions.md using assets/questions-template.md. Pick 4–18 questions (count scales with complexity) matched to the idea type — SaaS, marketplace, consumer, AI wrapper, dev tool, hardware, content. Group into three buckets: Who and Pain · Scope and Wedge · Constraints and Goals.
Every question must be actionable — the answer must change what you build. Generic questions ("who is your user?") are rejected; force specifics ("describe the last time your target user hit this pain — what did they do instead?").
After writing, stop and wait for answers. Do not research yet.
Skip questions.md entirely if the brief is already dense; state assumptions and proceed.
Write research.md using assets/research-template.md. Use WebSearch + WebFetch for real signal. Never hallucinate numbers; flag [assumption] or cut.
Pre-flight checklist (a floor):
Sections required (depth scales with complexity):
references/research-frameworks.mdreferences/research-frameworks.md#jtbdreferences/market-analysis.mdreferences/competitor-analysis.mdreferences/research-frameworks.md#swotreferences/distribution.mdWrite PRD.md using assets/PRD-template.md. Anatomy + tier-scaled sections + anti-patterns in references/prd-structure.md. Non-goals section is mandatory — it's where bad PRDs die.
Write plan.md using assets/plan-template.md. Must include:
references/user-journey.mdreferences/platform-recommendation.mdreferences/stack-recommendation.mdreferences/mvp-slicing.mdreferences/metrics-framework.mdWhen there's no live user (batch run, test, dry-run, or the user said "just proceed"): still write questions.md to make decisions legible, then simulate plausible answers consistent with a stated persona at the chosen tier. At the top of research.md, add an Assumptions in lieu of answers block listing each assumed answer and flagging the load-bearing ones. The reader should see exactly what you decided for them.
pwd unless user specifies[PRD](PRD.md))[assumption]More in references/prd-structure.md (PRD anti-patterns) and references/mvp-slicing.md (MVP traps).
examples/example-notion-for-dog-trainers.md is a reference-only walkthrough showing the quality bar. Do not load it unless you need calibration — it's 500+ lines.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 36,048 | 24,071 | -33% | 1 | 1 | 0% | 6,242 | 5,728 | -8% | 0 | 0 | — |
case-02 | fail→pass | 34,030 | 38,440 | +13% | 1 | 1 | 0% | 5,801 | 7,823 | +35% | 0 | 0 | — |
case-03 | fail→fail | 39,364 | 40,961 | +4% | 1 | 1 | 0% | 6,228 | 7,817 | +26% | 0 | 0 | — |
case-04 | fail→pass | 9,189 | 11,575 | +26% | 1 | 1 | 0% | 1,763 | 3,458 | +96% | 0 | 0 | — |
case-05 | fail→pass | 8,049 | 8,499 | +6% | 1 | 1 | 0% | 1,062 | 2,939 | +177% | 0 | 0 | — |
case-06 | fail→fail | 10,511 | 10,748 | +2% | 1 | 1 | 0% | 1,538 | 3,155 | +105% | 0 | 0 | — |
case-07 | pass→pass | 8,250 | 10,207 | +24% | 1 | 1 | 0% | 1,296 | 3,153 | +143% | 0 | 0 | — |
case-08 | fail→fail | 28,105 | 44,920 | +60% | 1 | 1 | 0% | 4,478 | 7,795 | +74% | 0 | 0 | — |
case-09 | fail→fail | 21,502 | 12,361 | -43% | 1 | 1 | 0% | 3,367 | 2,258 | -33% | 0 | 0 | — |
case-10 | pass→pass | 24,737 | 28,017 | +13% | 1 | 1 | 0% | 3,813 | 5,722 | +50% | 0 | 0 | — |
case-11 | pass→pass | 23,981 | 31,949 | +33% | 1 | 1 | 0% | 3,685 | 6,487 | +76% | 0 | 0 | — |
case-12 | fail→fail | 22,961 | 40,996 | +79% | 1 | 1 | 0% | 3,505 | 7,747 | +121% | 0 | 0 | — |
case-13 | pass→pass | 18,209 | 19,477 | +7% | 1 | 1 | 0% | 2,949 | 4,704 | +60% | 0 | 0 | — |
case-14 | fail→fail | 20,557 | 29,413 | +43% | 1 | 1 | 0% | 3,408 | 6,212 | +82% | 0 | 0 | — |
case-15 | pass→pass | 13,058 | 15,355 | +18% | 1 | 1 | 0% | 2,315 | 3,795 | +64% | 0 | 0 | — |
case-16 | pass→pass | 21,794 | 30,464 | +40% | 1 | 1 | 0% | 3,437 | 6,884 | +100% | 0 | 0 | — |
case-17 | fail→pass | 24,831 | 34,153 | +38% | 1 | 1 | 0% | 4,132 | 7,022 | +70% | 0 | 0 | — |
case-18 | pass→pass | 10,132 | 11,865 | +17% | 1 | 1 | 0% | 1,637 | 3,717 | +127% | 0 | 0 | — |
case-19 | pass→pass | 17,724 | 12,720 | -28% | 1 | 1 | 0% | 2,609 | 3,430 | +31% | 0 | 0 | — |
case-20 | pass→fail | 30,936 | 14,835 | -52% | 1 | 1 | 0% | 5,869 | 3,827 | -35% | 0 | 0 | — |
case-21 | pass→pass | 12,459 | 18,497 | +48% | 1 | 1 | 0% | 2,423 | 4,909 | +103% | 0 | 0 | — |
case-22 | pass→pass | 9,863 | 16,605 | +68% | 1 | 1 | 0% | 1,681 | 4,445 | +164% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.