Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Transform output-based feature lists into outcome-driven Now/Next/Later roadmaps using the "so what?" technique. Use when converting a feature-list roadmap to outcomes, communicating product strategy, or running quarterly planning.
.claude/skills/borghei-outcome-roadmap/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 187% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 288% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 39% | 0% |
The agent transforms output-based roadmaps ("build feature X") into outcome-driven roadmaps ("enable customers to achieve Y") using the "so what?" technique and Now/Next/Later framing. It produces roadmaps that communicate strategy and measurable impact, not just feature lists and dates.
Before transforming the roadmap, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
bashpython scripts/roadmap_transformer.py --input roadmap.json # transform a roadmap python scripts/roadmap_transformer.py --input roadmap.json --format markdown python scripts/roadmap_transformer.py --demo # run on built-in demo data
Each input initiative needs title, description, quarter (format "Q1-4] YYYY"), and type (feature/improvement/infrastructure). The tool emits outcome-statement templates — fill the placeholders with real customer and business data.
assets/outcome_roadmap_template.md — roadmap document template with Now/Next/Later sections.In Scope: transforming output-based feature lists into outcome-driven items, Now/Next/Later classification by quarter-to-current-date distance, "so what?" chain generation, strategic-question and metric suggestions by initiative type, markdown/text/JSON report output grouped by horizon.
Out of Scope: feature prioritization or scoring (execution/prioritization-frameworks/), sprint-level planning or capacity allocation (scrum-master/), product strategy or vision definition (roadmaps communicate strategy, they don't create it), cross-team dependency management (program-manager/).
Important Caveats: outcome roadmaps require a cultural shift — teams used to date-driven lists need coaching on commitment levels; the tool generates outcome-statement templates, not finished outcomes; Later items intentionally lack detailed metrics, and adding false precision undermines credibility.
| Integration | Direction | Description | |------------|-----------|-------------| | execution/brainstorm-okrs/ | Receives from | OKR key results become success metrics for Now/Next roadmap items | | execution/prioritization-frameworks/ | Receives from | RICE/ICE scores inform which initiatives move to Now vs. Next vs. Later | | execution/create-prd/ | Feeds into | Now items with validated outcomes become PRD candidates | | discovery/brainstorm-experiments/ | Receives from | Experiment results validate demand for Next/Later items, promoting them to Now | | senior-pm/ | Receives from | Portfolio strategic priorities influence roadmap horizon placement | | scrum-master/ | Receives from | Sprint capacity data determines how many Now items the team can support |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 12,545 | 19,892 | +59% | 1 | 1 | 0% | 2,133 | 4,443 | +108% | 0 | 0 | — |
case-02 | fail→fail | 16,033 | 22,139 | +38% | 1 | 1 | 0% | 2,710 | 4,928 | +82% | 0 | 0 | — |
case-03 | fail→fail | 15,971 | 22,495 | +41% | 1 | 1 | 0% | 2,630 | 4,577 | +74% | 0 | 0 | — |
case-04 | fail→pass | 10,044 | 21,813 | +117% | 1 | 1 | 0% | 1,607 | 4,619 | +187% | 0 | 0 | — |
case-05 | fail→pass | 3,413 | 6,032 | +77% | 1 | 1 | 0% | 552 | 2,141 | +288% | 0 | 0 | — |
case-06 | pass→pass | 18,369 | 16,160 | -12% | 1 | 1 | 0% | 2,529 | 3,387 | +34% | 0 | 0 | — |
case-07 | pass→pass | 9,739 | 9,567 | -2% | 1 | 1 | 0% | 1,395 | 2,567 | +84% | 0 | 0 | — |
case-08 | pass→pass | 5,955 | 8,494 | +43% | 1 | 1 | 0% | 829 | 2,403 | +190% | 0 | 0 | — |
case-09 | fail→fail | 5,944 | 12,511 | +110% | 1 | 1 | 0% | 998 | 3,210 | +222% | 0 | 0 | — |
case-10 | fail→pass | 18,754 | 19,780 | +5% | 1 | 1 | 0% | 3,264 | 4,223 | +29% | 0 | 0 | — |
case-11 | fail→fail | 17,301 | 21,093 | +22% | 1 | 1 | 0% | 2,665 | 4,358 | +64% | 0 | 0 | — |
case-12 | pass→fail | 6,550 | 18,674 | +185% | 1 | 1 | 0% | 1,155 | 4,274 | +270% | 0 | 0 | — |
case-13 | fail→pass | 10,795 | 6,850 | -37% | 1 | 1 | 0% | 1,842 | 2,260 | +23% | 0 | 0 | — |
case-14 | fail→pass | 8,934 | 5,068 | -43% | 1 | 1 | 0% | 1,358 | 1,890 | +39% | 0 | 0 | — |
case-15 | pass→pass | 15,221 | 12,753 | -16% | 1 | 1 | 0% | 2,596 | 3,394 | +31% | 0 | 0 | — |
case-21 | pass→pass | 10,440 | 4,335 | -58% | 1 | 1 | 0% | 1,521 | 1,810 | +19% | 0 | 0 | — |
case-16 | fail→pass | 7,886 | 10,751 | +36% | 1 | 1 | 0% | 1,324 | 2,535 | +91% | 0 | 0 | — |
case-17 | pass→pass | 9,419 | 14,503 | +54% | 1 | 1 | 0% | 1,374 | 3,295 | +140% | 0 | 0 | — |
case-18 | pass→pass | 13,860 | 10,815 | -22% | 1 | 1 | 0% | 2,006 | 2,681 | +34% | 0 | 0 | — |
case-19 | fail→pass | 8,393 | 4,583 | -45% | 1 | 1 | 0% | 1,225 | 1,795 | +47% | 0 | 0 | — |
case-20 | pass→pass | 14,141 | 16,216 | +15% | 1 | 1 | 0% | 2,063 | 3,540 | +72% | 0 | 0 | — |
case-22 | pass→pass | 15,485 | 13,559 | -12% | 1 | 1 | 0% | 2,449 | 3,121 | +27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.