Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build an Opportunity Solution Tree (OST) to structure product discovery — map a desired outcome to opportunities, solutions, and experiments. Based on Teresa Torres' Continuous Discovery Habits. Use when structuring discovery work, mapping opportunities to solutions, or deciding what to build next.
.claude/skills/phuryn-opportunity-solution-tree/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 78% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 142% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 24% | 0% |
A visual framework for structuring continuous product discovery. Connects a desired outcome to customer opportunities, possible solutions, and experiments to validate them.
The Opportunity Solution Tree (Teresa Torres, Continuous Discovery Habits) is the backbone of modern product discovery. It prevents teams from jumping to solutions by forcing them to first map the opportunity space.
Structure (4 levels):
Key principles:
You are helping a product team build an Opportunity Solution Tree for $ARGUMENTS.
Think step by step. Save as markdown if substantial.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 13,462 | 21,376 | +59% | 1 | 1 | 0% | 2,496 | 4,437 | +78% | 0 | 0 | — |
case-02 | fail→pass | 16,945 | 17,568 | +4% | 1 | 1 | 0% | 3,124 | 4,215 | +35% | 0 | 0 | — |
case-11 | pass→pass | 10,150 | 16,110 | +59% | 1 | 1 | 0% | 1,727 | 2,949 | +71% | 0 | 0 | — |
case-16 | fail→pass | 10,375 | 17,919 | +73% | 1 | 1 | 0% | 1,370 | 3,313 | +142% | 0 | 0 | — |
case-03 | fail→pass | 7,657 | 5,973 | -22% | 1 | 1 | 0% | 1,541 | 1,744 | +13% | 0 | 0 | — |
case-04 | pass→pass | 10,580 | 11,373 | +7% | 1 | 1 | 0% | 1,768 | 2,799 | +58% | 0 | 0 | — |
case-05 | fail→pass | 16,035 | 17,373 | +8% | 1 | 1 | 0% | 2,763 | 3,422 | +24% | 0 | 0 | — |
case-06 | pass→pass | 17,109 | 15,393 | -10% | 1 | 1 | 0% | 2,473 | 3,427 | +39% | 0 | 0 | — |
case-07 | pass→pass | 10,195 | 11,863 | +16% | 1 | 1 | 0% | 1,848 | 2,973 | +61% | 0 | 0 | — |
case-08 | fail→pass | 11,031 | 13,555 | +23% | 1 | 1 | 0% | 1,895 | 2,631 | +39% | 0 | 0 | — |
case-09 | pass→pass | 14,302 | 14,081 | -2% | 1 | 1 | 0% | 2,494 | 3,298 | +32% | 0 | 0 | — |
case-10 | fail→fail | 10,504 | 9,498 | -10% | 1 | 1 | 0% | 1,440 | 2,397 | +66% | 0 | 0 | — |
case-12 | pass→pass | 10,835 | 10,427 | -4% | 1 | 1 | 0% | 1,787 | 2,378 | +33% | 0 | 0 | — |
case-13 | fail→pass | 5,706 | 2,579 | -55% | 1 | 1 | 0% | 1,092 | 1,398 | +28% | 0 | 0 | — |
case-14 | pass→pass | 9,844 | 3,314 | -66% | 1 | 1 | 0% | 2,100 | 1,392 | -34% | 0 | 0 | — |
case-15 | fail→pass | 8,760 | 8,744 | -0% | 1 | 1 | 0% | 1,465 | 2,321 | +58% | 0 | 0 | — |
case-17 | pass→pass | 15,410 | 10,718 | -30% | 1 | 1 | 0% | 2,006 | 2,574 | +28% | 0 | 0 | — |
case-18 | pass→pass | 14,717 | 11,150 | -24% | 1 | 1 | 0% | 2,052 | 2,776 | +35% | 0 | 0 | — |
case-19 | pass→fail | 11,751 | 19,753 | +68% | 1 | 1 | 0% | 2,289 | 3,459 | +51% | 0 | 0 | — |
case-20 | pass→pass | 11,476 | 12,975 | +13% | 1 | 1 | 0% | 1,945 | 2,931 | +51% | 0 | 0 | — |
case-21 | pass→pass | 3,844 | 2,192 | -43% | 1 | 1 | 0% | 638 | 1,267 | +99% | 0 | 0 | — |
case-22 | pass→pass | 13,273 | 8,358 | -37% | 1 | 1 | 0% | 1,698 | 1,944 | +14% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.