Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Define a North Star Metric (NSM) and its input metric tree, with leading indicators, anti-metrics, and counter-metrics. Includes a Python tool that renders the metric tree as a Mermaid diagram.
.claude/skills/borghei-north-star-metric/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 79% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 109% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 92% | 0% |
| case-06 | ✓→✓ | = Same ✓ | 124% | 0% |
A North Star Metric (NSM) is the single number that best represents the value your product delivers to its customers. Sean Ellis popularized the framing; Amplitude codified the playbook; Lean Analytics calls a related concept the "One Metric That Matters" (OMTM). The NSM is one number, not a dashboard. Its job is to align the entire team -- engineering, marketing, sales, support -- on a shared definition of "we won this quarter."
This skill produces a complete NSM specification: the NSM itself, 3-5 input metrics the team can directly influence, the leading indicators that move days or weeks before the inputs, the anti-metrics (things that must NOT move in the wrong direction), and counter-metrics that guard against gaming. The Python tool (metric_tree_builder.py) emits the spec as JSON, Markdown, or a Mermaid tree diagram for a README or Confluence page.
This is the first artifact a team should produce after defining strategy and before writing OKRs. Once the NSM is set, OKRs map directly to moving the input metrics, and roadmaps justify themselves by which input metric they target.
When NOT to use: very early-stage discovery (use discovery/ first — you don't yet know what value you deliver); pure infrastructure work with an indirect user-value chain; before the org has aligned on strategy (the NSM exposes disagreement but does not resolve it).
Before defining the NSM, confirm these inputs. If any is unknown or vague, ASK — do not assume:
Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
bashpython scripts/metric_tree_builder.py --input nsm_spec.json --format mermaid # render the metric tree python scripts/metric_tree_builder.py --demo --format markdown # worked SaaS productivity NSM
Load the reference that matches the task — keep this file lean and pull detail on demand:
metric_tree_builder.py reference (flags, input JSON, Mermaid sample), troubleshooting, and success criteria. Read when selecting an NSM or building the tree.In Scope: NSM selection across 5 archetypes; input-metric tree decomposition with explicit math; leading-indicator selection per input; anti-/counter-metric definition with thresholds; the Python rendering tool; handoff to OKR drafting and roadmap prioritization.
Out of Scope: building actual analytics dashboards (BI tools — this produces the spec); statistical experiment design (discovery/brainstorm-experiments/); financial/revenue forecasting (finance/); OKR drafting (brainstorm-okrs/); data quality validation (data-analytics/).
Caveats: an NSM exposes strategic disagreement but does not resolve it — escalate the strategy decision, not the metric debate. Pure financial outputs (revenue, ARR) are usually too lagging; pick a customer-value proxy that revenue follows from. The NSM aligns; teams still need component metrics for diagnostics. A team without instrumentation cannot operate against an NSM — spend on telemetry first.
| Integration | Direction | Description | |-------------|-----------|-------------| | discovery/brainstorm-ideas/ | Receives from | Opportunity discovery defines what value to deliver; NSM measures it | | discovery/identify-assumptions/ | Receives from | NSM candidates surface assumptions about what customers value | | execution/brainstorm-okrs/ | Feeds into | NSM becomes the quarterly Objective; inputs become Key Results | | execution/outcome-roadmap/ | Feeds into | Roadmap items justify themselves by which input metric they target | | execution/prioritization-frameworks/ | Pairs with | NSM impact is one of the scoring criteria (e.g., RICE Impact, Weighted) | | execution/status-update-generator/ | Feeds into | NSM and input movements feature in Highlights of weekly updates | | data-analytics/ (domain) | Pairs with | NSM spec becomes the schema for dashboards and event taxonomies | | executive-reporting/ (senior-pm) | Feeds into | Monthly board packets lead with NSM trend |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 34,667 | 34,431 | -1% | 1 | 1 | 0% | 5,648 | 7,681 | +36% | 0 | 0 | — |
case-02 | fail→fail | 31,031 | 34,202 | +10% | 1 | 1 | 0% | 5,329 | 7,521 | +41% | 0 | 0 | — |
case-03 | fail→fail | 25,518 | 26,080 | +2% | 1 | 1 | 0% | 4,601 | 6,238 | +36% | 0 | 0 | — |
case-04 | fail→pass | 18,150 | 28,250 | +56% | 1 | 1 | 0% | 2,769 | 6,501 | +135% | 0 | 0 | — |
case-05 | pass→pass | 16,471 | 18,353 | +11% | 1 | 1 | 0% | 2,397 | 4,594 | +92% | 0 | 0 | — |
case-06 | pass→pass | 14,548 | 23,417 | +61% | 1 | 1 | 0% | 2,440 | 5,464 | +124% | 0 | 0 | — |
case-07 | pass→pass | 5,912 | 2,808 | -53% | 1 | 1 | 0% | 864 | 2,050 | +137% | 0 | 0 | — |
case-08 | pass→pass | 4,108 | 2,127 | -48% | 1 | 1 | 0% | 647 | 1,894 | +193% | 0 | 0 | — |
case-09 | fail→pass | 14,577 | 15,950 | +9% | 1 | 1 | 0% | 2,233 | 4,006 | +79% | 0 | 0 | — |
case-10 | fail→fail | 20,166 | 21,627 | +7% | 1 | 1 | 0% | 3,623 | 5,438 | +50% | 0 | 0 | — |
case-11 | fail→fail | 29,160 | 29,809 | +2% | 1 | 1 | 0% | 5,755 | 7,777 | +35% | 0 | 0 | — |
case-12 | fail→fail | 17,437 | 21,513 | +23% | 1 | 1 | 0% | 2,744 | 5,070 | +85% | 0 | 0 | — |
case-13 | fail→fail | 18,799 | 21,612 | +15% | 1 | 1 | 0% | 3,495 | 5,611 | +61% | 0 | 0 | — |
case-14 | pass→pass | 14,043 | 18,485 | +32% | 1 | 1 | 0% | 2,136 | 4,255 | +99% | 0 | 0 | — |
case-15 | pass→pass | 15,019 | 9,082 | -40% | 1 | 1 | 0% | 2,179 | 2,818 | +29% | 0 | 0 | — |
case-16 | fail→pass | 14,952 | 17,677 | +18% | 1 | 1 | 0% | 2,122 | 4,440 | +109% | 0 | 0 | — |
case-17 | fail→fail | 19,237 | 32,559 | +69% | 1 | 1 | 0% | 3,080 | 7,192 | +134% | 0 | 0 | — |
case-18 | pass→pass | 19,395 | 21,334 | +10% | 1 | 1 | 0% | 2,800 | 4,799 | +71% | 0 | 0 | — |
case-19 | pass→pass | 12,301 | 22,553 | +83% | 1 | 1 | 0% | 1,874 | 5,251 | +180% | 0 | 0 | — |
case-20 | pass→pass | 14,426 | 18,290 | +27% | 1 | 1 | 0% | 2,174 | 4,757 | +119% | 0 | 0 | — |
case-21 | pass→pass | 14,218 | 12,403 | -13% | 1 | 1 | 0% | 2,045 | 3,374 | +65% | 0 | 0 | — |
case-22 | pass→pass | 17,302 | 17,910 | +4% | 1 | 1 | 0% | 2,551 | 4,099 | +61% | 0 | 0 | — |
case-23 | pass→pass | 5,249 | 2,422 | -54% | 1 | 1 | 0% | 840 | 1,916 | +128% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +13 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.