Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Define and design a product metrics dashboard with key metrics, data sources, visualization types, and alert thresholds. Use when creating a metrics dashboard, defining KPIs, setting up product analytics, or building a data monitoring plan.
.claude/skills/phuryn-metrics-dashboard/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 138% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 118% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 116% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 162% | 0% |
Design a comprehensive product metrics dashboard with the right metrics, visualizations, and alert thresholds.
You are designing a metrics dashboard for $ARGUMENTS.
If the user provides files (existing dashboards, analytics data, OKRs, or strategy docs), read them first.
Metrics vs KPIs vs NSM: Metrics = all measurable things. KPIs = a few key quantitative metrics tracked over a longer period. North Star Metric = a single customer-centric KPI that is a leading indicator of business success.
4 criteria for a good metric (Ben Yoskovitz, Lean Analytics): (1) Understandable — creates a common language. (2) Comparative — over time, not a snapshot. (3) Ratio or Rate — more revealing than whole numbers. (4) Behavior-changing — the Golden Rule: "If a metric won't change how you behave, it's a bad metric."
8 metric types: Vanity vs Actionable (only actionable metrics change behavior), Qualitative vs Quantitative (WHAT vs WHY — you need both; never stop talking to customers), Exploratory vs Reporting (explore data to uncover unexpected insights), Lagging vs Leading (leading indicators enable faster learning cycles, e.g. customer complaints predict churn).
5 action steps: (1) Audit metrics against the 4 good-metric criteria. (2) Update dashboards — ensure all key metrics are good ones. (3) Identify vanity metrics — be careful how you use them. (4) Classify leading vs lagging indicators. (5) Pick one problem and dig deep into the data.
For case studies and more detail: Are You Tracking the Right Metrics? by Ben Yoskovitz
North Star Metric: The single metric that best captures core value delivery
Input Metrics (3-5): The levers that drive the North Star
Health Metrics: Guardrails that ensure overall product health
Business Metrics: Revenue, cost, and unit economics
| Metric | Definition | Data Source | Visualization | Target | Alert Threshold | |---|---|---|---|---|---| | Name] | Exact calculation: numerator/denominator, time window] | Where the data comes from] | Line chart / Bar / Number / Funnel] | Goal value] | When to trigger an alert] |
┌─────────────────────────────────────────────┐ │ NORTH STAR: [Metric] — [Current Value] │ │ Trend: [↑/↓ X% vs last period] │ ├──────────────────┬──────────────────────────┤ │ Input Metric 1 │ Input Metric 2 │ │ [Sparkline] │ [Sparkline] │ ├──────────────────┼──────────────────────────┤ │ Input Metric 3 │ Input Metric 4 │ │ [Sparkline] │ [Sparkline] │ ├──────────────────┴──────────────────────────┤ │ HEALTH: [Latency] [Error Rate] [NPS] │ ├─────────────────────────────────────────────┤ │ BUSINESS: [MRR] [CAC] [LTV] [Churn] │ └─────────────────────────────────────────────┘
Think step by step. Save the dashboard specification as a markdown document.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | pass→pass | 13,206 | 27,554 | +109% | 1 | 1 | 0% | 2,351 | 6,452 | +174% | 0 | 0 | — |
case-13 | pass→pass | 14,148 | 27,708 | +96% | 1 | 1 | 0% | 2,365 | 6,316 | +167% | 0 | 0 | — |
case-01 | pass→pass | 32,447 | 31,761 | -2% | 1 | 1 | 0% | 6,248 | 7,482 | +20% | 0 | 0 | — |
case-02 | fail→pass | 32,246 | 31,499 | -2% | 1 | 1 | 0% | 6,139 | 7,479 | +22% | 0 | 0 | — |
case-03 | pass→pass | 13,756 | 23,776 | +73% | 1 | 1 | 0% | 2,366 | 5,587 | +136% | 0 | 0 | — |
case-04 | pass→pass | 12,437 | 19,398 | +56% | 1 | 1 | 0% | 2,204 | 5,727 | +160% | 0 | 0 | — |
case-05 | fail→pass | 12,986 | 24,334 | +87% | 1 | 1 | 0% | 2,020 | 4,801 | +138% | 0 | 0 | — |
case-06 | fail→pass | 20,529 | 27,209 | +33% | 1 | 1 | 0% | 2,907 | 6,341 | +118% | 0 | 0 | — |
case-11 | fail→pass | 16,521 | 24,664 | +49% | 1 | 1 | 0% | 2,646 | 5,722 | +116% | 0 | 0 | — |
case-07 | pass→pass | 13,220 | 25,572 | +93% | 1 | 1 | 0% | 2,500 | 5,855 | +134% | 0 | 0 | — |
case-08 | pass→pass | 17,796 | 29,780 | +67% | 1 | 1 | 0% | 2,334 | 6,544 | +180% | 0 | 0 | — |
case-09 | fail→pass | 14,529 | 34,458 | +137% | 1 | 1 | 0% | 2,325 | 6,086 | +162% | 0 | 0 | — |
case-10 | pass→pass | 15,113 | 22,944 | +52% | 1 | 1 | 0% | 2,499 | 5,197 | +108% | 0 | 0 | — |
case-14 | fail→pass | 8,989 | 28,900 | +222% | 1 | 1 | 0% | 1,588 | 6,792 | +328% | 0 | 0 | — |
case-15 | pass→pass | 19,892 | 28,287 | +42% | 1 | 1 | 0% | 2,852 | 5,392 | +89% | 0 | 0 | — |
case-16 | pass→pass | 15,543 | 51,147 | +229% | 1 | 1 | 0% | 1,777 | 4,940 | +178% | 0 | 0 | — |
case-17 | pass→pass | 12,405 | 19,513 | +57% | 1 | 1 | 0% | 2,078 | 5,085 | +145% | 0 | 0 | — |
case-18 | pass→pass | 12,340 | 25,471 | +106% | 1 | 1 | 0% | 2,112 | 5,718 | +171% | 0 | 0 | — |
case-19 | pass→pass | 12,013 | 14,672 | +22% | 1 | 1 | 0% | 2,118 | 3,402 | +61% | 0 | 0 | — |
case-20 | fail→pass | 15,483 | 17,517 | +13% | 1 | 1 | 0% | 2,070 | 4,300 | +108% | 0 | 0 | — |
case-21 | fail→pass | 14,698 | 32,230 | +119% | 1 | 1 | 0% | 2,543 | 7,403 | +191% | 0 | 0 | — |
case-22 | fail→fail | 12,197 | 18,178 | +49% | 1 | 1 | 0% | 2,944 | 5,109 | +74% | 0 | 0 | — |
case-23 | pass→pass | 10,332 | 11,544 | +12% | 1 | 1 | 0% | 2,005 | 3,668 | +83% | 0 | 0 | — |
case-24 | pass→pass | 8,820 | 7,918 | -10% | 1 | 1 | 0% | 1,720 | 2,688 | +56% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +33 percentage points is the difference between those two pass rates over the 24 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.