Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Diagnose why a product metric changed (dropped, spiked, or plateaued) by orchestrating breakdowns, actors, paths, lifecycle, retention, and annotations queries. Use when the user reports an anomaly, asks \"why did X change?\", or needs root-cause analysis for a trend, funnel, retention, stickiness, or lifecycle metric.
.claude/skills/kunanonj-cursor-plugin-posthog-investigate-metric/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 35 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 139% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-12 | ✓→✗ | ▼ Worse | 55% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 110% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 30% | 0% |
For "why did X change?" questions about a saved insight, dashboard tile, or pasted query. Don't load this skill for plain "what is X?" questions — only when there's an observed change to explain.
Targets PostHog MCP v2. Typed query tools accept the query body directly — pass kind, series, dateRange as top-level fields, do not wrap in InsightVizNode.
| Tool | Purpose | | -------------------------------- | ------------------------------------------------ | | posthog:query-trends | Trends (count over time) | | posthog:query-funnel | Funnels (multi-step conversion) | | posthog:query-retention | Retention (cohort return rates) | | posthog:query-stickiness | Stickiness (active days per user) | | posthog:query-lifecycle | Lifecycle (new/returning/resurrecting/dormant) | | posthog:query-paths | Paths (navigation flow) | | posthog:query-trends-actors | Users behind a trend bucket (trends source only) | | posthog:execute-sql | HogQL — when no typed tool fits | | posthog:read-data-schema | Discover events, properties, sample values | | posthog:insight-get / -query | Fetch a saved insight's metadata / data |
Plus the standard PostHog tools the playbooks reference by name (feature-flag-get-all, experiment-get-all, annotations-list, query-error-tracking-issues-list, query-logs, query-session-recordings-list, cohorts-list/-create, annotation-create, insight-create).
compare_to_prior_periods.py — auto-detectsinterval and compares recent values to the natural cycle (day-of-week, hour-of-week, or sequential). Use to resolve step 2.2 cheaply.
breakdown_attribution.py — ranks breakdownsegments by absolute delta and flags offsetting moves.
bashpython3 scripts/compare_to_prior_periods.py < query_result.json WINDOW=7 python3 scripts/breakdown_attribution.py < breakdown_result.json
Read query.kind from the source the user pointed at:
short_id): posthog:insight-get → query.kind. Useposthog:insight-query if you also need the numbers.
kind directly.| kind | Playbook | | ----------------- | ------------------------------------------------------------- | | TrendsQuery | trend-playbook.md | | FunnelsQuery | funnel-playbook.md | | RetentionQuery | retention-playbook.md | | StickinessQuery | stickiness-playbook.md | | LifecycleQuery | lifecycle-playbook.md | | PathsQuery | paths-playbook.md | | HogQLQuery | route by what the SQL aggregates (see below) |
If kind === "TrendsQuery" and trendsFilter.display === "BoxPlot", use box-plot-playbook.md — distribution metric, no breakdowns.
For HogQLQuery insights, classify by the SQL's shape: count over time → trend playbook, multi-step conversion → funnel playbook, cohort return → retention playbook. Run the SQL through posthog:execute-sql to get the data, then follow the closest playbook's steps. See HogQL insights in shared-patterns.md.
If the user's question spans multiple kinds, run the playbooks in sequence.
Run the primary tool. Record baseline, current, delta (absolute and %), and the start of the anomaly window.
Widen to 3–4× the user's interval (or use compareFilter: {"compare": true} on TrendsQuery / StickinessQuery; for other kinds run two date ranges). Pipe the widened result through compare_to_prior_periods.py — it flags seasonality, partial right-edge buckets, and real anomalies. If the movement is normal variance, report that and stop.
In rough order of signal:
posthog:feature-flag-get-all → flags with updated_at near the anomaly start.posthog:experiment-get-all → start_date / end_date near the start.posthog:annotations-list → date_marker near the start.git log for the window if the repo is reachable (highest signal when available).Any match is a hypothesis to confirm in the playbook (usually via breakdown on $feature/<flag_key>, app_version, or utm_source).
Open the playbook for the kind from Step 1 and follow its numbered steps. Carry the record from 2.1 and any candidates from 2.3 into it.
Pick a segment the suspected cause should not have affected and rerun there. Stable in the control = strong hypothesis; moved too = expand the investigation. Skip when 2.2 already explained the movement.
Use the format below. Offer to save key charts via posthog:insight-create. If a cause is found and no annotation marks it, offer posthog:annotation-create. See common-causes.md for the cause taxonomy.
markdown# Investigation: <metric> **Anomaly**: <baseline> → <current> (<delta>) starting <date> ## Likely cause <one sentence> **Confidence**: low | medium | high — <one-line reason> **Evidence** - <query result> - <flag / experiment / annotation / commit if applicable> ## Possible causes (ruled out) - <hypothesis>: <why> ## Affected segment - <shared properties of affected users/events> ## Data gaps - <checks skipped and why> ## Suggested follow-ups - <concrete next action> - <offer to save chart / create annotation>
Confidence rule of thumb:
delta _and_ a flag/version aligns _and_ an error or annotation matches).
cross-check.
rules things _out_.
Link insights and dashboards inline: [Name](/insights/short_id).
box-plot, funnel, retention, stickiness, lifecycle, paths
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,562 | 3,950 | -29% | 1 | 1 | 0% | 407 | 2,209 | +443% | 0 | 0 | — |
case-02 | fail→fail | 31,801 | 4,237 | -87% | 1 | 1 | 0% | 4,950 | 2,209 | -55% | 0 | 0 | — |
case-03 | fail→fail | 17,755 | 4,080 | -77% | 1 | 1 | 0% | 3,404 | 2,215 | -35% | 0 | 0 | — |
case-04 | fail→fail | 9,058 | 10,461 | +15% | 1 | 1 | 0% | 1,489 | 2,805 | +88% | 0 | 0 | — |
case-05 | pass→pass | 7,692 | 3,127 | -59% | 1 | 1 | 0% | 1,482 | 2,591 | +75% | 0 | 0 | — |
case-06 | pass→pass | 7,723 | 2,353 | -70% | 1 | 1 | 0% | 1,323 | 2,370 | +79% | 0 | 0 | — |
case-07 | pass→pass | 10,516 | 8,755 | -17% | 1 | 1 | 0% | 1,943 | 3,604 | +85% | 0 | 0 | — |
case-08 | fail→fail | 7,555 | 6,037 | -20% | 1 | 1 | 0% | 1,261 | 2,454 | +95% | 0 | 0 | — |
case-09 | fail→fail | 7,382 | 1,838 | -75% | 1 | 1 | 0% | 1,376 | 2,279 | +66% | 0 | 0 | — |
case-10 | fail→fail | 6,553 | 2,654 | -59% | 1 | 1 | 0% | 1,177 | 2,427 | +106% | 0 | 0 | — |
case-11 | fail→fail | 13,528 | 7,901 | -42% | 1 | 1 | 0% | 2,942 | 3,491 | +19% | 0 | 0 | — |
case-12 | pass→fail | 8,349 | 6,010 | -28% | 1 | 1 | 0% | 1,506 | 2,341 | +55% | 0 | 0 | — |
case-13 | pass→pass | 3,945 | 3,219 | -18% | 1 | 1 | 0% | 750 | 2,581 | +244% | 0 | 0 | — |
case-14 | pass→pass | 8,065 | 3,815 | -53% | 1 | 1 | 0% | 1,339 | 2,637 | +97% | 0 | 0 | — |
case-15 | fail→pass | 5,255 | 2,119 | -60% | 1 | 1 | 0% | 1,003 | 2,394 | +139% | 0 | 0 | — |
case-16 | pass→pass | 13,684 | 9,508 | -31% | 1 | 1 | 0% | 2,531 | 3,662 | +45% | 0 | 0 | — |
case-17 | fail→pass | 10,677 | 3,359 | -69% | 1 | 1 | 0% | 1,842 | 2,649 | +44% | 0 | 0 | — |
case-18 | pass→pass | 3,613 | 1,084 | -70% | 1 | 1 | 0% | 723 | 2,150 | +197% | 0 | 0 | — |
case-19 | pass→pass | 7,064 | 1,124 | -84% | 1 | 1 | 0% | 1,335 | 2,153 | +61% | 0 | 0 | — |
case-20 | fail→fail | 3,490 | 2,753 | -21% | 1 | 1 | 0% | 719 | 2,154 | +200% | 0 | 0 | — |
case-21 | pass→fail | 4,994 | 5,264 | +5% | 1 | 1 | 0% | 1,081 | 2,272 | +110% | 0 | 0 | — |
case-22 | pass→fail | 8,332 | 4,089 | -51% | 1 | 1 | 0% | 1,759 | 2,286 | +30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 13 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 13 comparable cases. 6 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.