Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Guides a user from \"I want to watch recordings but don't know which ones\" to a short, high-signal list of sessions worth watching. Use when the user asks which sessions or replays to watch, wants help finding interesting / useful recordings, says they don't know where to start in session replay, or wants to watch sessions about a goal (signup, pricing, onboarding, checkout, a feature, rageclicks
.claude/skills/kunanonj-cursor-plugin-posthog-finding-sessions-to-watch/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 80% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 191% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 152% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 78% | 0% |
Most people open session replay with a goal ("why are signups dropping?") but no idea which of thousands of recordings to watch. A raw, unfiltered list is the worst possible answer — it buries the useful sessions in noise. Your job is to turn their intent into a focused filter, return a handful of high-signal recordings, and offer to dig into one.
The starting points below are the same ones the product surfaces as "filter templates" — they encode the jobs people actually use replay for. Treat them as a menu, not a script.
Never dump an unfiltered recording list. Always either (a) apply a goal-based filter, or (b) sort by a signal (activity, errors) so the first few rows are worth a click. If the user's goal is unclear, ask one short question or offer the menu before querying.
| Tool | Purpose | | ------------------------------------------- | ------------------------------------------------------------------------ | | posthog:query-session-recordings-list | Find/filter recordings (the workhorse). Returns metadata + id per row. | | posthog:read-data-schema | Confirm real event names, URLs, and property values before filtering. | | posthog:execute-sql | Collect $session_ids for sessions where a specific event happened. | | posthog:cohorts-list | Resolve a cohort name → id when scoping to a user segment. | | posthog:session-recording-playlist-create | Save the resulting filter as a saved filter view (type: 'filters'). |
Hand off to the investigating-replay skill once the user picks a recording to understand in depth.
Map the request to one of the starting points below. If it's vague ("show me something interesting"), offer 3-4 options rather than guessing, or default to most active sessions (high signal, no setup).
Event names and URLs vary per project — never assume $pageview paths, a signup_completed event, or a person property exists. Confirm with read-data-schema (event_properties, event_property_values, entity_property_values) before putting a value in a filter. If the needed event/property doesn't exist, say so and suggest the closest available signal.
Call query-session-recordings-list with only the filters that serve the goal. Recommended settings:
filter_test_accounts: true (the tool defaults to false) to exclude internal users, unless theuser is debugging their own session.
date_from of -7d to -30d for goal-based searches; -3d for "recent".order — activity_score for "interesting", console_error_count for "broken",start_time for "recent".
limit: 10 — you want a shortlist, not a dump.Don't relay raw rows. Pick the 3-5 most promising and say why each is worth watching (long active duration, many errors, reached the key page, high activity score). Deep-link each as {posthog_base_url}/replay/{id} — never /replay/home?sessionRecordingId={id}. Note total matches so the user knows how much is behind the shortlist.
investigating-replay.session-recording-playlist-create (type: 'filters' — a filter view, not a 'collection', which is for manually curated recordings and can't carry filters).
Two filter shapes cover almost everything:
visited_page ({ "type": "recording", "key": "visited_page","operator": "icontains", "value": "/pricing" }).
the recordings query, so first collect session IDs with execute-sql, then pass them as session_ids (see the two-step pattern below).
| User goal | Approach | | ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Signup / onboarding / pricing / checkout friction | visited_page icontains the relevant path (confirm the real path first). Order start_time, or console_error_count to surface broken ones. | | A specific feature | Two-step: execute-sql for $session_ids where the feature event fired, then session_ids. Pair with visited_page if the feature lives on one page. | | Rageclicks / frustration | Two-step on the $rageclick event → session_ids. | | Errors / something broken | properties: [{ "type": "recording", "key": "console_error_count", "operator": "gt", "value": 0 }], order console_error_count. | | A/B test / feature flag | { "type": "flag", "key": "<flag-key>", "operator": "flag_evaluates_to", "value": "<variant or true>" }. | | A specific person / segment | person_uuid, a person property filter (e.g. email), or a cohort filter (cohorts-list for the id). | | Mobile / responsive issues | { "type": "event", "key": "$device_type", "operator": "exact", "value": ["Mobile"] }, or { "type": "event", "key": "$screen_width", "operator": "lt", "value": 600 }. | | Most active users / "just show me good ones" | No filter; order: "activity_score". The reliable default when the user has no specific goal. | | Most active pages | execute-sql to rank $pageview by URL, then filter recordings by the hottest page's visited_page. |
The recordings query filters by event _properties_, not event _names_. To find sessions that contain a particular event, collect the session IDs first:
sqlposthog:execute-sql SELECT $session_id FROM events WHERE event = '$rageclick' -- or your signup/search/feature event (confirm via read-data-schema) AND timestamp > now() - INTERVAL 7 DAY AND $session_id != '' GROUP BY $session_id ORDER BY max(timestamp) DESC -- recent first: UUIDs aren't time-ordered, so the LIMIT must keep the freshest sessions LIMIT 100
Then fetch those recordings (some session IDs won't have a recording — that's expected). Pass the same date_from window as the SQL step — with only session_ids, the query falls back to its -3d default and would drop sessions whose event was older than that:
jsonposthog:query-session-recordings-list { "date_from": "-7d", "session_ids": ["<id1>", "<id2>", "..."] }
User: "Why are people bouncing on our pricing page? Show me some sessions."
visited_page approach.read-data-schema (event_property_values for $pathname) to confirm the path is /pricing.jsonposthog:query-session-recordings-list { "date_from": "-14d", "filter_test_accounts": true, "order": "activity_score", "limit": 10, "properties": [ { "type": "recording", "key": "visited_page", "operator": "icontains", "value": "/pricing" } ] }
{base}/replay/{id}, noting which lingered or hit errors.investigating-replay) or save it as a saved filter view (type: 'filters').nothing to watch; if it's still empty, recordings may not be captured for that flow (point the user to diagnosing-missing-recordings).
activity_score is a solid default proxy for "worth watching" when there's no sharper signal — but itrewards raw interaction volume, so prefer a goal-based filter (errors, a key page) when you have one.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 8,861 | 5,145 | -42% | 1 | 1 | 0% | 1,287 | 3,112 | +142% | 0 | 0 | — |
case-01 | fail→fail | 10,603 | 6,132 | -42% | 1 | 1 | 0% | 1,747 | 2,563 | +47% | 0 | 0 | — |
case-02 | fail→fail | 8,857 | 6,600 | -25% | 1 | 1 | 0% | 1,486 | 2,568 | +73% | 0 | 0 | — |
case-03 | fail→fail | 13,612 | 4,912 | -64% | 1 | 1 | 0% | 2,125 | 2,475 | +16% | 0 | 0 | — |
case-05 | pass→fail | 15,490 | 8,582 | -45% | 1 | 1 | 0% | 2,473 | 2,657 | +7% | 0 | 0 | — |
case-06 | pass→fail | 10,527 | 4,610 | -56% | 1 | 1 | 0% | 1,699 | 2,481 | +46% | 0 | 0 | — |
case-07 | fail→fail | 10,881 | 6,767 | -38% | 1 | 1 | 0% | 1,618 | 2,554 | +58% | 0 | 0 | — |
case-08 | pass→pass | 5,522 | 5,774 | +5% | 1 | 1 | 0% | 871 | 3,189 | +266% | 0 | 0 | — |
case-09 | pass→pass | 9,928 | 4,072 | -59% | 1 | 1 | 0% | 1,572 | 2,786 | +77% | 0 | 0 | — |
case-10 | fail→fail | 6,323 | 9,474 | +50% | 1 | 1 | 0% | 1,106 | 2,886 | +161% | 0 | 0 | — |
case-11 | pass→pass | 6,246 | 2,561 | -59% | 1 | 1 | 0% | 1,139 | 2,592 | +128% | 0 | 0 | — |
case-12 | fail→pass | 11,029 | 5,966 | -46% | 1 | 1 | 0% | 1,730 | 3,116 | +80% | 0 | 0 | — |
case-13 | fail→pass | 6,220 | 4,295 | -31% | 1 | 1 | 0% | 1,020 | 2,969 | +191% | 0 | 0 | — |
case-14 | pass→pass | 5,375 | 2,510 | -53% | 1 | 1 | 0% | 1,047 | 2,513 | +140% | 0 | 0 | — |
case-15 | fail→fail | 9,825 | 6,126 | -38% | 1 | 1 | 0% | 1,589 | 2,741 | +72% | 0 | 0 | — |
case-16 | pass→pass | 7,481 | 3,562 | -52% | 1 | 1 | 0% | 1,255 | 2,807 | +124% | 0 | 0 | — |
case-17 | fail→pass | 7,400 | 3,774 | -49% | 1 | 1 | 0% | 1,469 | 2,827 | +92% | 0 | 0 | — |
case-18 | fail→pass | 9,807 | 10,566 | +8% | 1 | 1 | 0% | 1,682 | 4,235 | +152% | 0 | 0 | — |
case-19 | fail→fail | 8,186 | 2,762 | -66% | 1 | 1 | 0% | 1,452 | 2,674 | +84% | 0 | 0 | — |
case-20 | fail→fail | 11,619 | 2,197 | -81% | 1 | 1 | 0% | 1,674 | 2,440 | +46% | 0 | 0 | — |
case-21 | pass→pass | 13,700 | 4,091 | -70% | 1 | 1 | 0% | 2,043 | 2,824 | +38% | 0 | 0 | — |
case-22 | fail→pass | 7,935 | 2,339 | -71% | 1 | 1 | 0% | 1,472 | 2,627 | +78% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 15 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.