Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.
.claude/skills/kunanonj-cursor-plugin-posthog-exploring-llm-clusters/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 182% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 166% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 53% | 0% |
Use this skill when investigating AI observability clusters — understanding what patterns exist in your AI/LLM traffic, comparing cluster behavior, and drilling into individual clusters.
| Tool | Purpose | | ---------------------------------- | ----------------------------------------------- | | posthog:llma-clustering-job-list | List clustering job configurations for the team | | posthog:llma-clustering-job-get | Get a specific clustering job by ID | | posthog:execute-sql | Query cluster run events and compute metrics | | posthog:query-llm-traces-list | Find traces belonging to a cluster | | posthog:query-llm-trace | Inspect a specific trace in detail |
PostHog clusters LLM traces (or individual generations) by embedding similarity. A Temporal workflow runs periodically or on-demand, producing cluster events stored as $ai_trace_clusters (trace-level) or $ai_generation_clusters (generation-level).
Each cluster event contains:
$ai_clustering_run_id — unique run identifier (format: <team_id>_<level>_<YYYYMMDD>_<HHMMSS>[_<job_id>])$ai_clustering_level — "trace" or "generation"$ai_window_start / $ai_window_end — time window analyzed$ai_total_items_analyzed — number of traces/generations processed$ai_clusters — JSON array of cluster objects$ai_clustering_params — algorithm parameters used$ai_clusters)json{ "cluster_id": 0, "size": 42, "title": "User authentication flows", "description": "Traces involving login, signup, and token refresh operations", "traces": { "<trace_or_generation_id>": { "distance_to_centroid": 0.123, "rank": 0, "x": -2.34, "y": 1.56, "timestamp": "2026-03-28T10:00:00Z", "trace_id": "abc-123", "generation_id": "gen-456" } }, "centroid_x": -2.1, "centroid_y": 1.4 }
cluster_id: -1 is the noise/outlier cluster (items that didn't fit any cluster)traces are keyed by trace ID (trace-level) or generation event UUID (generation-level)rank orders items by proximity to centroid (0 = closest)x, y are 2D coordinates for visualization (UMAP/PCA/t-SNE reduced)Each team can have up to 5 clustering jobs. A job defines:
"trace" or "generation"Default jobs named "Default - trace" and "Default - generation" are auto-created and disabled when a custom job is created for the same level.
sqlposthog:execute-sql SELECT JSONExtractString(properties, '$ai_clustering_run_id') as run_id, JSONExtractString(properties, '$ai_clustering_level') as level, JSONExtractString(properties, '$ai_window_start') as window_start, JSONExtractString(properties, '$ai_window_end') as window_end, JSONExtractInt(properties, '$ai_total_items_analyzed') as total_items, timestamp FROM events WHERE event IN ('$ai_trace_clusters', '$ai_generation_clusters') AND timestamp >= now() - INTERVAL 7 DAY ORDER BY timestamp DESC LIMIT 10
sqlposthog:execute-sql SELECT JSONExtractString(properties, '$ai_clustering_run_id') as run_id, JSONExtractString(properties, '$ai_clustering_level') as level, JSONExtractString(properties, '$ai_clustering_job_id') as job_id, JSONExtractString(properties, '$ai_clustering_job_name') as job_name, JSONExtractString(properties, '$ai_window_start') as window_start, JSONExtractString(properties, '$ai_window_end') as window_end, JSONExtractInt(properties, '$ai_total_items_analyzed') as total_items, JSONExtractRaw(properties, '$ai_clusters') as clusters, JSONExtractRaw(properties, '$ai_clustering_params') as params FROM events WHERE event IN ('$ai_trace_clusters', '$ai_generation_clusters') AND JSONExtractString(properties, '$ai_clustering_run_id') = '<run_id>' LIMIT 1
The clusters field is a JSON array. Parse it to see cluster titles, sizes, and descriptions.
Important: The clusters JSON can be very large (thousands of trace IDs with coordinates). When the result is too large for inline display, it auto-persists to a file. Use print_clusters.py from scripts/ to get a readable summary.
For trace-level clusters, compute cost/latency/token metrics:
sqlposthog:execute-sql SELECT JSONExtractString(properties, '$ai_trace_id') as trace_id, sum(toFloat(properties.$ai_total_cost_usd)) as total_cost, max(toFloat(properties.$ai_latency)) as latency, sum(toInt(properties.$ai_input_tokens)) as input_tokens, sum(toInt(properties.$ai_output_tokens)) as output_tokens, countIf(properties.$ai_is_error = 'true') as error_count FROM events WHERE event IN ('$ai_generation', '$ai_embedding', '$ai_span') AND timestamp >= parseDateTimeBestEffort('<window_start>') AND timestamp <= parseDateTimeBestEffort('<window_end>') AND JSONExtractString(properties, '$ai_trace_id') IN ('<trace_id_1>', '<trace_id_2>', ...) GROUP BY trace_id
For generation-level clusters, match by event UUID:
sqlposthog:execute-sql SELECT toString(uuid) as generation_id, toFloat(properties.$ai_total_cost_usd) as cost, toFloat(properties.$ai_latency) as latency, toInt(properties.$ai_input_tokens) as input_tokens, toInt(properties.$ai_output_tokens) as output_tokens, if(properties.$ai_is_error = 'true', 1, 0) as is_error FROM events WHERE event = '$ai_generation' AND timestamp >= parseDateTimeBestEffort('<window_start>') AND timestamp <= parseDateTimeBestEffort('<window_end>') AND toString(uuid) IN ('<gen_uuid_1>', '<gen_uuid_2>', ...)
Once you've identified interesting clusters, use the trace tools to inspect individual traces:
jsonposthog:query-llm-trace { "traceId": "<trace_id_from_cluster>", "dateRange": {"date_from": "<window_start>", "date_to": "<window_end>"} }
avg(cost), avg(latency), sum(cost) per clustertraces field)rank (closest to centroid = most representative)query-llm-trace to understand the patterntitle and description for the AI-generated summaryerror_countitems_with_errors / total_itemshttps://app.posthog.com/ai-observability/clustershttps://app.posthog.com/ai-observability/clusters/<url_encoded_run_id>https://app.posthog.com/ai-observability/clusters/<url_encoded_run_id>/<cluster_id>Always surface these links so the user can verify visually in the PostHog UI.
cluster_id: -1) contains outliers that didn't fit any patternllma-clustering-job-list to understand what clustering configs are activequery-llm-trace for deep inspection| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 11,709 | 10,875 | -7% | 1 | 1 | 0% | 1,923 | 3,736 | +94% | 0 | 0 | — |
case-01 | fail→fail | 12,502 | 6,531 | -48% | 1 | 1 | 0% | 1,518 | 2,902 | +91% | 0 | 0 | — |
case-02 | fail→fail | 17,470 | 5,798 | -67% | 1 | 1 | 0% | 3,421 | 2,779 | -19% | 0 | 0 | — |
case-03 | fail→fail | 6,314 | 5,825 | -8% | 1 | 1 | 0% | 566 | 2,766 | +389% | 0 | 0 | — |
case-05 | fail→pass | 5,290 | 2,636 | -50% | 1 | 1 | 0% | 1,036 | 2,924 | +182% | 0 | 0 | — |
case-06 | pass→pass | 4,506 | 3,215 | -29% | 1 | 1 | 0% | 754 | 2,932 | +289% | 0 | 0 | — |
case-07 | fail→pass | 10,271 | 3,314 | -68% | 1 | 1 | 0% | 2,195 | 3,032 | +38% | 0 | 0 | — |
case-08 | fail→pass | 16,773 | 4,489 | -73% | 1 | 1 | 0% | 2,214 | 3,457 | +56% | 0 | 0 | — |
case-13 | fail→fail | 6,993 | 2,007 | -71% | 1 | 1 | 0% | 1,145 | 2,689 | +135% | 0 | 0 | — |
case-09 | fail→pass | 17,153 | 7,514 | -56% | 1 | 1 | 0% | 1,406 | 3,740 | +166% | 0 | 0 | — |
case-10 | fail→pass | 15,781 | 11,118 | -30% | 1 | 1 | 0% | 2,897 | 4,442 | +53% | 0 | 0 | — |
case-11 | fail→pass | 10,577 | 3,528 | -67% | 1 | 1 | 0% | 2,027 | 3,107 | +53% | 0 | 0 | — |
case-12 | pass→pass | 11,716 | 5,987 | -49% | 1 | 1 | 0% | 2,519 | 3,694 | +47% | 0 | 0 | — |
case-14 | fail→pass | 11,826 | 2,390 | -80% | 1 | 1 | 0% | 2,332 | 2,908 | +25% | 0 | 0 | — |
case-15 | pass→pass | 16,390 | 2,663 | -84% | 1 | 1 | 0% | 3,168 | 2,959 | -7% | 0 | 0 | — |
case-16 | pass→pass | 9,223 | 2,515 | -73% | 1 | 1 | 0% | 1,512 | 2,899 | +92% | 0 | 0 | — |
case-17 | pass→pass | 10,152 | 5,457 | -46% | 1 | 1 | 0% | 1,780 | 3,324 | +87% | 0 | 0 | — |
case-18 | fail→pass | 12,456 | 6,325 | -49% | 1 | 1 | 0% | 2,292 | 3,678 | +60% | 0 | 0 | — |
case-19 | fail→fail | 14,650 | 6,248 | -57% | 1 | 1 | 0% | 2,547 | 2,815 | +11% | 0 | 0 | — |
case-20 | pass→pass | 13,821 | 14,676 | +6% | 1 | 1 | 0% | 2,632 | 5,238 | +99% | 0 | 0 | — |
case-21 | pass→pass | 13,203 | 9,575 | -27% | 1 | 1 | 0% | 2,554 | 4,420 | +73% | 0 | 0 | — |
case-22 | pass→pass | 9,944 | 8,376 | -16% | 1 | 1 | 0% | 2,346 | 4,254 | +81% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.