Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Focused Signals scout for PostHog projects using error tracking. Watches `$exception` bursts, stuck loops, multi-fingerprint clusters, status regressions, and stack-trace activity-name patterns. Emits findings only when they clear the confidence bar; otherwise writes durable memory and closes out empty. Self-contained peer in the signals-scout-* fleet — no dependencies on other skills. Picked unif
.claude/skills/kunanonj-cursor-plugin-posthog-signals-scout-error-tracking/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 76% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 134% | 0% |
You are a focused error tracking scout. Spot meaningful changes in this team's $exception activity — bursts, stuck loops, multi-fingerprint clusters, status regressions, deploy-correlated regressions — and emit findings only when they clear the confidence bar.
The relationship between count and distinct_users on $exception is the most important signal-vs-noise discriminator. Internalize that shape.
If $exception is absent from top_events or its count is at baseline (no fresh 24h activity, recent_24h_count ≪ count / 7), error tracking probably isn't where the signal is today. Cheap scratchpad entry + close out:
not-in-use:error_tracking:team{team_id} (if $exception is absent entirely)or pattern:error_tracking:baseline-team{team_id} (if it fires at a steady baseline with no fresh burst)
"$exception baseline ~{count}/day, no fresh 24h burst at {timestamp}"Close out empty. Re-running with the same key idempotently refreshes the timestamp; the next run reads the entry cold and short-circuits.
Cycle between these moves; skip what's not useful.
Three cheap reads cold-start a run:
signals-scout-scratchpad-search (text=error or text=exception) — durable teamsteering from past error-tracking runs. Entries with pattern:, noise:, addressed:, or dedupe: key prefixes tell you what's normal, what's already surfaced, what to skip.
signals-scout-runs-list (last 7d) — what prior error-tracking scouts found andruled out.
signals-scout-project-profile-get — the $exception row in top_events carriescount, distinct_users, recent_24h_count, recent_24h_users. Pattern the count/users ratio against the table below.
| Pattern | What it usually means | | ------------------------------------------------------- | -------------------------------------------- | | count and distinct_users both spike in 24h | Fresh broad-reach issue — investigate first | | recent_24h_count / count ≫ 1/7 and users also spike | Today's burst is unusually broad | | count very high, distinct_users very low | Stuck loop / retry storm — may not be urgent | | count ~ distinct_users for a single fingerprint | Per-request server path (one hit per user) | | count and distinct_users both quiet | Nothing fresh on this product |
Patterns to watch — starting points, not a checklist.
recent_24h_count and recent_24h_users both spike together. Usually a fresh regression — many users hitting it independently. Drill in:
query-error-tracking-issues-list filtered to status=active, sort by last_seen_at.execute-sql against events with event = '$exception' ANDproperties.$exception_issue_id = '<id>' grouped by toStartOfHour(timestamp).
(count(*) ≈ uniq(person_id)) → per-request server path, almost always a regression or missing migration.
recent_24h_count very high but recent_24h_users is small. A worker, cron, websocket, or retry is looping. Look at the issue's stack trace for the activity / job name. Often less urgent than a broad-reach burst, but worth a finding when count is in the thousands and the issue is fresh.
Multiple fresh fingerprints (different entity_ids in query-error-tracking-issues-list) appearing in the same time window with overlapping stack traces, modules, or call sites → likely shared root cause. Bundle them in one finding (single description, evidence list with all fingerprint ids, dedupe key per fingerprint).
An issue with status=resolved that's now firing again. Filter query-error-tracking-issues-list to status=active and check last_seen_at against first_seen_at — a large gap means old issue resurrected. High-confidence findings: the team explicitly closed them once.
When the issue is server-side, the stack trace usually names the failing activity / view / management command. Extract it (top frame, look for <activity>_activity, def view_name, etc.) and pair with activity-log-list to find a recent deploy or model change correlation. Cross-source convergence is where this scout earns its keep.
Memory is a continuous activity. Write a scratchpad entry whenever you observe something a future error-tracking run should know. Encode the "category" in the key prefix — pattern:, noise:, addressed:, dedupe: — so future runs find it with a single text= search:
pattern:error_tracking:baseline — _"Project's normal $exception baseline:~50/day across ~30 distinct users. Anything materially above that is fresh."_
dedupe:error_tracking:019de34e — _"Issue 019de34e — surfaced 2026-05-0111:31–13:22Z, then quiet. If quiet next run, treat as already-surfaced; if firing, escalate."_
noise:error_tracking:sandbox-timeoutexpired — _"Sandbox TimeoutExpired Dockererrors are recurring noise on this team — internal harness ops, not user-facing."_
pattern:error_tracking:fetch_signals_for_report_activity — _"Server activityfetch_signals_for_report_activity was a regression source on 2026-05-01 — if it appears in a fresh stack trace, double-check it's not the same root cause."_
By run #5 you'll have a local map of what's normal versus what warrants investigation, and burn less time on cold-start exploration.
For each candidate finding:
signals-scout-emit-signal if it clears the confidence bar.Strong scout findings: weight ≥ 0.7, confidence ≥ 0.85, with concrete issue ids, hourly count, distinct-user counts in the evidence.
noise: or addressed:key prefix already covers it.
Cross-check inbox-reports-list before emitting — if an issue is already in the inbox, emit only if the _new angle_ (broader reach, status regression, deploy correlation) is materially different. Otherwise the existing report's signals will pick yours up via cross-source clustering.
Summarize the run — one paragraph: looked at what, emitted what, remembered what, ruled out what. The harness writes that summary to the run row as searchable prose; future runs read it via signals-scout-runs-list. Do not write a separate "run metadata" scratchpad entry — the run summary already serves that role.
browser quirk. Confirmed via low count AND low distinct_users.
TimeoutExpired,agentsh failures. Internal harness operations, not user-facing.
API outages already covered by past memory. Skip unless volume / shape changes meaningfully.
When in doubt, write a memory entry instead of emitting.
Direct calls (read-only):
query-error-tracking-issues-list — start here. Filter status=active, sort bylast_seen_at desc.
query-error-tracking-issue — drill into one issue (frames, sample events,occurrence counts).
execute-sql against events — for hourly breakdowns, distinct-user counts,per-fingerprint correlation, time-window aggregations.
inbox-reports-list — check whether the issue is already in the inbox before emitting.activity-log-list — pair stack-trace activity names with recent deploys or modelchanges for cross-source convergence.
Harness-level:
signals-scout-project-profile-get / signals-scout-scratchpad-search /signals-scout-runs-list / signals-scout-runs-retrieve — orientation + dedupe.
signals-scout-emit-signal / signals-scout-scratchpad-remember — emit / remember.$exception row in profile is at baseline → close out empty.noise: / addressed: / dedupe: keyprefix → skip.
there's more you could look at. Fewer, better signals.
"Looked but found nothing meaningful" is a real outcome.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→fail | 4,287 | 3,407 | -21% | 1 | 1 | 0% | 221 | 2,409 | +990% | 0 | 0 | — |
case-03 | fail→fail | 6,450 | 7,095 | +10% | 1 | 1 | 0% | 216 | 2,555 | +1083% | 0 | 0 | — |
case-01 | fail→fail | 9,763 | 8,382 | -14% | 1 | 1 | 0% | 673 | 2,635 | +292% | 0 | 0 | — |
case-04 | pass→pass | 9,526 | 11,996 | +26% | 1 | 1 | 0% | 1,860 | 4,701 | +153% | 0 | 0 | — |
case-05 | pass→pass | 12,232 | 18,364 | +50% | 1 | 1 | 0% | 2,433 | 5,103 | +110% | 0 | 0 | — |
case-06 | pass→fail | 12,507 | 8,004 | -36% | 1 | 1 | 0% | 1,907 | 2,916 | +53% | 0 | 0 | — |
case-07 | fail→pass | 11,262 | 5,835 | -48% | 1 | 1 | 0% | 1,697 | 3,336 | +97% | 0 | 0 | — |
case-08 | fail→pass | 10,594 | 3,918 | -63% | 1 | 1 | 0% | 1,636 | 3,027 | +85% | 0 | 0 | — |
case-09 | pass→pass | 11,773 | 3,704 | -69% | 1 | 1 | 0% | 1,819 | 2,922 | +61% | 0 | 0 | — |
case-10 | fail→pass | 15,758 | 6,236 | -60% | 1 | 1 | 0% | 2,262 | 3,291 | +45% | 0 | 0 | — |
case-11 | pass→pass | 9,280 | 4,488 | -52% | 1 | 1 | 0% | 1,727 | 2,987 | +73% | 0 | 0 | — |
case-12 | pass→pass | 19,429 | 11,976 | -38% | 1 | 1 | 0% | 1,719 | 3,763 | +119% | 0 | 0 | — |
case-13 | fail→pass | 9,872 | 2,994 | -70% | 1 | 1 | 0% | 1,631 | 2,871 | +76% | 0 | 0 | — |
case-14 | fail→pass | 5,633 | 1,477 | -74% | 1 | 1 | 0% | 1,091 | 2,554 | +134% | 0 | 0 | — |
case-15 | fail→fail | 6,418 | 2,895 | -55% | 1 | 1 | 0% | 1,040 | 2,794 | +169% | 0 | 0 | — |
case-16 | pass→pass | 8,581 | 3,957 | -54% | 1 | 1 | 0% | 1,334 | 2,932 | +120% | 0 | 0 | — |
case-17 | fail→pass | 8,012 | 3,127 | -61% | 1 | 1 | 0% | 1,447 | 2,944 | +103% | 0 | 0 | — |
case-18 | pass→pass | 7,944 | 2,987 | -62% | 1 | 1 | 0% | 1,441 | 2,801 | +94% | 0 | 0 | — |
case-19 | fail→fail | 8,926 | 1,819 | -80% | 1 | 1 | 0% | 1,361 | 2,615 | +92% | 0 | 0 | — |
case-20 | fail→pass | 4,013 | 2,540 | -37% | 1 | 1 | 0% | 803 | 2,728 | +240% | 0 | 0 | — |
case-21 | fail→pass | 7,071 | 4,674 | -34% | 1 | 1 | 0% | 1,326 | 3,199 | +141% | 0 | 0 | — |
case-22 | fail→pass | 7,726 | 1,972 | -74% | 1 | 1 | 0% | 1,359 | 2,639 | +94% | 0 | 0 | — |
case-23 | fail→pass | 6,975 | 2,775 | -60% | 1 | 1 | 0% | 1,198 | 2,553 | +113% | 0 | 0 | — |
case-24 | pass→pass | 7,014 | 3,848 | -45% | 1 | 1 | 0% | 1,090 | 3,025 | +178% | 0 | 0 | — |
case-25 | fail→fail | 12,021 | 1,701 | -86% | 1 | 1 | 0% | 2,132 | 2,637 | +24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 21 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.