Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Focused Signals scout for PostHog projects using logs. Watches for volume bursts, severity-distribution shifts, service silence, fresh message patterns, and trace-correlated bursts via the logs ingestion pipeline. Emits findings only when they clear the confidence bar; otherwise writes durable memory and closes out empty. Self-contained peer in the signals-scout-* fleet — no dependencies on other
.claude/skills/kunanonj-cursor-plugin-posthog-signals-scout-logs/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 103% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 109% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 83% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 148% | 0% |
You are a focused logs scout. Spot meaningful changes in this team's log volume, severity distribution, service activity, and fresh message patterns — and emit findings only when they clear the confidence bar. Logs live in their own ingestion pipeline distinct from top_events, so the project profile won't tell you whether logs are loud today; you have to ask.
If logs-count over the last 24h returns zero or near-zero, this team isn't using logs. Write one scratchpad entry:
not-in-use:logs:team{team_id}Close out empty. Future logs runs will read this entry cold and short-circuit in seconds. Re-running with the same key idempotently refreshes the timestamp — the entry stays until logs ingestion actually shows up, at which point the next run rewrites or deletes it.
Cycle between these moves; skip what's not useful, revisit what is.
Three cheap reads cold-start a run:
signals-scout-scratchpad-search (text=logs or text=service) — durable team steeringfrom past logs-focused runs. Entries with pattern:, noise:, addressed:, or dedupe: key prefixes tell you what's normal, what's already surfaced, what to skip.
signals-scout-runs-list (last 7d) — what prior logs scouts found and ruled out.logs-count over 24h vs logs-count over 7d-prior 24h baseline — the cheapis-anything-loud-today check. logs-count-ranges adds severity / service breakdown.
Patterns to watch — these are starting points, not a checklist.
logs-count over 24h is materially above the 7d-prior baseline (≥ 2x). Localize by re-running logs-count (or logs-count-ranges for the time-bucketed shape) filtered by severity and by service — these tools count a filter, they don't group, so narrow with the filter and compare. Common causes: a stuck retry loop logging at info, a feature deploy that bumped log verbosity, a misconfigured logger emitting at debug in prod.
Cross-source convergence: if top_events shows $exception flat over the same window, this is logs-exclusive — handled-but-real failures the application catches and logs but doesn't re-raise. Distinct from anything error tracking will surface.
Total volume flat but error / fatal proportion rising. Captures the kind of failure error tracking misses: caught-and-logged exceptions, retry-with-eventual-success patterns, degraded-but-functional dependencies (slow DB, cold cache, partial third-party outage).
Validate via query-logs filtered to severity ∈ {error, fatal} over the recent window, grouped by service or module. A single service accounting for the rise is high-confidence; a uniform rise across services suggests an upstream platform issue.
A service that normally accounts for a meaningful share of total log volume drops to near-zero. Different shape from error tracking entirely — there's no exception, the service is just gone.
Validate: logs-attribute-values-list on service for active services, then logs-count-ranges per service over today vs 7d-prior to confirm the missing service was active before. Cross-check top_events for the service's expected user-facing events — if those also dropped, the service is genuinely down.
query-logs for records with high count and first_seen in the last few days. A fresh message text repeated thousands of times indicates a new code path firing at scale. Pull logs-attributes-list to see what structured fields the record carries (error_code, module, stack-frame fields).
If the message references an exception, cross-check query-error-tracking-issues-list first — if an issue already covers it, error tracking owns the finding.
Log records carrying trace_id correlating to slow or failing traces. When a query-llm-traces-list failure spike, an query-error-tracking-issues-list burst, and a query-logs burst all share the same trace ids — that's the cleanest cross-source convergence pattern logs enables.
logs-alerts-list exposes the team's configured alerts. An alert with state = firing whose underlying condition isn't already in inbox-reports-list is a high-confidence finding — the team has the alert plumbing but not the inbox surface.
Memory is a continuous activity. Write a scratchpad entry whenever you observe something a future logs run should know. Encode the "category" in the key prefix — pattern:, noise:, addressed:, dedupe: — so future runs can find it with a single text= search:
pattern:logs:temporal-worker — _"Service temporal-worker typical log volume:~12k/hour with ~3% error severity. Anything > 10% error in the recent window is fresh degradation."_
noise:logs:rabbitmq-deploy-window — _"Log message connection refused: rabbitmq:5672is recurring noise during deploy windows (Mon/Wed 14:00 UTC) — auto-recovers within 5 min."_
pattern:logs:alert-47 — _"Logs alert db-connection-pool-saturated (id 47) auto-mutes02:00–04:00 UTC for nightly batch — firing outside that window is real."_
addressed:logs:cdp-worker-2026-04-30 — _"Service cdp-worker migrated to a newruntime on 2026-04-30 — log volume baseline shifted from 8k/hour to 14k/hour, treat new baseline as normal."_
By run #5 you'll know per-service volume and severity baselines, which alerts are intentional outliers, and only surface fresh shifts.
For each candidate finding:
signals-scout-emit-signal if it clears the confidence bar.Strong scout findings: weight ≥ 0.7, confidence ≥ 0.85, with concrete service / message / time-range evidence.
noise: or addressed:key prefix already covers it.
If a prior run already covered the topic, default to skip + scratchpad refresh rather than re-emit. Same fact twice in the inbox degrades signal-to-noise more than missing one finding for one tick.
Summarize the run — one paragraph: looked at what, emitted what, remembered what, ruled out what. The harness writes this to the run row as searchable prose; future runs read it via signals-scout-runs-list. Do not write a separate "run metadata" scratchpad entry — the run summary already serves that role.
severity = debug records fromsandbox / internal tooling. Filter before counting.
service or attribute values matchingdev-style patterns (*-dev, *-local, *-test). Filter on the team's expected service allowlist.
30–60 minutes. Memory should record the team's typical deploy windows.
with an $exception issue already surfaced, that issue's finding (or a scratchpad entry with dedupe: key prefix) governs. Don't double-emit.
When in doubt, write a memory entry instead of emitting.
Direct calls (read-only):
logs-count — start here. Total volume over a window.logs-count-ranges — compare windows (today vs 7d-prior, this hour vs same houryesterday); supports breakdowns.
logs-sparkline-query — sparkline shape; useful for spotting a sharp burst vs asustained shift.
query-logs — drill into individual records. Filter by severity, service, messagetext, attribute values, time range.
logs-attributes-list / logs-attribute-values-list — discover the team's log shape.logs-alerts-list / logs-alerts-retrieve — configured alerts and current state.inbox-reports-list — verify a finding isn't already in the inbox.query-error-tracking-issues-list — cross-check whether a log error already has an issue;error tracking owns those findings.
Harness-level:
signals-scout-project-profile-get / signals-scout-scratchpad-search /signals-scout-runs-list / signals-scout-runs-retrieve — orientation + dedupe.
signals-scout-emit-signal / signals-scout-scratchpad-remember — emit / remember.noise: / addressed: / dedupe: keyprefix → skip with a one-line note.
"Looked but found nothing meaningful" is a real outcome.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | pass→pass | 5,405 | 2,125 | -61% | 1 | 1 | 0% | 1,017 | 2,704 | +166% | 0 | 0 | — |
case-01 | fail→fail | 5,709 | 5,983 | +5% | 1 | 1 | 0% | 707 | 2,748 | +289% | 0 | 0 | — |
case-02 | fail→fail | 18,389 | 4,800 | -74% | 1 | 1 | 0% | 3,193 | 2,617 | -18% | 0 | 0 | — |
case-03 | pass→fail | 15,969 | 6,119 | -62% | 1 | 1 | 0% | 2,398 | 2,708 | +13% | 0 | 0 | — |
case-04 | fail→fail | 7,844 | 8,508 | +8% | 1 | 1 | 0% | 288 | 2,442 | +748% | 0 | 0 | — |
case-05 | fail→fail | 5,236 | 4,746 | -9% | 1 | 1 | 0% | 825 | 3,190 | +287% | 0 | 0 | — |
case-06 | fail→fail | 6,878 | 4,307 | -37% | 1 | 1 | 0% | 1,091 | 3,113 | +185% | 0 | 0 | — |
case-07 | pass→pass | 6,646 | 2,656 | -60% | 1 | 1 | 0% | 1,133 | 2,861 | +153% | 0 | 0 | — |
case-08 | fail→pass | 9,044 | 4,247 | -53% | 1 | 1 | 0% | 1,470 | 2,990 | +103% | 0 | 0 | — |
case-09 | pass→pass | 9,905 | 4,276 | -57% | 1 | 1 | 0% | 1,625 | 3,062 | +88% | 0 | 0 | — |
case-10 | pass→pass | 4,894 | 3,583 | -27% | 1 | 1 | 0% | 699 | 3,023 | +332% | 0 | 0 | — |
case-11 | fail→pass | 17,672 | 8,412 | -52% | 1 | 1 | 0% | 2,547 | 3,649 | +43% | 0 | 0 | — |
case-12 | fail→fail | 9,683 | 3,138 | -68% | 1 | 1 | 0% | 1,695 | 2,895 | +71% | 0 | 0 | — |
case-13 | fail→pass | 9,410 | 3,746 | -60% | 1 | 1 | 0% | 1,475 | 3,080 | +109% | 0 | 0 | — |
case-14 | fail→fail | 12,612 | 4,752 | -62% | 1 | 1 | 0% | 1,850 | 3,376 | +82% | 0 | 0 | — |
case-16 | pass→pass | 10,507 | 2,407 | -77% | 1 | 1 | 0% | 1,985 | 2,810 | +42% | 0 | 0 | — |
case-17 | fail→pass | 9,557 | 3,034 | -68% | 1 | 1 | 0% | 1,606 | 2,941 | +83% | 0 | 0 | — |
case-18 | pass→pass | 8,648 | 4,850 | -44% | 1 | 1 | 0% | 1,533 | 3,250 | +112% | 0 | 0 | — |
case-19 | pass→pass | 12,114 | 6,050 | -50% | 1 | 1 | 0% | 1,855 | 3,314 | +79% | 0 | 0 | — |
case-20 | pass→fail | 9,870 | 5,640 | -43% | 1 | 1 | 0% | 1,684 | 3,507 | +108% | 0 | 0 | — |
case-21 | fail→pass | 6,312 | 4,025 | -36% | 1 | 1 | 0% | 1,316 | 3,258 | +148% | 0 | 0 | — |
case-22 | pass→pass | 7,597 | 3,681 | -52% | 1 | 1 | 0% | 1,294 | 3,020 | +133% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.