Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Detect quality and efficiency regressions over time using Agent Monitor data — rising error rate (APIError events), falling cache hit rate, growing compaction frequency, and climbing cost-per-session. Splits history into an earlier baseline window and a recent window and reports which metrics are getting worse, by how much, and where. Use when checking whether things are degrading or trending in the wrong direction.
.claude/skills/hoangsonww-regression-watch/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -23% | 0% |
Detect whether Claude Code sessions are getting worse over time across quality and efficiency metrics, using Agent Monitor data.
The user provides: $ARGUMENTS
This may be:
| Endpoint | Returns | |----------|---------| | GET /api/analytics | daily_events (365d), daily_sessions (365d), event_types, tokens (total_input, total_output, total_cache_read, total_cache_write — baselines pre-summed), avg_events_per_session | | GET /api/events?session_id=X | Event stream incl. APIError, Compaction, PreToolUse/PostToolUse — used to localize regressions to specific sessions | | GET /api/pricing/cost | { total_cost, breakdown[...] } — total cost to derive cost-per-session | | GET /api/pricing/cost/{sessionId} | Per-session cost — used to compare recent vs baseline session cost | | GET /api/workflows/{sessionId} | compaction (impact), errorPropagation (by depth), effectiveness — per-session quality signals | | GET /api/sessions?limit=N | Sessions with started_at, cost, metadata — to bucket sessions into time windows |
Split history into a baseline window (older) and a recent window (newer). Default: recent = last 30 days, baseline = the 30–90 day range before it. Use daily_events/daily_sessions for series metrics and GET /api/sessions?limit=N to assign sessions to each window by started_at.
APIError count / total events in the recent window(from event_types and daily_events, or per-session GET /api/events).
APIError events.
total_cache_read / (total_cache_read + total_input).the pricing breakdown). Flag a falling hit rate — that means more uncached input tokens and higher cost.
Compaction events / session per window (fromevent_types / daily_events, confirmed via per-session GET /api/workflows/{id} compaction). Flag a rising rate — context is overflowing more often.
GET /api/pricing/cost overall and GET /api/pricing/cost/{id} for the sessions in each window. Flag a climbing value.
Roll up which metrics regressed, rank by relative worsening, and name the most likely driver (e.g., cache hit rate fell → cost per session climbed).
Other measured skills in the registry, with their headline benchmark lift.