Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Benchmark one session (or a small recent set) against the rolling average using Agent Monitor data — cost, total tokens, tool count, and workflow complexity score — and report where each metric lands as a percentile of the population. Tells you whether a session was normal, cheap, or an outlier. Use when judging whether a session was typical or out of band.
.claude/skills/hoangsonww-benchmark/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 180% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -20% | 0% |
Score a session against the rolling population average and report its percentile on cost, tokens, tool count, and complexity using Agent Monitor data.
The user provides: $ARGUMENTS
This may be:
| Endpoint | Returns | |----------|---------| | GET /api/sessions?limit=N | Population of sessions with cost, model, started_at, metadata (turn_count, total_turn_duration_ms) — builds the rolling baseline | | GET /api/pricing/cost/{sessionId} | { total_cost, breakdown:[{ input_tokens, output_tokens, cache_read_tokens, cache_write_tokens, cost }] } — the target session's cost and tokens | | GET /api/workflows/{sessionId} | complexity (score), stats (tool/event counts), toolFlow (distinct tools used) — the target session's tool count and complexity | | GET /api/analytics | avg_events_per_session, tool_usage, daily_sessions — corroborates population-level averages |
Fetch the population with GET /api/sessions?limit=200 (the rolling set). For each session gather cost (GET /api/pricing/cost/{id} or the list cost field), total tokens (sum of the 4 token types from the pricing breakdown), tool count and complexity (GET /api/workflows/{id}). Compute mean, median, and standard deviation for each metric across the population.
For the requested session, pull the same four metrics:
total_cost from GET /api/pricing/cost/{id}.input + output + cache_read + cache_write summed from the breakdown.GET /api/workflows/{id} stats/toolFlow.complexity.score from GET /api/workflows/{id}.For each metric report the target's percentile within the population (share of sessions at or below it) and its z-score (value − mean) / stddev. Label each: below average / typical / above average / outlier (|z| > 2).
State whether the session was normal overall. If it is an outlier, name which metric drove it (e.g., complexity p96, cost p91 → an unusually heavy session).
Other measured skills in the registry, with their headline benchmark lift.