Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audit model delegation and subagent effectiveness for a session — which models handled which subagent types, per-type success rates and average durations, and wasted delegations (heavy models on trivial work or types that consistently fail) — using the Agent Monitor workflow intelligence API. Use when reviewing how a session delegated work across models and subagents.
.claude/skills/hoangsonww-delegation-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 239% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -64% | 0% |
Audit how a Claude Code session delegated work: model-to-subagent mapping and whether each delegation paid off.
The user provides: $ARGUMENTS
A session ID. If empty, fetch GET /api/sessions?limit=1 and audit the most recent session, stating which one.
| Endpoint | Returns | |----------|---------| | GET /api/workflows/{sessionId} | The modelDelegation dataset (which models are delegated which subagent types) and the effectiveness dataset (per-type completion/success rate, avg duration, task success) | | GET /api/agents | Raw subagent records (type, model, status, depth, parent) to corroborate counts and statuses |
From modelDelegation: a model × subagent-type table of how many agents of each type each model ran. | Model | explore | code-review | debugger | ... | Total | |-------|---------|-------------|----------|-----|-------|
From effectiveness: per type, the success rate and average duration. | Subagent type | Count | Success rate | Avg duration | Verdict | |---------------|-------|--------------|--------------|---------| Mark types below ~70% success as low-yield.
Flag, with evidence:
Concrete model reassignments grounded in the matrix and effectiveness data. State the type, the model used, the success rate, and the suggested model — only where the data supports it.
1m 12s).effectiveness dataset does not provide.npm start from the repo root.Other measured skills in the registry, with their headline benchmark lift.