Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Diagnose agent failures by identifying missing capabilities. Use when an agent fails a task, produces wrong output, or needs escalation. Provides a decision tree for failure analysis.
.claude/skills/majiayu000-capability-diagnostic/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 2% | 0% |
Run this diagnostic when an agent:
allowed-tools in agent spec (if any)skills field)escalation skill)After diagnosis, report:
Diagnosis: [tool gap | context gap | model gap | spec gap | infrastructure]
Root cause: [one sentence]
Fix: [specific action to take]| Symptom | Likely Cause | Fix | |---------|-------------|-----| | Hook blocks Write/Edit | Orchestrator trying to code | Delegate to coder agent | | Agent ignores instructions | Context overload or wrong skills | Check preloaded skills, reduce context | | Tests fail on agent's code | Model tier too low for task complexity | Escalate model tier | | Agent asks many questions | Task spec is ambiguous | Rewrite spec with zero ambiguity | | Agent produces wrong format | Missing output schema | Add structured output to task prompt | | MCP tool call fails silently | Server not connected or rate limited | Check MCP status, retry after delay | | Agent edits wrong file | Missing file organization context | Add project structure to task prompt | | classifyHandoffIfNeeded error | Transient infrastructure bug | Respawn agent with same prompt | | Agent stalls mid-task | Context window exhaustion | Break task into smaller pieces | | Git conflict on commit | Multiple agents on same file | Enforce file ownership per agent |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,476 | 5,305 | -3% | 1 | 1 | 0% | 899 | 1,747 | +94% | 0 | 0 | — |
case-02 | fail→fail | 5,316 | 9,149 | +72% | 1 | 1 | 0% | 793 | 2,533 | +219% | 0 | 0 | — |
case-03 | fail→pass | 6,675 | 2,942 | -56% | 1 | 1 | 0% | 1,101 | 1,297 | +18% | 0 | 0 | — |
case-04 | pass→pass | 15,818 | 12,286 | -22% | 1 | 1 | 0% | 2,402 | 2,674 | +11% | 0 | 0 | — |
case-05 | pass→pass | 21,293 | 20,113 | -6% | 1 | 1 | 0% | 3,795 | 4,178 | +10% | 0 | 0 | — |
case-06 | pass→pass | 10,261 | 7,922 | -23% | 1 | 1 | 0% | 1,870 | 2,071 | +11% | 0 | 0 | — |
case-07 | fail→fail | 6,553 | 5,071 | -23% | 1 | 1 | 0% | 1,089 | 1,645 | +51% | 0 | 0 | — |
case-08 | pass→pass | 6,552 | 2,843 | -57% | 1 | 1 | 0% | 929 | 1,167 | +26% | 0 | 0 | — |
case-09 | fail→pass | 6,943 | 3,751 | -46% | 1 | 1 | 0% | 1,024 | 1,353 | +32% | 0 | 0 | — |
case-10 | pass→pass | 6,508 | 3,487 | -46% | 1 | 1 | 0% | 983 | 1,315 | +34% | 0 | 0 | — |
case-11 | fail→fail | 7,112 | 3,951 | -44% | 1 | 1 | 0% | 1,049 | 1,370 | +31% | 0 | 0 | — |
case-12 | fail→pass | 4,684 | 2,859 | -39% | 1 | 1 | 0% | 754 | 1,206 | +60% | 0 | 0 | — |
case-13 | fail→pass | 6,362 | 2,812 | -56% | 1 | 1 | 0% | 1,017 | 1,169 | +15% | 0 | 0 | — |
case-14 | fail→pass | 6,093 | 2,231 | -63% | 1 | 1 | 0% | 1,088 | 1,106 | +2% | 0 | 0 | — |
case-15 | pass→pass | 4,961 | 3,573 | -28% | 1 | 1 | 0% | 735 | 1,359 | +85% | 0 | 0 | — |
case-16 | fail→pass | 4,245 | 3,586 | -16% | 1 | 1 | 0% | 591 | 1,294 | +119% | 0 | 0 | — |
case-17 | fail→pass | 8,704 | 2,715 | -69% | 1 | 1 | 0% | 1,318 | 1,205 | -9% | 0 | 0 | — |
case-18 | fail→pass | 7,358 | 4,465 | -39% | 1 | 1 | 0% | 1,181 | 1,440 | +22% | 0 | 0 | — |
case-19 | fail→pass | 5,183 | 2,795 | -46% | 1 | 1 | 0% | 830 | 1,161 | +40% | 0 | 0 | — |
case-20 | pass→pass | 4,236 | 3,306 | -22% | 1 | 1 | 0% | 591 | 1,269 | +115% | 0 | 0 | — |
case-21 | fail→pass | 7,166 | 7,023 | -2% | 1 | 1 | 0% | 1,032 | 1,913 | +85% | 0 | 0 | — |
case-22 | fail→pass | 7,503 | 4,033 | -46% | 1 | 1 | 0% | 1,019 | 1,369 | +34% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.