Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Loop-breaking self-diagnosis. Use when 3+ consecutive failures occur, circular retries persist, or context overwhelms the session.
.claude/skills/hashgraph-online-agent-introspection/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 110% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 6% | 0% |
NO UNSUPPORTED SELF-HEALING CLAIMS. Only assert recovery actions you can actually perform with available tools.
Stop immediately and record:
Match against known patterns:
| Pattern | Symptoms | Root Cause | |---------|----------|------------| | Loop trap | Same error 3+ times | Wrong approach, not wrong parameters | | Context overflow | Increasingly confused responses | Too much information, need compaction | | Environment drift | "Works locally" failures | Missing env var, different tool version | | Cascade failure | Fix A breaks B | Underlying assumption is wrong | | Tool mismatch | Wrong tool for the job | Need a different approach entirely |
Execute ONLY the smallest safe action:
/compact or summarize current state, then continue.which, --version).Edit fails 3 times, try Write. If Bash fails, try Read first.Generate a structured report:
markdown## Introspection Report - **Failure type**: [type] - **Root cause**: [diagnosis] - **Recovery action**: [what was done] - **Confidence**: [high/medium/low] - **Next step**: [what to do if this happens again]
Save to memory:
bashepic mem add \ --title "Self-diagnosis: {error_type} in {file}" \ --type error \ --body "Pattern: ...\nRoot cause: ...\nRecovery: ...\n"
| Excuse | Rebuttal | What to do instead | |--------|----------|-------------------| | "One more try might work" | 3 failures means the approach is wrong, not unlucky. | Stop and run the full 4-step introspection process. | | "I just need to tweak the parameters" | Tweaking a failing approach is not debugging. Step back and reassess. | Abandon the current approach and try a fundamentally different strategy. | | "I can fix this myself" | Asking for help is not weakness. Escalation saves everyone time. | Escalate to the user with a clear summary of what was tried and what failed. | | "The error is clear, I know the fix" | You said that 3 times already. Prove it with a different approach. | Run the introspection report and verify the fix with a test before claiming success. |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,837 | 17,514 | +4% | 1 | 1 | 0% | 1,384 | 3,004 | +117% | 0 | 0 | — |
case-02 | fail→pass | 14,265 | 11,256 | -21% | 1 | 1 | 0% | 1,331 | 2,790 | +110% | 0 | 0 | — |
case-03 | fail→fail | 18,955 | 30,069 | +59% | 1 | 1 | 0% | 2,147 | 3,053 | +42% | 0 | 0 | — |
case-04 | fail→pass | 21,233 | 14,437 | -32% | 1 | 1 | 0% | 2,076 | 2,524 | +22% | 0 | 0 | — |
case-05 | fail→fail | 23,004 | 11,296 | -51% | 1 | 1 | 0% | 2,401 | 2,695 | +12% | 0 | 0 | — |
case-06 | fail→fail | 21,990 | 15,224 | -31% | 1 | 1 | 0% | 2,340 | 2,327 | -1% | 0 | 0 | — |
case-07 | fail→pass | 14,723 | 13,803 | -6% | 1 | 1 | 0% | 1,288 | 2,029 | +58% | 0 | 0 | — |
case-08 | fail→pass | 18,404 | 7,953 | -57% | 1 | 1 | 0% | 2,064 | 2,188 | +6% | 0 | 0 | — |
case-09 | fail→pass | 22,057 | 17,778 | -19% | 1 | 1 | 0% | 2,111 | 2,220 | +5% | 0 | 0 | — |
case-10 | fail→pass | 13,215 | 8,253 | -38% | 1 | 1 | 0% | 1,675 | 2,204 | +32% | 0 | 0 | — |
case-11 | fail→pass | 6,957 | 7,932 | +14% | 1 | 1 | 0% | 973 | 2,229 | +129% | 0 | 0 | — |
case-12 | fail→pass | 12,248 | 8,206 | -33% | 1 | 1 | 0% | 1,637 | 1,427 | -13% | 0 | 0 | — |
case-13 | fail→fail | 13,602 | 4,364 | -68% | 1 | 1 | 0% | 1,343 | 1,486 | +11% | 0 | 0 | — |
case-14 | fail→fail | 15,478 | 11,042 | -29% | 1 | 1 | 0% | 1,389 | 1,883 | +36% | 0 | 0 | — |
case-15 | fail→pass | 20,201 | 9,064 | -55% | 1 | 1 | 0% | 1,998 | 2,357 | +18% | 0 | 0 | — |
case-16 | fail→pass | 13,187 | 9,405 | -29% | 1 | 1 | 0% | 1,357 | 1,692 | +25% | 0 | 0 | — |
case-17 | fail→pass | 10,049 | 10,141 | +1% | 1 | 1 | 0% | 1,639 | 2,297 | +40% | 0 | 0 | — |
case-18 | fail→fail | 17,114 | 4,137 | -76% | 1 | 1 | 0% | 1,598 | 1,342 | -16% | 0 | 0 | — |
case-19 | pass→fail | 22,453 | 19,455 | -13% | 1 | 1 | 0% | 3,713 | 4,031 | +9% | 0 | 0 | — |
case-20 | pass→pass | 13,408 | 11,840 | -12% | 1 | 1 | 0% | 1,646 | 2,130 | +29% | 0 | 0 | — |
case-21 | pass→pass | 18,130 | 21,527 | +19% | 1 | 1 | 0% | 2,488 | 3,468 | +39% | 0 | 0 | — |
case-22 | fail→fail | 15,262 | 8,191 | -46% | 1 | 1 | 0% | 1,334 | 1,938 | +45% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.