Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Sutando introspection — read logs + git + memory + build log for a chosen time window and produce a concise narrative of what the agent has been doing, what's broken, and what to prioritize next.
.claude/skills/sonichi-self-diagnose/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-10 | ✓→✓ | = Same ✓ | 13% | 0% |
| case-11 | ✓→✓ | = Same ✓ | 1% | 0% |
| case-12 | ✓→✓ | = Same ✓ | -9% | 0% |
Read Sutando's own observable state (logs, git, memory, build log, pending questions, health check, cold-review log) over a chosen window and produce a structured narrative:
Usage: /self-diagnose [--since 24h]
ARGUMENTS: $ARGUMENTS
bash skills/self-diagnose/scripts/gather.sh [window] — collects log tails, git log, build_log.md tail, pending-questions, health-check output, cold-review-log, and recent Discord activity into /tmp/sutando-diagnose-<ts>/.notes/diagnose-YYYY-MM-DD-HHMM.md with frontmatter (title, date, tags: [diagnose, self]).When Sutando works on one machine but is broken on another, the helper script runs gather.sh on both sides over SSH and produces a structured comparison:
bashbash skills/self-diagnose/scripts/gather-remote.sh <ssh-target> [window] # e.g. bash skills/self-diagnose/scripts/gather-remote.sh mac-mini bash skills/self-diagnose/scripts/gather-remote.sh user@macbook.local 6h
Output: /tmp/sutando-diagnose-cross-<ts>/{local,remote,diff.md} plus a persisted copy at notes/diagnose-cross-node-<YYYY-MM-DD>.md. The comparison report surfaces commit drift, health-check differences, voice-agent error counts, quota state, and PR-view divergence per side.
Override the remote sutando path via SUTANDO_REMOTE_REPO=/path/on/peer if the failing node's checkout isn't at the default ~/Desktop/sutando.
Security posture (per #421): read-only on the remote (only gather.sh runs there, no mutation), allowlist enforced by gather.sh's existing scope (no .env, no tokens), per-session SSH (no daemon, no persistent state on the remote).
24 hours if unspecified. Accept: 24h, 3d, 1w. Longer windows = broader scope, higher token cost.
18:28:16 — transport 1006 after NoteView injection) rather than "some errors happened"#394 mergeable no reviews, #354 retention sweep open)notes/cold-review-capability.md — sister capability for PR-level reviewnotes/voice-transport-1006-hypothesis.md — example of the depth of analysis a self-diagnose report should matchbuild_log.md — the canonical "what has been built" log; complements but doesn't replace| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | pass→pass | 10,232 | 5,913 | -42% | 1 | 1 | 0% | 1,701 | 1,925 | +13% | 0 | 0 | — |
case-11 | pass→pass | 10,391 | 4,199 | -60% | 1 | 1 | 0% | 1,635 | 1,646 | +1% | 0 | 0 | — |
case-01 | fail→fail | 12,170 | 6,767 | -44% | 1 | 1 | 0% | 1,817 | 1,277 | -30% | 0 | 0 | — |
case-02 | fail→fail | 8,575 | 4,420 | -48% | 1 | 1 | 0% | 440 | 1,225 | +178% | 0 | 0 | — |
case-03 | fail→fail | 24,979 | 5,038 | -80% | 1 | 1 | 0% | 4,038 | 1,269 | -69% | 0 | 0 | — |
case-04 | fail→fail | 8,852 | 4,683 | -47% | 1 | 1 | 0% | 1,394 | 1,183 | -15% | 0 | 0 | — |
case-05 | fail→fail | 5,232 | 5,643 | +8% | 1 | 1 | 0% | 739 | 1,275 | +73% | 0 | 0 | — |
case-06 | fail→pass | 8,115 | 3,310 | -59% | 1 | 1 | 0% | 1,384 | 1,621 | +17% | 0 | 0 | — |
case-07 | fail→fail | 29,840 | 2,217 | -93% | 1 | 1 | 0% | 1,822 | 1,360 | -25% | 0 | 0 | — |
case-08 | fail→fail | 16,524 | 3,894 | -76% | 1 | 1 | 0% | 2,636 | 1,606 | -39% | 0 | 0 | — |
case-09 | fail→fail | 8,337 | 5,013 | -40% | 1 | 1 | 0% | 1,439 | 1,209 | -16% | 0 | 0 | — |
case-12 | pass→pass | 9,415 | 3,165 | -66% | 1 | 1 | 0% | 1,567 | 1,433 | -9% | 0 | 0 | — |
case-13 | pass→pass | 8,370 | 2,474 | -70% | 1 | 1 | 0% | 1,328 | 1,349 | +2% | 0 | 0 | — |
case-14 | pass→pass | 11,198 | 4,656 | -58% | 1 | 1 | 0% | 1,730 | 1,728 | -0% | 0 | 0 | — |
case-15 | pass→pass | 10,888 | 5,047 | -54% | 1 | 1 | 0% | 1,570 | 1,638 | +4% | 0 | 0 | — |
case-16 | fail→pass | 10,157 | 3,761 | -63% | 1 | 1 | 0% | 1,567 | 1,584 | +1% | 0 | 0 | — |
case-17 | fail→fail | 12,268 | 4,399 | -64% | 1 | 1 | 0% | 1,821 | 1,199 | -34% | 0 | 0 | — |
case-18 | fail→fail | 13,914 | 5,247 | -62% | 1 | 1 | 0% | 2,685 | 1,204 | -55% | 0 | 0 | — |
case-19 | pass→pass | 13,573 | 5,726 | -58% | 1 | 1 | 0% | 2,092 | 1,814 | -13% | 0 | 0 | — |
case-20 | fail→fail | 5,387 | 6,088 | +13% | 1 | 1 | 0% | 147 | 1,249 | +750% | 0 | 0 | — |
case-21 | fail→fail | 11,579 | 5,629 | -51% | 1 | 1 | 0% | 2,169 | 1,123 | -48% | 0 | 0 | — |
case-22 | fail→fail | 11,181 | 26,406 | +136% | 1 | 1 | 0% | 1,442 | 3,121 | +116% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 12 counted toward the lift figure. The other 10 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 12 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.