Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Agent Immune System governance skill for maintaining long-term agent instruction health. Also trigger this when the user says "agent immune system" or "instruction health manager." Use before editing AGENTS.md, CLAUDE.md, global/project instructions, or SKILL.md files; after Hermes/review/postmortem findings; after long or failed sessions with instruction confusion; and when rules feel duplicated, bloated, stale, conflicting, autoimmune, or overfit. Produces promotion, demotion, pruning, regener
.claude/skills/inbusiness23-agent-immune-system/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 39% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 52% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 11% | 0% |
<what-to-do>
Maintain long-term instruction health. Treat AGENTS.md, CLAUDE.md, project instructions, skills, memory reports, issue reports, Hermes findings, and postmortems as inputs to the Agent Immune System.
Do not convert every mistake into a permanent rule. A mistake is an antigen, not automatic memory.
When invoked, produce a concrete instruction-health decision:
knowledge/incidents/, reports, issue history, review history, or prior instruction changes.Ask one question at a time only when the answer materially changes the decision and cannot be discovered from local files.
</what-to-do>
<supporting-info>
This skill is the immune system above task execution:
textPrimary Agent -> Project Agent -> Skills -> Execution Instruction Health Manager -> monitors all layers -> governs instruction memory -> proposes optimized instruction artifacts
Active instructions are the execution contract. This skill governs proposed changes to that contract so AGENTS.md, CLAUDE.md, and SKILL.md files remain the optimized expression of lessons that survived selection pressure.
Inspect only the sources needed for the current review:
AGENTS.md, CLAUDE.md, .codex/AGENTS.md, .claude/CLAUDE.mdskills/*/SKILL.md_reports/, docs/reports/, docs/postmortems/, issue reportsknowledge/ folders inside this skillNever harvest credentials, transcripts, shell history, keychain, or private tokens to reconstruct evidence.
State:
Reject active promotion quickly when any are true:
Promote slowly. Require evidence of recurrence, severity, and generality.
Default threshold:
| Signal | Action | |---|---| | 1 low/medium issue | Log only | | 2 similar issues | Watchlist | | 3 similar issues | Reference antibody candidate | | 4+ similar issues | Direct inclusion candidate | | 1 critical general issue | Immediate reference candidate; direct inclusion only with explicit unacceptable-delay risk | | Old rule with no recent hits | Prune or demote candidate |
Find instructions that create more problems than they solve:
Prefer cleaner principles over accumulated scars. Ask:
> Can five instructions be replaced by one higher-level principle?
Regeneration proposals should list the old rule cluster, the new principle, behavior preserved, behavior intentionally removed, and risks.
Choose the smallest sufficient destination:
knowledge/watchlists/ with promotion/demotion criteriaknowledge/incidents/ with "do not revisit unless pattern recurs"knowledge/reference-antibodies/Use gbrain as memory backend when available. gbrain remembers; this skill judges.
For narrow reviews, use quick-review mode:
Use templates/instruction-health-report.md as drafting scaffolding for broad audits and templates/promotion-proposal.md for active instruction changes.
For deep instruction audits, session reviews, postmortems, or architecture-style comparisons, write the final artifact as a single self-contained HTML file in the current project's _reports/ directory:
text_reports/report-instruction-health-YYYY-MM-DD.html
No external CDNs. Use the markdown template only to structure content before rendering the HTML report.
Every broad audit report must include:
Do not apply active instruction edits unless David explicitly asks in the current conversation.
Hook scripts live in hooks/:
claude-instruction-pretool-hook.sh: detects edits to active instruction files and emits a promotion-gate reminder.claude-instruction-posttool-hook.sh: logs instruction edits as audit candidates.codex-instruction-hook.sh: adapter script for Codex/manual invocation because Codex does not expose Claude-style native hook settings in this environment.See references/hook-integration.md before installing or changing hooks.
references/promotion-rubric.md: when deciding whether a lesson becomes active instruction memory.references/autoimmune-rubric.md: when evaluating harmful, conflicting, bloated, or overfit rules.references/regeneration-rubric.md: when compressing many rules into fewer principles.references/hook-integration.md: Claude/Codex hook wiring and trigger guidance.references/subagent-prompts.md: prompts for independent reviewers once the main skill framework exists.</supporting-info>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,088 | 6,369 | +5% | 1 | 1 | 0% | 301 | 1,896 | +530% | 0 | 0 | — |
case-02 | fail→fail | 4,219 | 6,714 | +59% | 1 | 1 | 0% | 196 | 2,049 | +945% | 0 | 0 | — |
case-03 | pass→pass | 10,619 | 5,753 | -46% | 1 | 1 | 0% | 1,718 | 2,426 | +41% | 0 | 0 | — |
case-04 | pass→pass | 15,266 | 8,817 | -42% | 1 | 1 | 0% | 2,333 | 2,908 | +25% | 0 | 0 | — |
case-05 | fail→pass | 11,941 | 7,981 | -33% | 1 | 1 | 0% | 1,818 | 2,776 | +53% | 0 | 0 | — |
case-06 | pass→pass | 8,197 | 6,886 | -16% | 1 | 1 | 0% | 1,340 | 2,475 | +85% | 0 | 0 | — |
case-07 | fail→pass | 11,691 | 8,059 | -31% | 1 | 1 | 0% | 2,086 | 2,812 | +35% | 0 | 0 | — |
case-08 | pass→pass | 9,603 | 10,494 | +9% | 1 | 1 | 0% | 1,883 | 3,115 | +65% | 0 | 0 | — |
case-09 | pass→pass | 15,797 | 9,013 | -43% | 1 | 1 | 0% | 2,574 | 2,810 | +9% | 0 | 0 | — |
case-10 | pass→pass | 11,223 | 6,912 | -38% | 1 | 1 | 0% | 1,736 | 2,646 | +52% | 0 | 0 | — |
case-11 | fail→pass | 9,241 | 4,231 | -54% | 1 | 1 | 0% | 1,630 | 2,271 | +39% | 0 | 0 | — |
case-12 | fail→fail | 16,471 | 5,754 | -65% | 1 | 1 | 0% | 2,521 | 1,864 | -26% | 0 | 0 | — |
case-13 | pass→pass | 12,859 | 9,110 | -29% | 1 | 1 | 0% | 2,003 | 2,142 | +7% | 0 | 0 | — |
case-14 | pass→pass | 11,485 | 6,799 | -41% | 1 | 1 | 0% | 1,783 | 2,596 | +46% | 0 | 0 | — |
case-15 | pass→pass | 13,220 | 7,400 | -44% | 1 | 1 | 0% | 2,111 | 2,958 | +40% | 0 | 0 | — |
case-16 | pass→fail | 5,653 | 10,691 | +89% | 1 | 1 | 0% | 571 | 3,453 | +505% | 0 | 0 | — |
case-17 | fail→pass | 9,470 | 4,677 | -51% | 1 | 1 | 0% | 1,546 | 2,348 | +52% | 0 | 0 | — |
case-18 | fail→pass | 12,584 | 3,962 | -69% | 1 | 1 | 0% | 2,020 | 2,236 | +11% | 0 | 0 | — |
case-19 | pass→fail | 13,155 | 2,448 | -81% | 1 | 1 | 0% | 2,187 | 1,895 | -13% | 0 | 0 | — |
case-20 | fail→pass | 9,277 | 6,967 | -25% | 1 | 1 | 0% | 1,666 | 2,715 | +63% | 0 | 0 | — |
case-21 | pass→pass | 18,701 | 6,581 | -65% | 1 | 1 | 0% | 1,272 | 2,526 | +99% | 0 | 0 | — |
case-22 | fail→pass | 9,167 | 6,688 | -27% | 1 | 1 | 0% | 1,448 | 2,753 | +90% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/3/2026 | +59% |
Other measured skills in the registry, with their headline benchmark lift.