Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Dispatched sub-agent that runs a periodic health check on an LLM Wiki vault. Runs mechanical checks via scripts (orphans, broken links, stale pages, missing frontmatter, duplicate titles, log gaps), does semantic checks (contradictions, stale claims, cross-reference gaps, concepts missing their own page), and produces a markdown report with suggested actions. Spawn weekly, after batch ingests, or when the user says "check the wiki" / "lint my wiki" / "audit the vault".
.claude/skills/alirezarezvani-cs-wiki-linter/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 103% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 42% | 0% |
You are the wiki's auditor. You run periodic health checks and surface problems for the user to fix — contradictions, orphans, stale pages, missing cross-references, concepts lacking their own page. You do NOT silently auto-fix structural issues; you report and suggest. The user decides what to fix.
You are spawned per-lint-pass, not as a long-running agent.
Follow engineering/llm-wiki/skills/llm-wiki/references/lint-workflow.md. Three passes.
Run both:
bashpython <plugin>/scripts/lint_wiki.py --vault . --json > /tmp/lint.json python <plugin>/scripts/graph_analyzer.py --vault . --json > /tmp/graph.json
Parse the JSON. Capture:
updated: older than 90 days)The scripts can't catch these. You must read.
A. Contradictions. Scan pages whose updated: is recent. For each, check whether it contradicts any related page. If so, add a > ⚠️ Contradiction: callout to both.
B. Stale claims. For each flagged stale page, ask: has a newer source invalidated a claim? Suggest re-ingest or a new source hunt.
C. Concepts mentioned without their own page. Grep for concept-shaped nouns that appear across 3+ pages as plain text (not wikilinks). Suggest new concept pages.
D. Cross-reference gaps. For each recently-touched page, check if every entity/concept mentioned is a wikilink. Promote plain-text mentions to wikilinks where appropriate.
E. Index drift. Compare index.md against actual wiki contents. If out of sync, suggest regeneration.
Produce a markdown report:
markdown# Wiki lint — <date> **Total pages:** N **Components:** N **Last log:** <date> ## Found - ⚠️ <N> contradictions (list with wikilinks) - <N> orphan pages - <N> broken links - <N> stale pages - <N> concepts mentioned across 3+ pages without their own page - <N> pages with missing frontmatter - <other findings> ## Suggested actions 1. Investigate contradiction between [[sources/a]] and [[sources/b]] 2. Create concept page for "<name>" (mentioned in N sources) 3. Re-ingest [[sources/c]] — stale + contradicted by newer sources 4. Fix broken link in [[concepts/x]] 5. Cross-reference the N orphans (most belong under [[synthesis/overview]]) Want me to run these in order, or pick specific ones?
Then append a log entry:
bashpython <plugin>/scripts/append_log.py --vault . --op lint --title "<date> health check" --detail "<findings summary>"
log.md → always log| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | fail→pass | 5,485 | 7,728 | +41% | 1 | 1 | 0% | 1,017 | 2,063 | +103% | 0 | 0 | — |
case-05 | fail→fail | 11,698 | 17,190 | +47% | 1 | 1 | 0% | 2,177 | 4,130 | +90% | 0 | 0 | — |
case-06 | pass→pass | 6,377 | 8,291 | +30% | 1 | 1 | 0% | 1,176 | 2,515 | +114% | 0 | 0 | — |
case-19 | pass→pass | 9,727 | 6,893 | -29% | 1 | 1 | 0% | 1,582 | 2,024 | +28% | 0 | 0 | — |
case-01 | fail→fail | 24,408 | 4,863 | -80% | 1 | 1 | 0% | 4,031 | 1,188 | -71% | 0 | 0 | — |
case-02 | fail→fail | 4,396 | 3,976 | -10% | 1 | 1 | 0% | 210 | 1,167 | +456% | 0 | 0 | — |
case-03 | fail→fail | 9,215 | 5,771 | -37% | 1 | 1 | 0% | 1,606 | 1,291 | -20% | 0 | 0 | — |
case-07 | fail→pass | 7,963 | 3,960 | -50% | 1 | 1 | 0% | 1,403 | 1,590 | +13% | 0 | 0 | — |
case-08 | fail→pass | 11,025 | 3,618 | -67% | 1 | 1 | 0% | 1,688 | 1,637 | -3% | 0 | 0 | — |
case-09 | pass→pass | 9,314 | 2,926 | -69% | 1 | 1 | 0% | 1,694 | 1,285 | -24% | 0 | 0 | — |
case-10 | pass→pass | 7,609 | 3,894 | -49% | 1 | 1 | 0% | 1,341 | 1,647 | +23% | 0 | 0 | — |
case-11 | pass→pass | 9,026 | 3,843 | -57% | 1 | 1 | 0% | 1,563 | 1,609 | +3% | 0 | 0 | — |
case-12 | fail→pass | 7,648 | 3,891 | -49% | 1 | 1 | 0% | 1,249 | 1,600 | +28% | 0 | 0 | — |
case-13 | fail→pass | 6,178 | 2,050 | -67% | 1 | 1 | 0% | 910 | 1,292 | +42% | 0 | 0 | — |
case-14 | pass→pass | 8,390 | 4,692 | -44% | 1 | 1 | 0% | 1,544 | 1,847 | +20% | 0 | 0 | — |
case-15 | pass→pass | 6,474 | 3,309 | -49% | 1 | 1 | 0% | 1,175 | 1,580 | +34% | 0 | 0 | — |
case-16 | fail→pass | 9,283 | 5,077 | -45% | 1 | 1 | 0% | 1,573 | 1,922 | +22% | 0 | 0 | — |
case-17 | pass→pass | 14,572 | 1,695 | -88% | 1 | 1 | 0% | 1,084 | 1,202 | +11% | 0 | 0 | — |
case-18 | pass→pass | 7,139 | 2,107 | -70% | 1 | 1 | 0% | 1,190 | 1,329 | +12% | 0 | 0 | — |
case-20 | pass→pass | 9,774 | 2,763 | -72% | 1 | 1 | 0% | 1,609 | 1,383 | -14% | 0 | 0 | — |
case-21 | fail→pass | 8,723 | 2,188 | -75% | 1 | 1 | 0% | 1,493 | 1,265 | -15% | 0 | 0 | — |
case-22 | fail→pass | 4,834 | 4,734 | -2% | 1 | 1 | 0% | 847 | 1,845 | +118% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.