Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review generated or changed documentation before it ships, including READMEs, API references, docstrings, changelogs, tutorials, and documentation sites.
.claude/skills/sickn33-docs-guard/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 192% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 177% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 96% | 0% |
You are reviewing generated or changed documentation before it ships. Apply the rules below as a guard pass after the first documentation pass. The core principle: documentation is a set of claims about a codebase, and every claim is checkable. Your job is to check them.
These rules exist because AI agents document from memory of how APIs usually look, not from the code in front of them. Published research: half of AI answers to programming questions contain incorrect information, and models produce valid invocations for infrequent APIs barely a third of the time — yet the prose sounds authoritative either way. Readers cannot tell verified docs from hallucinated docs. You can, because you have the source.
Use this skill when reviewing generated or changed documentation before it ships. Activate it reactively after an agent writes or updates READMEs, API references, docstrings, PHPDoc/JSDoc, changelogs, tutorials, or doc sites.
Guard-pass mode (recommended): after documentation or docstrings have been generated or edited, verify every claim against the source and run the self-check before delivery.
Live mode (explicit): when the user invokes this skill before writing docs, verify before you write — read the actual implementation, then document what it does. Run the self-check before delivery.
Review mode (the user asks you to review, audit, or fact-check docs): walk references/review-checklist.md against the target docs and produce a findings report with file:line evidence. Do not rewrite in review mode unless asked.
get_user_by_id), sections that restate their heading, marketing adjectives in technical prose ("powerful", "seamless", "blazingly fast"), and intro padding ("In this section, we will explore…"). A docstring earns its place by adding contracts the signature cannot express: units, ranges, error conditions, side effects, threading/ordering guarantees.If any answer is wrong, fix it before showing the user.
**Rule N violation** in `docs/path.md:<line or section>`
- Claim: <what the docs say>
- Reality: <what the code/CLI/schema actually has, with file:line>
- Fix: <one sentence>Lead with Rule 1–4 findings (false claims), then drift, then substance. If a doc is clean, say so in one line — accuracy deserves credit.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,000 | 1,473 | -51% | 1 | 1 | 0% | 355 | 1,935 | +445% | 0 | 0 | — |
case-02 | fail→fail | 3,110 | 5,926 | +91% | 1 | 1 | 0% | 482 | 2,065 | +328% | 0 | 0 | — |
case-08 | fail→fail | 8,501 | 3,306 | -61% | 1 | 1 | 0% | 1,515 | 2,098 | +38% | 0 | 0 | — |
case-03 | fail→fail | 4,807 | 4,359 | -9% | 1 | 1 | 0% | 247 | 2,009 | +713% | 0 | 0 | — |
case-04 | pass→fail | 10,706 | 7,238 | -32% | 1 | 1 | 0% | 1,989 | 2,905 | +46% | 0 | 0 | — |
case-05 | pass→pass | 18,418 | 12,752 | -31% | 1 | 1 | 0% | 3,166 | 3,731 | +18% | 0 | 0 | — |
case-06 | fail→fail | 3,906 | 3,456 | -12% | 1 | 1 | 0% | 633 | 2,263 | +258% | 0 | 0 | — |
case-07 | fail→fail | 4,287 | 2,268 | -47% | 1 | 1 | 0% | 721 | 2,028 | +181% | 0 | 0 | — |
case-09 | fail→fail | 23,171 | 5,098 | -78% | 1 | 1 | 0% | 248 | 1,990 | +702% | 0 | 0 | — |
case-10 | fail→fail | 11,888 | 2,801 | -76% | 1 | 1 | 0% | 2,020 | 1,999 | -1% | 0 | 0 | — |
case-11 | fail→fail | 4,111 | 9,961 | +142% | 1 | 1 | 0% | 546 | 3,506 | +542% | 0 | 0 | — |
case-12 | fail→pass | 5,145 | 5,126 | -0% | 1 | 1 | 0% | 900 | 2,631 | +192% | 0 | 0 | — |
case-13 | fail→pass | 13,771 | 7,254 | -47% | 1 | 1 | 0% | 2,331 | 2,991 | +28% | 0 | 0 | — |
case-14 | fail→fail | 8,027 | 2,403 | -70% | 1 | 1 | 0% | 1,458 | 1,994 | +37% | 0 | 0 | — |
case-15 | fail→pass | 8,435 | 7,666 | -9% | 1 | 1 | 0% | 1,549 | 3,214 | +107% | 0 | 0 | — |
case-16 | fail→pass | 6,662 | 8,613 | +29% | 1 | 1 | 0% | 1,098 | 3,040 | +177% | 0 | 0 | — |
case-17 | pass→fail | 8,845 | 5,083 | -43% | 1 | 1 | 0% | 1,299 | 2,538 | +95% | 0 | 0 | — |
case-18 | fail→fail | 16,970 | 2,232 | -87% | 1 | 1 | 0% | 2,528 | 2,033 | -20% | 0 | 0 | — |
case-19 | fail→fail | 9,222 | 6,679 | -28% | 1 | 1 | 0% | 1,702 | 3,243 | +91% | 0 | 0 | — |
case-20 | fail→pass | 8,306 | 5,539 | -33% | 1 | 1 | 0% | 1,322 | 2,587 | +96% | 0 | 0 | — |
case-21 | pass→fail | 2,517 | 6,655 | +164% | 1 | 1 | 0% | 394 | 2,731 | +593% | 0 | 0 | — |
case-22 | fail→fail | 5,371 | 2,098 | -61% | 1 | 1 | 0% | 807 | 2,004 | +148% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 20 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.