Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when auditing DX, docs, and delivery in any codebase — the 30-minute-contributor test, docs-vs-reality drift, onboarding friction, error-message quality. Assess by default, fix docs on request; docs are judged by use, and every failure is a misfiled or missing Diátaxis quadrant.
.claude/skills/automagik-dev-dx-docs/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -22% | 0% |
| case-22 | ✓→✗ | ▼ Worse | 231% | 0% |
Runtime syntax: invoke the plugin copy through the active runtime's owner-qualified skill selector; use a bare selector only when intentionally selecting a user-tier copy (a separately installed personal copy; Genie no longer seeds this tier). Cross-skill prose below uses bare names as portable semantic routes; the orchestrator resolves the selector for the active runtime.
This lane treats documentation as four different things — tutorials (learning-oriented), how-to guides (task-oriented), reference (information-oriented), explanation (understanding-oriented) — and nearly every documentation failure is one quadrant's content misfiled in another, or a quadrant missing entirely. Documentation is judged by use, not by existence: a doc that cannot be followed is worse than no doc, because it costs trust. Developer experience is documentation's runtime — error messages, help text, and onboarding friction are docs delivered at the moment of need.
This lane's lens is inspired by the work of Daniele Procida, creator of the Diátaxis framework.
Assess and report by default. Apply doc fixes only when the invocation explicitly asks — and only through the repo's documented docs workflow if it has one (submodules, docs repos, review gates). Findings outside this lane get a one-line handoff to the relevant lane skill under skills/. Judgments come from using the docs and the product, never from reading them approvingly.
Map the docs estate before judging it: where docs live (in-repo, submodule, separate site), which are public vs internal, what the contribution/onboarding path claims to be (README, CONTRIBUTING, CLAUDE.md/AGENTS.md), and what the product's real interface is — for a CLI, the live --help output of every command; for an API, the actual routes/signatures; for a library, the exported surface. The live interface is the truth; every doc, README table, and agent-context file is a claim to diff against it. Note the repo's stated DX bar (e.g. a 30-minute-contributor promise) — hold it to its own standard.
Genie-framework repos: the lifecycle skills (brainstorm → wish → work → review, plus their kin) are part of the user-facing surface. Their SKILL.md descriptions, inputs, and outputs must chain coherently — does wish consume what brainstorm produces, does review validate what work emits — and match what the docs claim about them.
Repo profile — recall, verify, persist. Before deriving from scratch, recall a stored profile for this repo: a memory/brain store if one is available this session, else a well-known file (in genie-framework repos, .genie/repo-profile.md). For this lane the profile records the docs topology, the live-interface inventory, past stumble logs, and open drift findings. Recalled drift may have been fixed since — re-check each entry against the live interface before reporting, and report new drift as a finding. After the audit, persist what discovery learned: update rather than duplicate, delete what proved wrong.
Profile write boundary. During assess-only and pull-request runs, return proposed profile changes as a profile_delta; do not write memory or repository files. Persist a profile only when the user explicitly asks.
Every drift claim quotes both sides; every stumble names the exact step and what actually happened; skipped steps (e.g. no fresh clone was feasible) are stated, with affected conclusions marked partial.
Lead with a one-sentence verdict: did the repo pass its contributor test, and what is the worst drift. Then findings ranked as above with evidence and concrete fixes — routed through the repo's docs workflow where one exists. Include what works well; a review that only lists friction misleads. In a genie-framework repo, use CRITICAL/HIGH/MEDIUM/LOW for finding severities and SHIP/FIX-FIRST/BLOCKED only for the overall verdict and offer to crystallize a docs-overhaul into a wish via wish.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | pass→pass | 9,878 | 9,230 | -7% | 1 | 1 | 0% | 1,432 | 2,912 | +103% | 0 | 0 | — |
case-14 | pass→pass | 8,989 | 8,733 | -3% | 1 | 1 | 0% | 1,397 | 2,808 | +101% | 0 | 0 | — |
case-01 | fail→fail | 5,260 | 7,402 | +41% | 1 | 1 | 0% | 314 | 1,855 | +491% | 0 | 0 | — |
case-02 | fail→fail | 6,342 | 6,249 | -1% | 1 | 1 | 0% | 372 | 1,867 | +402% | 0 | 0 | — |
case-03 | fail→fail | 4,645 | 6,603 | +42% | 1 | 1 | 0% | 182 | 1,768 | +871% | 0 | 0 | — |
case-04 | pass→fail | 15,153 | 6,873 | -55% | 1 | 1 | 0% | 2,445 | 1,918 | -22% | 0 | 0 | — |
case-05 | fail→pass | 13,177 | 5,413 | -59% | 1 | 1 | 0% | 2,060 | 2,227 | +8% | 0 | 0 | — |
case-06 | pass→pass | 13,603 | 10,209 | -25% | 1 | 1 | 0% | 2,014 | 3,022 | +50% | 0 | 0 | — |
case-07 | pass→pass | 10,700 | 7,277 | -32% | 1 | 1 | 0% | 1,666 | 2,577 | +55% | 0 | 0 | — |
case-08 | pass→pass | 9,471 | 16,963 | +79% | 1 | 1 | 0% | 1,358 | 2,632 | +94% | 0 | 0 | — |
case-09 | pass→pass | 10,277 | 3,576 | -65% | 1 | 1 | 0% | 1,633 | 1,974 | +21% | 0 | 0 | — |
case-10 | pass→pass | 9,122 | 4,444 | -51% | 1 | 1 | 0% | 1,355 | 2,183 | +61% | 0 | 0 | — |
case-11 | pass→pass | 8,264 | 5,412 | -35% | 1 | 1 | 0% | 1,381 | 2,244 | +62% | 0 | 0 | — |
case-12 | pass→pass | 9,862 | 4,548 | -54% | 1 | 1 | 0% | 1,347 | 2,136 | +59% | 0 | 0 | — |
case-15 | pass→pass | 8,876 | 4,509 | -49% | 1 | 1 | 0% | 1,278 | 2,116 | +66% | 0 | 0 | — |
case-16 | fail→pass | 12,830 | 3,510 | -73% | 1 | 1 | 0% | 1,817 | 1,931 | +6% | 0 | 0 | — |
case-17 | fail→pass | 14,691 | 5,405 | -63% | 1 | 1 | 0% | 2,020 | 2,269 | +12% | 0 | 0 | — |
case-18 | pass→pass | 12,068 | 7,525 | -38% | 1 | 1 | 0% | 1,800 | 2,570 | +43% | 0 | 0 | — |
case-19 | pass→pass | 13,657 | 8,492 | -38% | 1 | 1 | 0% | 1,947 | 2,661 | +37% | 0 | 0 | — |
case-20 | fail→fail | 5,451 | 2,421 | -56% | 1 | 1 | 0% | 1,067 | 1,822 | +71% | 0 | 0 | — |
case-21 | pass→pass | 12,408 | 13,319 | +7% | 1 | 1 | 0% | 2,359 | 3,687 | +56% | 0 | 0 | — |
case-22 | pass→fail | 5,082 | 8,421 | +66% | 1 | 1 | 0% | 880 | 2,911 | +231% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +5 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.