Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create or revise ContextOS MDX documentation while preserving frontmatter, foundation templates, navigation, component registration, canonical examples, URL stability, and spec/reference alignment. Use for docs pages, not blog posts or contract implementation alone.
.claude/skills/contextosai-contextos-docs-author/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 124% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 68% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-08 | ✓→✗ | ▼ Worse | -61% | 0% |
Publish a documentation page that fits the reader journey and does not create spec drift.
Read the repository instructions, the nearest related docs, current navigation source, MDX loader/renderer, component registry, redirects, and relevant tests. Verify actual imports before following a possibly stale statement about where navigation is hardcoded; if repository guidance conflicts with code, report the conflict rather than editing both blindly.
Read references/docs-map.md for page classes, coupled files, and validation routing.
title and description frontmatter.## sections: Definition, Why it exists, How it works, Interfaces, Failure modes, Operational concerns, Evaluation metrics, Example, Common misconceptions. Do not expand an allowlist to avoid the template.Optional metadata such as status, review date, primitive planes, lifecycle, inputs, outputs, and key types should match neighboring pages and real consumers.
Run the frontmatter and docs-template tests for content changes. Run spec-reference drift tests when canonical examples or runtime terms change; MDX loader tests when rendering changes; typecheck/lint/build when TypeScript, components, navigation, or routes change. Preview the affected route when practical.
Report the page's role in the reader journey, integration points changed, and any preview/build/live check not performed.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,598 | 4,548 | -1% | 1 | 1 | 0% | 216 | 811 | +275% | 0 | 0 | — |
case-02 | fail→fail | 11,345 | 5,540 | -51% | 1 | 1 | 0% | 236 | 795 | +237% | 0 | 0 | — |
case-03 | fail→fail | 12,146 | 5,519 | -55% | 1 | 1 | 0% | 2,207 | 829 | -62% | 0 | 0 | — |
case-04 | fail→fail | 8,190 | 6,794 | -17% | 1 | 1 | 0% | 1,273 | 886 | -30% | 0 | 0 | — |
case-05 | fail→fail | 6,177 | 6,797 | +10% | 1 | 1 | 0% | 910 | 890 | -2% | 0 | 0 | — |
case-06 | fail→fail | 11,876 | 9,036 | -24% | 1 | 1 | 0% | 2,029 | 1,067 | -47% | 0 | 0 | — |
case-07 | fail→fail | 14,131 | 5,172 | -63% | 1 | 1 | 0% | 2,377 | 880 | -63% | 0 | 0 | — |
case-08 | pass→fail | 14,600 | 3,953 | -73% | 1 | 1 | 0% | 2,097 | 825 | -61% | 0 | 0 | — |
case-09 | fail→fail | 10,141 | 4,746 | -53% | 1 | 1 | 0% | 1,603 | 866 | -46% | 0 | 0 | — |
case-10 | fail→pass | 19,979 | 44,090 | +121% | 1 | 1 | 0% | 3,535 | 7,925 | +124% | 0 | 0 | — |
case-11 | pass→pass | 10,745 | 9,849 | -8% | 1 | 1 | 0% | 1,406 | 1,644 | +17% | 0 | 0 | — |
case-12 | fail→pass | 9,809 | 11,364 | +16% | 1 | 1 | 0% | 1,515 | 2,551 | +68% | 0 | 0 | — |
case-13 | fail→fail | 8,233 | 9,981 | +21% | 1 | 1 | 0% | 1,340 | 1,049 | -22% | 0 | 0 | — |
case-14 | fail→fail | 12,932 | 7,615 | -41% | 1 | 1 | 0% | 2,188 | 1,101 | -50% | 0 | 0 | — |
case-15 | fail→pass | 9,763 | 11,741 | +20% | 1 | 1 | 0% | 1,377 | 1,856 | +35% | 0 | 0 | — |
case-16 | fail→pass | 8,583 | 8,530 | -1% | 1 | 1 | 0% | 1,347 | 1,801 | +34% | 0 | 0 | — |
case-17 | pass→fail | 10,695 | 6,500 | -39% | 1 | 1 | 0% | 1,705 | 930 | -45% | 0 | 0 | — |
case-18 | pass→fail | 9,909 | 7,194 | -27% | 1 | 1 | 0% | 1,564 | 807 | -48% | 0 | 0 | — |
case-19 | fail→fail | 7,132 | 13,209 | +85% | 1 | 1 | 0% | 1,035 | 963 | -7% | 0 | 0 | — |
case-20 | pass→fail | 15,165 | 6,134 | -60% | 1 | 1 | 0% | 2,666 | 903 | -66% | 0 | 0 | — |
case-21 | pass→fail | 14,374 | 5,958 | -59% | 1 | 1 | 0% | 3,066 | 854 | -72% | 0 | 0 | — |
case-22 | pass→fail | 8,605 | 4,181 | -51% | 1 | 1 | 0% | 1,594 | 868 | -46% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 5 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -9 percentage points is the difference between those two pass rates over the 5 comparable cases. 10 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.