Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Design, implement, backfill, audit, and release in-game changelogs with contiguous versioning, deployment provenance, menu-state navigation, accessible toggle, close, and Escape behavior, and responsive release-ledger UI. Use when Codex needs to add or revise a changelog or version screen in a game, reconstruct release history from deployments and Git, define version-bump rules, keep displayed versions synchronized with live builds, or test changelog mechanics across desktop and mobile.
.claude/skills/mengto-build-game-changelog/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 166% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 34% | 0% |
Build the changelog as a release system, not a decorative modal. Keep one authoritative ledger connected to production history, game-menu navigation, and release verification.
Inspect before editing:
Prefer repository, deployment, and live-runtime evidence over chat summaries or commit-message guesses. Do not invent missing release dates, source revisions, features, or version mappings.
Treat one successful player-visible production deployment as one changelog version.
Retain an established version scheme. For a new pre-1.0 game without one, default to 0.9.0 and increment the patch once per production release: 0.9.1, 0.9.2, and so on. Do not jump to a new minor or major version without an explicit product milestone.
Work oldest-to-newest when reconstructing history, then store entries newest-first for rendering.
0.9.0 when adopting the default scheme.If evidence is incomplete, mark provenance unknown in internal metadata or stop for clarification. Never fabricate a neat history.
Keep changelog content in typed or schema-validated structured data. Give each entry:
Derive the displayed current version from the newest ledger entry. Do not maintain a second hand-written version constant.
Store an explicit mapping between game versions and deployment versions. Only use a formula such as patch = deploymentOrdinal - 1 when the project guarantees it and tests it.
Avoid self-referential source metadata. A commit cannot contain its own final hash. Either:
See references/reference-architecture.md when implementing the schema, state machine, or tests.
Make the changelog an explicit menu route or screen state. Use an overlay only when the game’s existing menu architecture uses overlays.
On open, focus the close control or panel heading. On close, restore focus to the trigger when practical. Never strand keyboard or controller focus inside an unmounted panel.
Fit the existing game UI rather than introducing a new visual language.
aria-expanded, aria-controls, a labelled panel, and dynamic open or close labels on web-based games.Prefer concise, specific notes such as “Restore vitality at checkpoints” over “Various fixes.” Avoid developer-only jargon unless players need it.
Use this order:
Do not create another changelog version merely for filling in the previous release’s provenance unless that metadata update itself is deployed to players. Bundle bookkeeping with the next real release when possible.
Require automated checks for:
Finish with production proof when the task includes a release. Test the live build, not only a local preview.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 22,750 | 2,337 | -90% | 1 | 1 | 0% | 4,970 | 1,826 | -63% | 0 | 0 | — |
case-04 | fail→pass | 10,685 | 8,862 | -17% | 1 | 1 | 0% | 2,022 | 3,205 | +59% | 0 | 0 | — |
case-05 | pass→pass | 8,396 | 5,853 | -30% | 1 | 1 | 0% | 1,405 | 2,701 | +92% | 0 | 0 | — |
case-01 | fail→fail | 18,962 | 4,386 | -77% | 1 | 1 | 0% | 4,375 | 1,988 | -55% | 0 | 0 | — |
case-02 | fail→fail | 15,602 | 5,198 | -67% | 1 | 1 | 0% | 3,543 | 1,933 | -45% | 0 | 0 | — |
case-06 | pass→pass | 8,066 | 5,175 | -36% | 1 | 1 | 0% | 1,519 | 2,660 | +75% | 0 | 0 | — |
case-07 | pass→pass | 10,112 | 8,467 | -16% | 1 | 1 | 0% | 1,725 | 3,183 | +85% | 0 | 0 | — |
case-08 | fail→pass | 9,809 | 4,219 | -57% | 1 | 1 | 0% | 1,724 | 2,488 | +44% | 0 | 0 | — |
case-09 | pass→pass | 14,065 | 12,250 | -13% | 1 | 1 | 0% | 2,572 | 3,813 | +48% | 0 | 0 | — |
case-10 | pass→pass | 7,465 | 4,044 | -46% | 1 | 1 | 0% | 1,269 | 2,417 | +90% | 0 | 0 | — |
case-17 | pass→pass | 11,070 | 9,151 | -17% | 1 | 1 | 0% | 1,908 | 3,152 | +65% | 0 | 0 | — |
case-11 | fail→pass | 5,210 | 4,345 | -17% | 1 | 1 | 0% | 981 | 2,611 | +166% | 0 | 0 | — |
case-12 | pass→pass | 9,673 | 6,869 | -29% | 1 | 1 | 0% | 1,896 | 3,007 | +59% | 0 | 0 | — |
case-13 | pass→pass | 10,693 | 4,875 | -54% | 1 | 1 | 0% | 1,835 | 2,562 | +40% | 0 | 0 | — |
case-14 | pass→pass | 12,538 | 9,181 | -27% | 1 | 1 | 0% | 2,123 | 3,472 | +64% | 0 | 0 | — |
case-15 | pass→pass | 13,541 | 12,916 | -5% | 1 | 1 | 0% | 2,713 | 4,452 | +64% | 0 | 0 | — |
case-16 | fail→pass | 10,303 | 7,055 | -32% | 1 | 1 | 0% | 1,756 | 2,807 | +60% | 0 | 0 | — |
case-18 | pass→pass | 8,276 | 4,230 | -49% | 1 | 1 | 0% | 1,448 | 2,398 | +66% | 0 | 0 | — |
case-19 | fail→pass | 11,248 | 6,248 | -44% | 1 | 1 | 0% | 2,189 | 2,933 | +34% | 0 | 0 | — |
case-20 | pass→pass | 4,212 | 3,411 | -19% | 1 | 1 | 0% | 881 | 2,426 | +175% | 0 | 0 | — |
case-21 | pass→pass | 2,111 | 2,395 | +13% | 1 | 1 | 0% | 451 | 2,147 | +376% | 0 | 0 | — |
case-22 | pass→pass | 3,094 | 4,244 | +37% | 1 | 1 | 0% | 528 | 2,417 | +358% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.