Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Prepare, verify, and publish Sutando engine releases. Use when asked to prepare a release proposal, determine a version bump, audit release documentation, validate release gates, create a release PR, tag a confirmed release, or publish GitHub release notes. Preparation is safe and non-publishing; tagging and publishing require explicit owner confirmation.
.claude/skills/sonichi-release/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 31% | 0% |
Use docs/release-process.md as policy. Do not duplicate or reinterpret it here. This skill turns that policy into an evidence-backed workflow.
gates, and write a proposal. This is the default.
confirmation in the current conversation.
Never infer publish permission from “prepare,” “get ready,” a green CI run, or an approved release PR.
docs/release-process.md completely.bash REPO="$(git rev-parse --show-toplevel)" WORKSPACE="$(bash "$REPO/scripts/sutando-config.sh" workspace)"
to the candidate; do not rely only on PR titles.
bash python3 skills/release/scripts/docs_audit.py
identify the canonical document in docs/catalog.json. Update that document, its last_verified date, README navigation when appropriate, migration or upgrade guidance, and CHANGELOG.md. Do not copy canonical text into a second document.
Do not mix product fixes into it; list unresolved fixes as blockers. It records the entry after the tag (see Publish step 2) and is not a precondition for publishing.
docs/release-process.md, including thefresh-clone health check, headline-feature smoke, migration idempotency, and prior-release-to-candidate upgrade smoke. Record exact commands, candidate SHA, exit status, and observable evidence. Never invent or summarize an unexecuted gate as passing.
$WORKSPACE/notes/release-proposals/proposed-vX.Y.Z.md with:
publish: pending owner confirmation.Proceed only when the owner explicitly confirms the exact version and candidate SHA in the current conversation.
docs/release-process.md.green, the worktree is clean, the tag does not exist, and every blocker is resolved.
Do not gate the tag on the release PR being merged. This repo tags the candidate commit and records the CHANGELOG entry afterwards — v0.7.0 through v0.10.0 were each tagged on a tree containing no entry for themselves, and v0.10.0's entry landed two days later in #2826. Holding the tag for the changelog PR blocks the release on review latency for no benefit.
the signing with their configured key; never impersonate an owner signature.
external writes requiring explicit owner confirmation.
CHANGELOG.md authoritative; generated GitHub notes are a completenessaid, not the release narrative.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→pass | 7,654 | 1,788 | -77% | 1 | 1 | 0% | 1,089 | 1,283 | +18% | 0 | 0 | — |
case-14 | pass→pass | 29,189 | 2,476 | -92% | 1 | 1 | 0% | 1,402 | 1,390 | -1% | 0 | 0 | — |
case-01 | fail→fail | 24,225 | 7,580 | -69% | 1 | 1 | 0% | 3,096 | 1,368 | -56% | 0 | 0 | — |
case-02 | fail→fail | 5,925 | 6,795 | +15% | 1 | 1 | 0% | 1,033 | 1,366 | +32% | 0 | 0 | — |
case-03 | fail→fail | 13,179 | 7,031 | -47% | 1 | 1 | 0% | 1,914 | 1,556 | -19% | 0 | 0 | — |
case-04 | fail→fail | 4,422 | 7,768 | +76% | 1 | 1 | 0% | 709 | 1,488 | +110% | 0 | 0 | — |
case-06 | fail→pass | 14,409 | 3,740 | -74% | 1 | 1 | 0% | 1,905 | 1,605 | -16% | 0 | 0 | — |
case-07 | pass→pass | 10,210 | 3,337 | -67% | 1 | 1 | 0% | 1,480 | 1,407 | -5% | 0 | 0 | — |
case-08 | pass→pass | 10,020 | 5,016 | -50% | 1 | 1 | 0% | 1,530 | 1,792 | +17% | 0 | 0 | — |
case-09 | fail→pass | 11,236 | 5,446 | -52% | 1 | 1 | 0% | 1,723 | 1,822 | +6% | 0 | 0 | — |
case-10 | fail→pass | 12,970 | 5,690 | -56% | 1 | 1 | 0% | 1,932 | 1,761 | -9% | 0 | 0 | — |
case-11 | pass→pass | 12,602 | 2,061 | -84% | 1 | 1 | 0% | 1,734 | 1,234 | -29% | 0 | 0 | — |
case-12 | fail→pass | 8,218 | 5,434 | -34% | 1 | 1 | 0% | 1,404 | 1,841 | +31% | 0 | 0 | — |
case-13 | pass→pass | 14,972 | 5,418 | -64% | 1 | 1 | 0% | 2,183 | 1,686 | -23% | 0 | 0 | — |
case-15 | fail→fail | 8,952 | 9,586 | +7% | 1 | 1 | 0% | 1,273 | 2,440 | +92% | 0 | 0 | — |
case-16 | fail→fail | 9,925 | 3,308 | -67% | 1 | 1 | 0% | 1,578 | 1,481 | -6% | 0 | 0 | — |
case-17 | fail→pass | 3,342 | 1,622 | -51% | 1 | 1 | 0% | 461 | 1,198 | +160% | 0 | 0 | — |
case-18 | fail→pass | 14,452 | 1,386 | -90% | 1 | 1 | 0% | 959 | 1,199 | +25% | 0 | 0 | — |
case-19 | pass→pass | 4,835 | 1,799 | -63% | 1 | 1 | 0% | 652 | 1,274 | +95% | 0 | 0 | — |
case-20 | pass→pass | 3,814 | 4,549 | +19% | 1 | 1 | 0% | 679 | 1,738 | +156% | 0 | 0 | — |
case-21 | pass→pass | 6,480 | 3,602 | -44% | 1 | 1 | 0% | 1,035 | 1,543 | +49% | 0 | 0 | — |
case-22 | pass→pass | 7,159 | 4,206 | -41% | 1 | 1 | 0% | 1,311 | 1,750 | +33% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.