Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run the Harnss release workflow — review staged diff, bump version, commit, tag, push, and create a GitHub release. Use when releasing, bumping version, tagging, or creating a release. Argument: major, minor, or patch.
.claude/skills/opensource03-release/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 114% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -42% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 171% | 0% |
Run the full release pipeline. Bump type is passed as $ARGUMENTS (major, minor, or patch).
Run these commands and read every line of output:
bashgit status git diff --cached --stat git diff --cached
If the diff is too large to read in one shot, read it in chunks (e.g., per-directory or line ranges). You must read the entire diff before proceeding.
Review for:
test-*.ts, scratch.*, *.tmp, random-*.md).env, credentials, large binaries)If you find any:
If there are no staged changes but unstaged changes exist, ask the user if they want to stage anything first. If the working tree is completely clean (nothing to release), tell the user and stop.
package.json (the "version" field)git tag --sort=-v:refname | head -1MAJOR.MINOR.PATCH$ARGUMENTS:major → (MAJOR+1).0.0minor → MAJOR.(MINOR+1).0patch → MAJOR.MINOR.(PATCH+1)$ARGUMENTS is empty or invalid, ask the user which bump type they wantbashnpm view @anthropic-ai/claude-agent-sdk version
Compare with the version in package.json under dependencies["@anthropic-ai/claude-agent-sdk"] (strip the ^ prefix for comparison). If a newer version exists, update the dependency version (keep the ^ prefix) and tell the user.
package.json to set the new version numbergit add package.jsonbash pnpm install git add package.json pnpm-lock.yaml
Choose the commit message based on what's staged:
feat: short summary (2-4 key themes)
- Change description 1
- Change description 2
- ...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>Use fix: instead of feat: if all changes are bug fixes.
chore: bump version to X.Y.Z
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>If the SDK was also updated:
chore: bump version to X.Y.Z and update claude-agent-sdk to A.B.C
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>bashgit commit -m "$(cat <<'EOF' <message here> Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> EOF )"
bashgit tag vX.Y.Z HEAD git push origin master && git push origin vX.Y.Z
If push fails, report the error and stop.
Get the previous release tag:
bashgit tag --sort=-v:refname | head -2 | tail -1
Read the full diff and commit log since the previous release:
bashgit log v{prev}...HEAD --oneline git diff v{prev}...HEAD --stat git diff v{prev}...HEAD
Read ALL of this output. For the full diff, read it in chunks if needed — every line matters for writing accurate release notes.
Load the template from references/release-notes-template.md and follow its format exactly.
bashgh release create vX.Y.Z --title "vX.Y.Z — Short Descriptive Phrase" --notes "$(cat <<'EOF' <release notes> EOF )"
The title uses an em dash (—), not a hyphen.
Output the release URL when done so the user can verify.
Co-Authored-By trailer is mandatory on every commit.https://github.com/OpenSource03/harnssmasterhttps://github.com/OpenSource03/harnss/compare/v{prev}...v{current}pnpm (never use npm or yarn for installs)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,822 | 4,005 | -17% | 1 | 1 | 0% | 214 | 1,383 | +546% | 0 | 0 | — |
case-02 | fail→fail | 11,029 | 4,018 | -64% | 1 | 1 | 0% | 1,449 | 1,385 | -4% | 0 | 0 | — |
case-03 | fail→fail | 5,073 | 4,411 | -13% | 1 | 1 | 0% | 258 | 1,392 | +440% | 0 | 0 | — |
case-04 | pass→fail | 7,519 | 8,129 | +8% | 1 | 1 | 0% | 1,465 | 1,568 | +7% | 0 | 0 | — |
case-05 | pass→pass | 15,556 | 11,361 | -27% | 1 | 1 | 0% | 2,718 | 3,214 | +18% | 0 | 0 | — |
case-06 | pass→pass | 10,703 | 5,164 | -52% | 1 | 1 | 0% | 1,878 | 2,104 | +12% | 0 | 0 | — |
case-07 | pass→fail | 7,766 | 2,575 | -67% | 1 | 1 | 0% | 1,267 | 1,600 | +26% | 0 | 0 | — |
case-08 | fail→pass | 3,825 | 1,332 | -65% | 1 | 1 | 0% | 657 | 1,404 | +114% | 0 | 0 | — |
case-09 | pass→fail | 10,137 | 1,083 | -89% | 1 | 1 | 0% | 1,678 | 1,379 | -18% | 0 | 0 | — |
case-10 | fail→pass | 6,683 | 3,774 | -44% | 1 | 1 | 0% | 1,088 | 1,450 | +33% | 0 | 0 | — |
case-11 | fail→pass | 15,935 | 2,053 | -87% | 1 | 1 | 0% | 2,690 | 1,573 | -42% | 0 | 0 | — |
case-12 | pass→pass | 5,650 | 1,986 | -65% | 1 | 1 | 0% | 928 | 1,466 | +58% | 0 | 0 | — |
case-13 | fail→pass | 13,849 | 5,112 | -63% | 1 | 1 | 0% | 2,383 | 2,096 | -12% | 0 | 0 | — |
case-14 | fail→pass | 3,237 | 1,726 | -47% | 1 | 1 | 0% | 526 | 1,424 | +171% | 0 | 0 | — |
case-15 | fail→pass | 9,311 | 2,832 | -70% | 1 | 1 | 0% | 1,613 | 1,712 | +6% | 0 | 0 | — |
case-16 | fail→pass | 10,097 | 1,796 | -82% | 1 | 1 | 0% | 1,653 | 1,504 | -9% | 0 | 0 | — |
case-17 | fail→pass | 3,495 | 1,392 | -60% | 1 | 1 | 0% | 464 | 1,358 | +193% | 0 | 0 | — |
case-18 | fail→pass | 12,420 | 1,638 | -87% | 1 | 1 | 0% | 1,979 | 1,428 | -28% | 0 | 0 | — |
case-19 | fail→pass | 3,670 | 1,264 | -66% | 1 | 1 | 0% | 648 | 1,381 | +113% | 0 | 0 | — |
case-20 | pass→pass | 8,077 | 2,901 | -64% | 1 | 1 | 0% | 1,456 | 1,631 | +12% | 0 | 0 | — |
case-21 | fail→fail | 10,198 | 1,889 | -81% | 1 | 1 | 0% | 1,598 | 1,532 | -4% | 0 | 0 | — |
case-22 | pass→pass | 10,034 | 2,253 | -78% | 1 | 1 | 0% | 1,712 | 1,517 | -11% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 18 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.