Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use before claiming CodeWhale release work is done: run the full gate sweep and list the manual QA targets.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -5% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 70% | 0% |
Run this before claiming any CodeWhale release work is "done." A green automated gate sweep plus the three manual QA targets is the evidence bar. No sweep, no "done" — report exactly what was run and the result of each step.
merge-ready.
(e.g. <release-branch>), which is often local-only.
Run from the repo root, in order. Stop on the first failure and report it.
bash# 0. Confirm you are on the real release head, not a main-based assumption. git branch --show-current # expect e.g. <release-branch> git status --short # working tree should be clean # 1. Formatting + stray whitespace/conflict markers cargo fmt --all --check git diff --check # 2. Library/protocol/cli/flow/state tests, locked cargo test -p codewhale-config -p codewhale-protocol -p codewhale-cli \ -p codewhale-workflow -p codewhale-state --locked # 3. TUI test binaries, locked cargo test -p codewhale-tui --bins --locked # 4. Real-PTY release runtime QA (sealed HOME + loopback providers) cargo test -p codewhale-tui --test release_runtime_qa --locked -- --test-threads=1 # 5. TUI debug build, locked cargo build -p codewhale-tui --locked # 6. Release build for the shipped binaries, locked cargo build --release --locked -p codewhale-cli -p codewhale-tui # 7. Version-drift gate (workspace ↔ npm ↔ Cargo.lock ↔ changelog ↔ README) ./scripts/release/check-versions.sh # 8. Binary smoke ./target/release/codewhale --version
If you are validating a PR for landing, also test mergeability against the actual release head, never the main-based clean flag:
bashgit merge-tree $(git merge-base <release-branch> <pr-head>) <release-branch> <pr-head>
A PR that is clean against main can still conflict with the release branch.
Unit/build gates do not cover the live TUI. Exercise all three and record what you saw:
The repeatable local baseline is release_runtime_qa: it boots real TUI processes in pseudo-terminals with sealed homes and loopback mock providers, then asserts each scenario below. Run it even when doing a separate hands-on visual pass; the test leaves no provider traffic or credentials behind.
typing, render, cancel, and the sidebar stay live throughout, and that Esc cancels mid-fanout (prompt interrupt, not a wedged ~24s burst or freeze). For the Windows Terminal retest path from #3289, start in plan mode, add follow-up input to the plan, press Esc, switch to yolo/accept flow, trigger at least two auto/Fleet worker spawns, and keep typing/cancel/mode-switch checks live for several minutes. Attach logs if the freeze reproduces.
distinct provider/model routes. Confirm zero cross-terminal contamination and no provider+model mismatch — each terminal honors its own route.
queues a typed follow-up, the preview advertises Enter send now, and an empty Enter promotes the oldest queued follow-up. Confirm Ctrl+Enter steers typed text directly, Shift+Enter inserts a newline, and Ctrl+G/Ctrl+S only stash drafts.
Report a checklist: each command, pass/fail, and the salient output line (test counts, the --version string, check-versions.sh verdict). For manual QA, state what you actually observed per target, citing the issue number. If a step was skipped or could not be run (e.g. no display for TUI QA), say so explicitly — do not imply coverage you do not have.
Assertions without command output are not acceptable.
git merge-tree against the real head.
route-mismatch, and steering regressions live in the runtime, not the gates.
any PR or issue without Hunter's explicit approval. A green sweep is readiness evidence, not permission.
comments, and checks.
keeps the original author, otherwise add Co-authored-by: Name <email> and Harvested-from: PR #N by @handle so the auto-close-at-main workflow credits the contributor.
dry-run/advisory unless Hunter approves enforcement.
Other measured skills in the registry, with their headline benchmark lift.