Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Final repo-harness closeout workflow. Runs review/check gates, commits finished contract worktrees, pushes codex branches, and creates GitHub PRs by default.
.claude/skills/ancienttwo-repo-harness-ship/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -36% | 0% |
Use this command when implementation is complete and the user wants the harness to close out worktrees and create reviewable PRs.
git status --short --branch -uall and git worktree list --porcelain.repo-harness run ship-worktrees; it validates review/check evidence, runs repo-harness run contract-worktree finish --no-merge, pushes codex/<slug>, and creates a draft PR with gh pr create --base main --head codex/<slug>.--slug <slug> so the script creates codex/<slug>-main-closeout and opens a PR instead of committing to main.repo-harness run ship-worktrees --local-merge; this preserves the older finish -> fast-forward main -> cleanup path.main contains the branch, run repo-harness run ship-worktrees --cleanup-merged to remove only proven-merged local worktrees and branches. If the branch is merged but the linked worktree is dirty, pick/apply/commit useful changes first; use --discard-scaffold-only only when the dirty paths are generated plan/contract/review/notes scaffold.verify-sprint evidence are present and passing.--discard-scaffold-only only for generated scaffold files.main, merge PRs, publish releases, or tag versions.git reset --hard, git clean, or automatic stash._ops tgz archives as a successful closeout path; merged dirty worktrees must be committed/picked/applied or explicitly discarded as scaffold-only./check, external acceptance, or verify-sprint evidence is missing or failing.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | pass→pass | 3,263 | 1,659 | -49% | 1 | 1 | 0% | 576 | 792 | +38% | 0 | 0 | — |
case-01 | fail→fail | 3,506 | 4,352 | +24% | 1 | 1 | 0% | 394 | 781 | +98% | 0 | 0 | — |
case-02 | fail→fail | 3,337 | 2,352 | -30% | 1 | 1 | 0% | 342 | 854 | +150% | 0 | 0 | — |
case-03 | fail→fail | 9,104 | 5,258 | -42% | 1 | 1 | 0% | 1,681 | 846 | -50% | 0 | 0 | — |
case-04 | fail→pass | 8,314 | 5,862 | -29% | 1 | 1 | 0% | 1,266 | 1,194 | -6% | 0 | 0 | — |
case-05 | fail→pass | 10,333 | 3,171 | -69% | 1 | 1 | 0% | 1,643 | 1,165 | -29% | 0 | 0 | — |
case-06 | pass→pass | 10,464 | 4,283 | -59% | 1 | 1 | 0% | 1,814 | 1,400 | -23% | 0 | 0 | — |
case-07 | fail→pass | 8,977 | 5,099 | -43% | 1 | 1 | 0% | 1,646 | 1,502 | -9% | 0 | 0 | — |
case-08 | fail→pass | 16,281 | 2,567 | -84% | 1 | 1 | 0% | 1,256 | 1,069 | -15% | 0 | 0 | — |
case-09 | pass→pass | 6,619 | 4,164 | -37% | 1 | 1 | 0% | 1,397 | 1,327 | -5% | 0 | 0 | — |
case-10 | fail→pass | 11,662 | 6,115 | -48% | 1 | 1 | 0% | 1,797 | 1,159 | -36% | 0 | 0 | — |
case-11 | pass→pass | 11,310 | 3,817 | -66% | 1 | 1 | 0% | 1,817 | 1,270 | -30% | 0 | 0 | — |
case-12 | pass→pass | 6,003 | 3,874 | -35% | 1 | 1 | 0% | 1,080 | 1,265 | +17% | 0 | 0 | — |
case-13 | fail→pass | 6,509 | 4,460 | -31% | 1 | 1 | 0% | 1,266 | 1,436 | +13% | 0 | 0 | — |
case-14 | fail→pass | 12,181 | 4,154 | -66% | 1 | 1 | 0% | 2,131 | 1,354 | -36% | 0 | 0 | — |
case-15 | pass→pass | 2,632 | 3,003 | +14% | 1 | 1 | 0% | 472 | 1,190 | +152% | 0 | 0 | — |
case-16 | pass→fail | 6,680 | 5,469 | -18% | 1 | 1 | 0% | 1,358 | 903 | -34% | 0 | 0 | — |
case-17 | pass→pass | 10,731 | 5,639 | -47% | 1 | 1 | 0% | 1,985 | 1,701 | -14% | 0 | 0 | — |
case-19 | pass→pass | 4,514 | 1,922 | -57% | 1 | 1 | 0% | 824 | 891 | +8% | 0 | 0 | — |
case-20 | fail→pass | 10,035 | 1,810 | -82% | 1 | 1 | 0% | 1,782 | 878 | -51% | 0 | 0 | — |
case-21 | fail→pass | 7,504 | 2,760 | -63% | 1 | 1 | 0% | 1,226 | 1,052 | -14% | 0 | 0 | — |
case-22 | pass→pass | 4,077 | 4,152 | +2% | 1 | 1 | 0% | 723 | 1,332 | +84% | 0 | 0 | — |
case-23 | fail→pass | 10,934 | 2,742 | -75% | 1 | 1 | 0% | 1,901 | 1,052 | -45% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 20 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.