Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Intent-anchored ship pipeline — capture intent, then run challenge → validate → verify → code-review → docs-gap → create-pr, then watch CI and triage review comments, as gated stages that pause only on findings needing a human decision. A nudgebee-native take on the "no-mistakes" pipeline.
.claude/skills/nudgebee-ship/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 71% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 171% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 184% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 213% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 206% | 0% |
Orchestrate the existing nudgebee skills into one intent-anchored pipeline that takes a change from committed-on-a-branch to PR-ready. This is a thin driver — it does not re-implement service detection, validation, review, or PR formatting. It invokes the skills that already own those (/challenge, /validate, /code-review, /create-pr, /pr-comments) and adds the two ideas worth borrowing from the no-mistakes tool: an explicit intent that every stage is judged against, and a gate taxonomy so the agent only stops the user for decisions that are genuinely theirs.
This is opt-in and skill-only — it blocks nothing on its own. There is no git proxy and no push hook; a manual git push still works. The pipeline's authority ends at "PR opened."
Optional argument: $ARGUMENTS — the target base branch (defaults to main).
bashgit branch --show-current git status --short git log --oneline -5
main/test/prod, stop — ship runs from a feature branch. Tell the user to branch first./commit). Ship reviews committed work; it does not commit for the user.Intent is the anchor the whole pipeline is judged against — it is not a diff summary. A good intent states: what the user set out to accomplish, the key decisions/tradeoffs made, and anything deliberately left out of scope.
user and ask them to confirm or correct it before proceeding. Do not silently invent it — a wrong intent poisons every downstream stage.
Hold the confirmed intent in context; pass it into the review stages below.
Every finding surfaced by any stage below is classified into exactly one bucket:
import your change created, an obviously-correct one-line correction). The agent applies it, notes it, and continues. No pause.
cross-service change, a design tradeoff, scope creep, or a finding that contradicts the stated intent. Stop and escalate. Never auto-approve an ask-user finding.
When unsure whether a finding is auto-fix or ask-user, treat it as ask-user.
Invoke the challenge skill against the intent + diff. It emits PROCEED / REVISE / REDESIGN.
PROCEED → continue.REVISE / REDESIGN → this is an ask-user gate. Surface the objections and stop;do not paper over a structural objection by proceeding.
Skip this stage only for changes the repo's AI principles say skip /challenge (typos, formatting, 1-line fixes, docs-only) — state that you skipped it and why.
Invoke the validate skill (no argument — let it detect affected services from the diff).
fix-lint, re-runvalidate once, continue if green.
verbatim and stop. Do not "fix" a failing test by weakening it.
Green tests prove nothing broke; they do not prove the change does what the intent says. If the diff has runtime surface (product code, not docs/tests/config-only), exercise the affected flow end-to-end and observe its behavior against the Step 1 intent: run the service locally (or the relevant e2e test) and drive the changed path with a real request/event, capturing what was exercised and what was observed as evidence. Per CLAUDE.md → Definition of Done, behavior must be observed, not assumed.
observed) for the final report and PR body.
non-trivial → ask-user gate. Do not proceed on unverified behavior.
evidence for a change you did not actually drive.
Invoke the code-review skill on the diff. Route each finding through the Step 2 gate model: apply auto-fix findings, collect ask-user findings, record no-op findings. If any ask-user findings exist, stop and present them before opening a PR.
Cross-check every finding against the intent from Step 1 — a change that review flags as unexplained but the intent justifies is a no-op; a change the intent does not justify is scope creep and an ask-user finding.
A code change often obsoletes or requires a doc. Map the changed paths to their likely doc targets and check each for a gap:
| Changed | Likely doc target | |---|---| | app/src/lib/actions.yaml (new/renamed action) | naming vs docs/rpc-action-naming.md | | api-server/migrations/** | api-server/migrations/README.md | | A service's public behavior / setup | that service's CLAUDE.md | | User-facing feature or install flow | nudgebee-docs/ (if present in the workspace) | | Shared type / API contract / cross-service change | docs/architecture-decisions.md entry |
Route findings through the Step 2 gate model:
→ auto-fix: update the doc, include it in the change.
this the right doc home?) → ask-user.
principles are surgical-changes and simplicity-first.
Only reached when the prior steps left no open ask-user gate. Invoke the create-pr skill with the target base branch ($ARGUMENTS, default main). It already handles the issue-link requirement, the template, self-review, and validation — do not duplicate that here.
Seed the PR's intent/summary from the Step 1 intent so the human reviewer starts from the same anchor the pipeline used — the Step 1 one-liner is usually the best available draft of the PR lead, because it was written before the implementation detail piled up.
The verify-stage evidence (Step 5) is author-side proof and belongs inside the PR body's <details> fold. What goes above the fold is how a reviewer or QA confirms the change — see docs/writing-for-readers.md, which create-pr enforces. If the pipeline never produced reader-side verification steps, say so in one honest line rather than dressing up the commands you ran as a test plan.
Capture the PR number/URL from create-pr for the next step.
Once the PR exists, watch it settle instead of declaring victory at "PR opened."
CI status — watch the checks for this PR as a background task (never a foreground sleep/poll loop; keep working while it runs):
bashgh pr checks <pr-number> --watch # run in the background; read the final matrix when it reports gh pr checks <pr-number> # final state, one line per check
Route the outcome through the Step 2 gate model:
slipped, a flaky-looking build step) → auto-fix: reproduce with /validate, fix, push the follow-up commit, and re-watch once. Do not loop indefinitely — one auto-fix attempt, then escalate.
gate. Report the failing check name and its log tail verbatim, and stop.
Do not merge — merging stays with the human reviewer and the repo's main → test → prod automation.
Review comments — fetch any human/bot review feedback already on the PR:
bashgh pr view <pr-number> --json reviews,comments,reviewDecision
If there are actionable review comments, invoke the pr-comments skill to enumerate and triage them. Apply auto-fix-class comments (obvious, unambiguous corrections) and push a follow-up commit; collect judgment-class comments as ask-user items to present to the user. If the PR is fresh and has no reviews yet, note that and move on — don't wait for a reviewer.
Emit a compact summary:
(counts per gate bucket).
ask-user gate that stopped the run (if it stopped early).ask-user gate.stays with the human reviewer and the repo's main → test → prod automation.
formatting here, stop — that logic lives in validate / create-pr and must stay there.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,049 | 6,351 | -58% | 1 | 1 | 0% | 740 | 2,824 | +282% | 0 | 0 | — |
case-02 | fail→fail | 20,704 | 7,590 | -63% | 1 | 1 | 0% | 538 | 2,902 | +439% | 0 | 0 | — |
case-03 | fail→fail | 15,877 | 10,861 | -32% | 1 | 1 | 0% | 306 | 2,804 | +816% | 0 | 0 | — |
case-04 | pass→fail | 14,103 | 6,949 | -51% | 1 | 1 | 0% | 1,462 | 2,901 | +98% | 0 | 0 | — |
case-05 | fail→fail | 8,062 | 6,851 | -15% | 1 | 1 | 0% | 968 | 2,866 | +196% | 0 | 0 | — |
case-06 | fail→fail | 8,262 | 11,340 | +37% | 1 | 1 | 0% | 1,123 | 2,775 | +147% | 0 | 0 | — |
case-07 | fail→pass | 13,989 | 4,130 | -70% | 1 | 1 | 0% | 1,903 | 3,257 | +71% | 0 | 0 | — |
case-08 | fail→pass | 10,615 | 21,902 | +106% | 1 | 1 | 0% | 1,295 | 3,504 | +171% | 0 | 0 | — |
case-09 | pass→pass | 13,698 | 5,267 | -62% | 1 | 1 | 0% | 1,852 | 3,349 | +81% | 0 | 0 | — |
case-10 | pass→fail | 9,156 | 4,819 | -47% | 1 | 1 | 0% | 1,273 | 3,287 | +158% | 0 | 0 | — |
case-11 | pass→pass | 11,052 | 3,870 | -65% | 1 | 1 | 0% | 1,664 | 3,160 | +90% | 0 | 0 | — |
case-12 | fail→fail | 28,875 | 6,197 | -79% | 1 | 1 | 0% | 1,736 | 3,414 | +97% | 0 | 0 | — |
case-13 | pass→pass | 20,609 | 4,366 | -79% | 1 | 1 | 0% | 1,845 | 3,132 | +70% | 0 | 0 | — |
case-14 | fail→pass | 9,256 | 6,010 | -35% | 1 | 1 | 0% | 1,244 | 3,531 | +184% | 0 | 0 | — |
case-15 | fail→pass | 8,275 | 3,260 | -61% | 1 | 1 | 0% | 990 | 3,097 | +213% | 0 | 0 | — |
case-16 | fail→pass | 8,313 | 5,474 | -34% | 1 | 1 | 0% | 1,134 | 3,474 | +206% | 0 | 0 | — |
case-17 | fail→fail | 10,456 | 3,158 | -70% | 1 | 1 | 0% | 1,576 | 2,994 | +90% | 0 | 0 | — |
case-18 | fail→pass | 15,292 | 4,256 | -72% | 1 | 1 | 0% | 2,298 | 3,207 | +40% | 0 | 0 | — |
case-19 | fail→pass | 11,127 | 5,732 | -48% | 1 | 1 | 0% | 1,504 | 3,232 | +115% | 0 | 0 | — |
case-20 | pass→pass | 10,288 | 4,383 | -57% | 1 | 1 | 0% | 1,383 | 3,262 | +136% | 0 | 0 | — |
case-21 | fail→pass | 15,066 | 11,065 | -27% | 1 | 1 | 0% | 2,408 | 3,332 | +38% | 0 | 0 | — |
case-22 | fail→fail | 15,769 | 128,356 | +714% | 1 | 1 | 0% | 1,064 | 2,941 | +176% | 0 | 0 | — |
case-23 | fail→fail | 5,452 | 9,295 | +70% | 1 | 1 | 0% | 713 | 3,120 | +338% | 0 | 0 | — |
case-24 | fail→fail | 62,417 | 22,517 | -64% | 1 | 1 | 0% | 7,584 | 3,074 | -59% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 15 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +25 percentage points is the difference between those two pass rates over the 15 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.