Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create a pull request with proper formatting, validation, and conventions for this monorepo
.claude/skills/nudgebee-create-pull-request/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 281% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 347% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 379% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 376% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 463% | 0% |
Create a pull request for the current branch, then watch it through to a green, review-clean state — polling CI and review threads, fixing what can be auto-fixed, and surfacing what only a human can resolve (Step 11). Optional argument: $ARGUMENTS (target base branch, defaults to main).
Before writing any prose — the PR body, review replies, the final report — read docs/writing-for-readers.md. It is the binding style contract for anything this skill posts to GitHub. A correct PR that nobody finishes reading has failed.
Run these commands to understand the current state:
bash# Current branch and tracking info git branch --show-current git status # Commits on this branch not in base git log main..HEAD --oneline # Full diff against base git diff main...HEAD --stat git diff main...HEAD
If $ARGUMENTS specifies a different base branch (e.g., test or prod), use that instead of main.
This step applies ONLY when the base branch is main. Skip this step entirely for PRs targeting test or prod (cherry-picks / promotions).
Goal: traceability. The PR must link back to the ticket that motivated this work — the sprint ticket or existing issue — NOT a freshly minted issue. Creating a new issue per PR destroys traceability and litters the tracker with unassigned orphans. Never create a new issue in this step.
Find the original ticket, in this order — stop at the first hit:
$ARGUMENTS, the branch name (fix/nb-1234-description), commit messages, or mentioned anywhere in the conversation (including the ticket the session started from, e.g. via /my-tickets). Verify it exists:bash gh issue view <issue_number> --json number,title,state 2>&1
bash gh issue list --assignee "@me" --state open --limit 30 --json number,title
bash gh issue list --state all --search "<keywords from branch/commits>" --limit 10 --json number,title,state
If candidates are found, present them and ask the user which one to link (or none). A closed issue is still a valid link target for follow-up work.
If no related ticket exists, stop and ask the user how to proceed. Do NOT run gh issue create and do NOT invoke /create-issue on your own — only the user can decide a new ticket is warranted. If they confirm one, create it via /create-issue (which assigns it and adds it to the sprint board) and reference the motivating context in its body.
Once an issue number is confirmed, store it for the PR body:
Fixes #<number> if this PR fully resolves the ticket (auto-closes it on merge).Part of #<number> if this PR is one piece of a larger ticket that must stay open.Map changed files to services using this table:
| Path prefix | Service | Type | Validation | |---|---|---|---| | api-server/services/ | api-server | Go | make validate | | ticket-server/ | ticket-server | Go | make validate | | collector-server/cloud-collector/ | cloud-collector | Go | make validate | | collector-server/k8s-collector/relay-server/ | relay-server | Go | make validate | | collector-server/k8s-collector/app/ | k8s-collector-app | Python | make lint && make test | | llm/code-analysis/ | code-analysis | Go | make check | | llm/llm-server/ | llm-server | Go | make validate | | llm/rag-server/ | rag-server | Python | make lint && make test | | llm/benchmark/ | benchmark | Python | poetry run pytest | | ml-k8s-server/ | ml-k8s-server | Python | make lint && make test | | auto-pilot/ | auto-pilot | Python | poetry run black --check . && poetry run flake8 . | | auto-pilot/sidecar/ | auto-pilot-sidecar | Python | poetry run black --check . && poetry run flake8 . | | notifications-server/ | notifications-server | Python | poetry run black --check . && poetry run flake8 . | | app/ | frontend | TypeScript | npm run lint2 | | deploy/ | infrastructure | — | Manual review |
For each affected service, run its validation command. Report results to the user. If validation fails, ask the user whether to fix the issues or proceed anyway.
Mandatory. Before pushing, read the diff and run an AI first-pass review. The goal is to catch issues before a human reviewer ever sees them, and to surface residual risks explicitly rather than hoping the reviewer finds them.
Run:
bashgit diff {base_branch}...HEAD
Read the diff in full and evaluate against these dimensions:
CLAUDE.md → AI Coding Principles, surgical changes: every changed line traces directly to the request).slog + testify, Python black 120 + flake8 + mypy, TypeScript oxlint + prettier, commit scope correctness.Categorize each finding into one of three buckets:
git commit -am "chore: fix issues found during self-review"). This commit will be squashed in Step 6. Do not push a PR with known issues that you could have fixed. A clean working tree is required for the rebase in Step 5 — uncommitted changes will cause it to fail.If the PR touches shared contracts, DB schema, cross-service behavior, or any architectural decision, also run the logic from /challenge against the diff itself: what are the three strongest reasons this diff is wrong? Include the surviving counterarguments in the Risks section of the PR body. Skip this sub-step for typo / docs / 1-line fixes.
MANDATORY: The branch MUST be rebased on the latest base branch before creating or updating a PR. This ensures a clean, linear history.
bash# Fetch latest from remote git fetch origin {base_branch} # Rebase onto the latest base branch git rebase origin/{base_branch}
If there are conflicts, resolve them and continue the rebase (git rebase --continue). If conflicts are too complex, inform the user and ask how to proceed.
PRs MUST contain exactly one commit. Check the commit count:
bashgit log origin/{base_branch}..HEAD --oneline | wc -l
If there is more than one commit, squash them into a single commit:
bash# Squash all commits into one git reset --soft origin/{base_branch} git commit -m "<combined commit message covering all changes>"
The squashed commit message should summarize all changes coherently. Use the PR title as the first line, and list key changes as bullet points in the body.
bash# Push with force-with-lease (required after rebase/squash) git push --force-with-lease -u origin $(git branch --show-current) 2>&1 || true
Always use --force-with-lease since rebase and squash rewrite history.
Based on the commits and diff, generate:
Title format: type(scope): subject (per .github/semantic.yml)
Allowed types: | Type | Use when | |---|---| | feat | New feature or functionality | | fix | Bug fix | | docs | Documentation only | | style | Formatting, whitespace, no code change | | refactor | Code restructure, no behavior change | | perf | Performance improvement | | test | Adding or updating tests only | | chore | Maintenance, deps, config | | revert | Reverting a previous commit | | ci | CI/CD workflow changes | | infra | Infrastructure, Helm, K8s changes | | release | Release-related changes |
Allowed scopes (required): | Scope | Services / paths | |---|---| | ui | app/ (frontend) | | autopilot | auto-pilot/, auto-pilot/sidecar/ | | ml | ml-k8s-server/ | | llm | ml-k8s-server/, llm/code-analysis/, llm/llm-server/, llm/rag-server/, llm/benchmark/ | | workflow | workflow-server/ | | notifications | notifications-server/ | | tickets | ticket-server/ | | relay | collector-server/k8s-collector/relay-server/ | | collector | collector-server/cloud-collector/, collector-server/k8s-collector/app/ | | deps | Dependency updates | | NB-xxx | Ticket number — use for api-server/services/, api-server/migrations/, deploy/, .github/, or any cross-service change |
Examples: fix(ui): handle null state in settings, feat(NB-1234): add Azure onboarding flow
Semantic type → PR "Type of change" mapping: | Semantic type | PR checkbox | |---|---| | feat | New feature | | fix | Bug fix | | docs | Documentation | | style | Chore | | refactor | Refactor | | perf | Performance | | test | Chore | | chore | Chore | | ci | Build / CI | | infra | Build / CI | | revert | Bug fix | | release | Chore |
Body MUST follow the repo's PR template (.github/pull_request_template.md), written to the rules in docs/writing-for-readers.md — read that file before drafting. The short version: one screen above the fold, everything else demoted into a <details> block, and "How Has This Been Tested?" answers how the reader verifies this, not which commands you ran.
markdown# Description {THE LEAD — max 3 sentences, no symbol names, no file paths, no SHAs. What changed, who it affects, why it mattered. One number if you have one. A PM must be able to read this and stop.} Fixes #{issue_number} ← MANDATORY for PRs to main (from Step 2; use "Part of #N" if the ticket stays open). Remove this line ONLY for PRs to test/prod. ## Type of change - [x] {Matching type from mapping above} # How Has This Been Tested? {READER-SIDE STEPS. Numbered. What someone else does to confirm this works — where to click, what to run, what they should see, and what they'd have seen before. No internal symbols. If the change has no observable surface, write exactly one line saying so — e.g. "No user-visible surface; verified via unit tests and the log line in the fold below" — and put the developer probe in the fold. Do not invent a UI flow.} 1. {step} 2. {step — expected result} # Risks {Max 3 bullets, only things a reviewer should actively evaluate: trade-offs accepted, risks mitigated but not eliminated, surviving counterarguments from Step 4.5 / `/challenge`. Each bullet: the risk, why it was accepted, what would trigger a revisit. DELETE THIS HEADING ENTIRELY if there is nothing real — do not write "None".} <details> <summary>Engineering detail</summary> {No length limit. Everything the fold exists to hold, in whatever structure fits: - Root cause, with file:line, symbols, commit SHAs, migration ids. - Evidence you ran: validation commands and their output, test counts, benchmark numbers. - Measurements, comparison tables, per-model / per-service breakdowns. - Cross-service impact, performance, UX notes — only where non-obvious. Omit what doesn't apply. - "Not in this PR" — adjacent problems deliberately left alone, and where they're tracked.} </details>
Rules for filling the template:
[x] — only include the checked types, delete all unchecked optionsmain: the issue link is mandatory — the number comes from Step 2; Fixes # closes the ticket on merge, Part of # keeps it open for multi-PR ticketstest/prod: remove the Fixes # line entirelyRe-read your own draft and answer these four. Any failure is a rewrite, not a caveat — and the fix is almost always "move it into the fold", not "delete it".
Then check the budget: wc -w on everything above <details> must be ≤ 200.
Ask the user to confirm the title and body, then create:
bashgh pr create \ --base {base_branch} \ --title "type(scope): subject" \ --body "$(cat <<'EOF' {body} EOF )"
Print the PR URL and a summary:
PR created: {url}
Title: type(scope): subject
Base: {base} <- {head}
Services: {list}
Validation: {pass/fail status per service}Creating the PR is not the finish line. Drive it to a green, review-clean state: poll CI and the review threads, fix everything you can, and surface only what a human must do. Never fabricate approval, self-approve, or mark a human-only gate as done.
bashgh pr checks {pr_number} --repo {owner/repo} --watch # run as a BACKGROUND task
Do not foreground-poll with sleep N && gh pr checks loops — that burns the session doing nothing. Start the watch as a background task (or arm a Monitor on the check status) and continue with other work; when it completes, read the final matrix. Cap the total wait at ~25 min of wall-clock; if checks are still pending past that, report the current matrix and hand back rather than spinning. Treat the pass as "settled" once no check is pending.
For every fail, fetch the failing job's log and classify it:
bashgh run view --repo {owner/repo} --job {job_id} --log-failed | tail -80
| Failure | Class | Action | |---|---|---| | lint / prettier / gofmt / black | auto-fix | run the service's auto-format/fix (make fmt, npm run lint2:fix, poetry run black .), re-validate | | build / vet / type error | auto-fix | fix the code, re-run the service's validation | | unit test | auto-fix if cause is clear | reproduce locally, fix, re-run; if the failure is unrelated/flaky, re-run the job once | | migration version collision (Vlabel-reuse) | auto-fix | rebase on the base branch, then ./api-server/migrations/new-migration.sh to renumber (fresh V + timestamp), move the SQL, delete the old files (see CLAUDE.md → Migrations) | | PR-title / semantic-scope check | auto-fix | correct the title via gh pr edit / gh api (scopes must be lowercase) | | screenshots-for-UI gate (label-prs) | needs human | the rule (external binary) wants real images dragged into the PR body — you cannot upload them; surface it | | flaky / external infra (timeouts, registry 5xx) | retry | gh run rerun; if it re-fails, surface it | | check needing secrets/env you don't have | needs human | surface it |
Fix all auto-fixable failures together, re-run the affected service's validation locally, then follow the single-commit discipline (Steps 5–7): rebase if needed, amend/squash into the one commit, git push --force-with-lease. Do not add fix-up commits.
Fetch review threads, including bot reviewers (e.g. gemini-code-assist, CodeRabbit):
bashgh api repos/{owner/repo}/pulls/{pr_number}/comments --paginate # inline comments gh pr view {pr_number} --repo {owner/repo} --json reviews,comments # review states + top-level
Triage each actionable comment — and verify before acting (bots hallucinate; confirm the symbol/line/behavior actually exists in your diff):
bash gh api -X POST repos/{owner/repo}/pulls/{pr_number}/comments/{comment_id}/replies -f body="..."
Pushing fixes re-runs the checks and re-triggers bot review on synchronize. Return to 11.1. Repeat until every required check is green AND every actionable review comment is fixed or answered. Bound the loop (≤ ~4 fix-and-push cycles); if the same check keeps failing after a genuine fix, stop and surface it — don't loop blindly.
You cannot self-approve or upload screenshots, so stop at the first stable state and report which one it is:
Print the final check matrix, the review-comment disposition (fixed / replied / needs-user), and a one-sentence next action for the user.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,986 | 7,119 | +19% | 1 | 1 | 0% | 579 | 5,563 | +861% | 0 | 0 | — |
case-02 | fail→fail | 17,126 | 29,871 | +74% | 1 | 1 | 0% | 2,334 | 5,420 | +132% | 0 | 0 | — |
case-03 | fail→fail | 9,752 | 6,556 | -33% | 1 | 1 | 0% | 1,518 | 5,448 | +259% | 0 | 0 | — |
case-04 | fail→fail | 7,154 | 12,577 | +76% | 1 | 1 | 0% | 716 | 5,929 | +728% | 0 | 0 | — |
case-05 | fail→fail | 7,474 | 7,561 | +1% | 1 | 1 | 0% | 303 | 5,489 | +1712% | 0 | 0 | — |
case-06 | fail→fail | 7,721 | 7,456 | -3% | 1 | 1 | 0% | 1,006 | 5,477 | +444% | 0 | 0 | — |
case-07 | fail→fail | 12,238 | 4,293 | -65% | 1 | 1 | 0% | 1,702 | 5,763 | +239% | 0 | 0 | — |
case-08 | fail→fail | 8,146 | 7,779 | -5% | 1 | 1 | 0% | 1,172 | 5,520 | +371% | 0 | 0 | — |
case-09 | fail→pass | 8,397 | 4,718 | -44% | 1 | 1 | 0% | 1,447 | 5,515 | +281% | 0 | 0 | — |
case-10 | fail→pass | 8,528 | 3,351 | -61% | 1 | 1 | 0% | 1,228 | 5,491 | +347% | 0 | 0 | — |
case-11 | pass→pass | 10,654 | 8,299 | -22% | 1 | 1 | 0% | 1,727 | 5,579 | +223% | 0 | 0 | — |
case-12 | pass→pass | 9,507 | 4,925 | -48% | 1 | 1 | 0% | 1,406 | 5,556 | +295% | 0 | 0 | — |
case-13 | fail→pass | 8,756 | 5,334 | -39% | 1 | 1 | 0% | 1,241 | 5,949 | +379% | 0 | 0 | — |
case-14 | fail→pass | 7,966 | 4,001 | -50% | 1 | 1 | 0% | 1,206 | 5,735 | +376% | 0 | 0 | — |
case-15 | fail→pass | 7,699 | 4,038 | -48% | 1 | 1 | 0% | 1,041 | 5,857 | +463% | 0 | 0 | — |
case-16 | fail→pass | 13,156 | 7,756 | -41% | 1 | 1 | 0% | 1,836 | 6,473 | +253% | 0 | 0 | — |
case-17 | pass→fail | 12,194 | 8,001 | -34% | 1 | 1 | 0% | 1,774 | 5,517 | +211% | 0 | 0 | — |
case-18 | fail→pass | 15,865 | 5,330 | -66% | 1 | 1 | 0% | 2,624 | 5,972 | +128% | 0 | 0 | — |
case-19 | pass→pass | 14,621 | 7,549 | -48% | 1 | 1 | 0% | 2,188 | 6,176 | +182% | 0 | 0 | — |
case-20 | pass→pass | 14,517 | 5,045 | -65% | 1 | 1 | 0% | 1,988 | 6,116 | +208% | 0 | 0 | — |
case-21 | fail→pass | 13,025 | 5,206 | -60% | 1 | 1 | 0% | 1,840 | 5,966 | +224% | 0 | 0 | — |
case-22 | pass→pass | 10,133 | 4,876 | -52% | 1 | 1 | 0% | 1,470 | 5,797 | +294% | 0 | 0 | — |
case-23 | pass→fail | 8,175 | 6,043 | -26% | 1 | 1 | 0% | 1,237 | 6,048 | +389% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 15 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +26 percentage points is the difference between those two pass rates over the 15 comparable cases. 6 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/27/2026 | +27% |
| gemini-3.6-flash | verified | 8/22/2026 | +52% |
Other measured skills in the registry, with their headline benchmark lift.