Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create a pull request with proper formatting, validation, and conventions for this monorepo
.claude/skills/nudgebee-create-pr/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 197% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 160% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 261% | 0% |
Create a pull request for the current branch. Optional argument: $ARGUMENTS (target base branch, defaults to main).
Run these commands to understand the current state:
bash# Current branch and tracking info git branch --show-current git status # Commits on this branch not in base git log main..HEAD --oneline # Full diff against base git diff main...HEAD --stat git diff main...HEAD
If $ARGUMENTS specifies a different base branch (e.g., test or prod), use that instead of main.
Map changed files to services using this table:
| Path prefix | Service | Type | Validation | |---|---|---|---| | api-server/services/ | api-server | Go | make validate | | ticket-server/ | ticket-server | Go | make validate | | collector-server/cloud-collector/ | cloud-collector | Go | make validate | | collector-server/k8s-collector/relay-server/ | relay-server | Go | make validate | | collector-server/k8s-collector/app/ | k8s-collector-app | Python | make lint && make test | | llm/code-analysis/ | code-analysis | Go | make check | | llm/llm-server/ | llm-server | Go | make validate | | llm/rag-server/ | rag-server | Python | make lint && make test | | llm/benchmark/ | benchmark | Python | poetry run pytest | | ml-k8s-server/ | ml-k8s-server | Python | make lint && make test | | auto-pilot/ | auto-pilot | Python | poetry run black --check . && poetry run flake8 . | | auto-pilot/sidecar/ | auto-pilot-sidecar | Python | poetry run black --check . && poetry run flake8 . | | notifications-server/ | notifications-server | Python | poetry run black --check . && poetry run flake8 . | | app/ | frontend | TypeScript | npm run lint2 | | deploy/ | infrastructure | — | Manual review |
For each affected service, run its validation command. Report results to the user. If validation fails, ask the user whether to fix the issues or proceed anyway.
Mandatory. Before pushing, read the diff and run an AI first-pass review. The goal is to catch issues before a human reviewer ever sees them, and to surface residual risks explicitly rather than hoping the reviewer finds them.
Run:
bashgit diff {base_branch}...HEAD
Read the diff in full and evaluate against these dimensions:
CLAUDE.md → AI Coding Principles, surgical changes: every changed line traces directly to the request).slog + testify, Python black 120 + flake8 + mypy, TypeScript oxlint + prettier, commit scope correctness.Categorize each finding into one of three buckets:
If the PR touches shared contracts, DB schema, cross-service behavior, or any architectural decision, also run the logic from /challenge against the diff itself: what are the three strongest reasons this diff is wrong? Include the surviving counterarguments in the Risks & Counterarguments section of the PR body. Skip this sub-step for typo / docs / 1-line fixes.
Search for an existing GitHub issue (open or closed) to link to this PR:
bash# Search using keywords from the branch name and commit messages gh issue list --search "<keywords from branch/commits>" --state all --limit 10 --json number,title,state,url
If related issues are found, present them to the user and ask which to link. If none found, ask the user if they have an issue number. Do NOT auto-create a new issue — only create one if the user explicitly asks.
bash# Ensure branch is pushed git remote -v git push -u origin $(git branch --show-current) 2>&1 || true
If there are unpushed commits, push them before creating the PR.
Based on the commits and diff, generate:
Title format: type(scope): subject (per .github/semantic.yml)
Allowed types: | Type | Use when | |---|---| | feat | New feature or functionality | | fix | Bug fix | | docs" | Documentation only | | style | Formatting, whitespace, no code change | | refactor | Code restructure, no behavior change | | perf | Performance improvement | | test | Adding or updating tests only | | chore | Maintenance, deps, config | | revert | Reverting a previous commit | | ci | CI/CD workflow changes | | infra | Infrastructure, Helm, K8s changes | | release | Release-related changes |
Allowed scopes (required): | Scope | Services / paths | |---|---| | ui | app/ (frontend) | | autopilot | auto-pilot/, auto-pilot/sidecar/ | | ml | ml-k8s-server/, llm/code-analysis/, llm/llm-server/, llm/rag-server/, llm/benchmark/ | | notifications | notifications-server/ | | tickets | ticket-server/ | | relay | collector-server/k8s-collector/relay-server/ | | collector | collector-server/cloud-collector/, collector-server/k8s-collector/app/ | | deps | Dependency updates | | NB-xxx | Ticket number — use for api-server/services/, api-server/migrations/, deploy/, .github/, or any cross-service change |
Examples: fix(ui): handle null state in settings, feat(NB-1234): add Azure onboarding flow
Semantic type → PR "Type of change" mapping: | Semantic type | PR checkbox | |---|---| | feat | New feature | | fix | Bug fix | | docs | Documentation | | style | Chore | | refactor | Refactor | | perf | Performance | | test | Chore | | chore | Chore | | ci | Build / CI | | infra | Build / CI | | revert | Bug fix | | release | Chore |
Body MUST follow the repo's PR template (.github/pull_request_template.md):
markdown# Description {Summary of the changes and the related issue. Include relevant motivation and context. List any dependencies that are required for this change.} Fixes # (issue) ← include if there's a linked issue, otherwise remove this line ## Type of change - [x] {Matching type from mapping above} # How Has This Been Tested? {Describe the tests that you ran to verify your changes. Provide instructions so we can reproduce.} - [x] {Test A description} - [x] {Test B description} --- # Risks & Counterarguments {Residual risks from the AI self-review in Step 3.5 — things that were considered and mitigated but not eliminated, plus any counterarguments from /challenge that the implementation accepted rather than resolved. Format as a bulleted list, each bullet stating: the risk, why it was accepted, and what would trigger a revisit. State "None — fully resolved during self-review" if there are no residual concerns. Do NOT put trivial concerns here; save this section for things a human reviewer should actively evaluate.}
Rules for filling the template:
[x] — only include the checked types, delete all unchecked optionsFixes # line entirelyRisks & Counterarguments section is required — state "None — fully resolved during self-review" if there are no residual concernsAsk the user to confirm the title and body, then create:
bashgh pr create --base {base_branch} --title "type(scope): subject" --body "$(cat <<'EOF' {body} EOF )"
Print the PR URL and a summary:
PR created: {url}
Title: type(scope): subject
Base: {base} <- {head}
Services: {list}
Validation: {pass/fail status per service}| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-12 | pass→pass | 5,460 | 2,438 | -55% | 1 | 1 | 0% | 743 | 2,777 | +274% | 0 | 0 | — |
case-10 | fail→pass | 6,443 | 3,458 | -46% | 1 | 1 | 0% | 1,028 | 3,053 | +197% | 0 | 0 | — |
case-01 | fail→fail | 5,212 | 4,260 | -18% | 1 | 1 | 0% | 326 | 2,680 | +722% | 0 | 0 | — |
case-02 | fail→fail | 6,346 | 4,802 | -24% | 1 | 1 | 0% | 483 | 2,699 | +459% | 0 | 0 | — |
case-03 | fail→fail | 9,226 | 4,729 | -49% | 1 | 1 | 0% | 431 | 2,766 | +542% | 0 | 0 | — |
case-04 | fail→fail | 8,239 | 5,742 | -30% | 1 | 1 | 0% | 1,358 | 2,850 | +110% | 0 | 0 | — |
case-11 | fail→pass | 11,005 | 4,846 | -56% | 1 | 1 | 0% | 1,567 | 3,164 | +102% | 0 | 0 | — |
case-05 | fail→pass | 8,734 | 2,963 | -66% | 1 | 1 | 0% | 1,499 | 2,893 | +93% | 0 | 0 | — |
case-06 | fail→fail | 8,108 | 5,043 | -38% | 1 | 1 | 0% | 1,303 | 2,727 | +109% | 0 | 0 | — |
case-07 | fail→pass | 6,203 | 2,636 | -58% | 1 | 1 | 0% | 1,098 | 2,855 | +160% | 0 | 0 | — |
case-08 | fail→pass | 5,348 | 2,547 | -52% | 1 | 1 | 0% | 809 | 2,921 | +261% | 0 | 0 | — |
case-09 | fail→pass | 4,867 | 3,472 | -29% | 1 | 1 | 0% | 832 | 3,118 | +275% | 0 | 0 | — |
case-13 | pass→pass | 7,375 | 3,512 | -52% | 1 | 1 | 0% | 959 | 3,086 | +222% | 0 | 0 | — |
case-14 | pass→pass | 6,036 | 2,765 | -54% | 1 | 1 | 0% | 863 | 2,969 | +244% | 0 | 0 | — |
case-15 | pass→pass | 8,922 | 3,485 | -61% | 1 | 1 | 0% | 1,474 | 2,964 | +101% | 0 | 0 | — |
case-16 | pass→pass | 6,889 | 3,686 | -46% | 1 | 1 | 0% | 942 | 3,095 | +229% | 0 | 0 | — |
case-17 | fail→fail | 24,294 | 3,614 | -85% | 1 | 1 | 0% | 3,354 | 2,898 | -14% | 0 | 0 | — |
case-18 | fail→pass | 7,198 | 3,043 | -58% | 1 | 1 | 0% | 1,111 | 3,011 | +171% | 0 | 0 | — |
case-19 | fail→pass | 6,438 | 2,514 | -61% | 1 | 1 | 0% | 1,088 | 2,837 | +161% | 0 | 0 | — |
case-20 | fail→fail | 7,691 | 7,323 | -5% | 1 | 1 | 0% | 1,250 | 2,922 | +134% | 0 | 0 | — |
case-21 | fail→pass | 7,783 | 2,637 | -66% | 1 | 1 | 0% | 1,315 | 2,873 | +118% | 0 | 0 | — |
case-22 | fail→fail | 26,665 | 6,696 | -75% | 1 | 1 | 0% | 2,712 | 2,733 | +1% | 0 | 0 | — |
case-23 | fail→fail | 6,518 | 6,410 | -2% | 1 | 1 | 0% | 861 | 2,725 | +216% | 0 | 0 | — |
case-24 | fail→fail | 34,276 | 46,232 | +35% | 1 | 1 | 0% | 309 | 2,730 | +783% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 15 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +38 percentage points is the difference between those two pass rates over the 15 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/22/2026 | +55% |
Other measured skills in the registry, with their headline benchmark lift.