Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review a pull request in an isolated worktree, post comments on the PR, and clean up
.claude/skills/nudgebee-review-pr/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 318% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 109% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 365% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 61% | 0% |
Review the pull request specified by $ARGUMENTS (PR number or URL).
If $ARGUMENTS is empty or not provided, you MUST stop and ask the user for a PR number. Do NOT proceed without an explicit PR number or URL. Do NOT fall back to the current branch's PR.
Example prompt: "Please provide a PR number or URL. Usage: /review-pr 123"
Run these commands sequentially in a single bash call:
bashREPO_ROOT=$(git rev-parse --show-toplevel) PR_NUMBER=<pr number from arguments> WORKTREE_DIR="${REPO_ROOT}-pr-review-${PR_NUMBER}" # Fetch PR metadata gh pr view $PR_NUMBER --json number,title,body,baseRefName,headRefName,additions,deletions,changedFiles,author # Fetch latest and create worktree git fetch origin HEAD_BRANCH=$(gh pr view $PR_NUMBER --json headRefName -q .headRefName) git worktree add "$WORKTREE_DIR" "origin/${HEAD_BRANCH}" 2>/dev/null || echo "Worktree already exists" echo "WORKTREE_DIR=${WORKTREE_DIR}"
Store the WORKTREE_DIR path. You will use it for ALL subsequent operations.
From this point forward, you MUST:
$WORKTREE_DIR as the working directory for ALL bash commands (cd $WORKTREE_DIR && ...)$WORKTREE_DIR as the path for ALL Glob and Grep tool callsgh commands can run from anywhere (they query the GitHub API), but even those should cd $WORKTREE_DIR first for consistencyGet the list of changed files from the PR:
bashgh pr diff $PR_NUMBER --name-only
Map changed files to services:
| Path prefix | Service | Type | |---|---|---| | api-server/services/ | api-server | Go | | ticket-server/ | ticket-server | Go | | collector-server/cloud-collector/ | cloud-collector | Go | | collector-server/k8s-collector/relay-server/ | relay-server | Go | | collector-server/k8s-collector/app/ | k8s-collector-app | Python | | llm/code-analysis/ | code-analysis | Go | | llm/llm-server/ | llm-server | Go | | llm/rag-server/ | rag-server | Python | | llm/benchmark/ | benchmark | Python | | ml-k8s-server/ | ml-k8s-server | Python | | auto-pilot/ | auto-pilot | Python | | auto-pilot/sidecar/ | auto-pilot-sidecar | Python | | notifications-server/ | notifications-server | Python | | app/ | frontend | TypeScript | | deploy/ | infrastructure | Helm/K8s |
For each affected service, use the Read tool to check if $WORKTREE_DIR/{service}/CLAUDE.md exists and read it for service-specific conventions.
Read each changed file from the worktree using the Read tool with paths like $WORKTREE_DIR/path/to/changed/file.
Also fetch the full diff for line number context:
bashcd $WORKTREE_DIR && gh pr diff $PR_NUMBER
Analyze the diff and the files you read against these dimensions:
slog for logging, testify for teststype(scope): subject (per .github/semantic.yml)feat, fix, docs, style, refactor, perf, test, chore, revert, ci, infra, releaseautopilot, ml, notifications, ui, tickets, relay, collector, deps, NB-\d+For each finding, post a review comment directly on the PR using gh.
Each inline comment MUST start with a severity icon on the first line, followed by a blank line, then the explanation. Use these exact icon URLs based on severity:
| Severity | Icon markdown | |---|---| | Critical bug / must fix |  | | High importance |  | | Medium importance |  | | Low / nit |  | | Security critical |  followed by  | | Security medium |  followed by  | | Positive / looks good | > ✅ **Looks Good** (blockquote format — the positive.svg URL is a 404) |
Comment structure:
suggestion fenced blocks for single-line replacements, or regular fenced blocks for multi-line examples)Example inline comment body:

The dependency array for this `useEffect` is missing `mode`. This violates the `react-hooks/exhaustive-deps` rule and can lead to bugs from stale closures.
` ` `suggestion
}, [value, mode]);
` ` `IMPORTANT: Posting inline comments
Always use --input with a JSON payload to avoid gh api -f escaping the exclamation mark in image markdown:
bashcd $WORKTREE_DIR && cat <<'JSONEOF' | gh api repos/{owner}/{repo}/pulls/$PR_NUMBER/comments --input - { "body": "<comment-with-icon>", "path": "<file-path-relative-to-repo-root>", "commit_id": "<head-commit-sha>", "position": <diff-hunk-position> } JSONEOF
Note: Use position (the line offset within the diff hunk, starting from 1), NOT line + subject_type. Get the head commit SHA via: gh pr view $PR_NUMBER --json headRefOid -q .headRefOid
Post a single summary comment on the PR with this structure:
bashcd $WORKTREE_DIR && gh pr comment $PR_NUMBER --body "$(cat <<'EOF' ## Summary of Changes {1-3 sentence summary of what this PR does and its motivation} ### Highlights - **{Feature/Fix 1}**: {Brief description} - **{Feature/Fix 2}**: {Brief description} - ... <details> <summary>Changelog</summary> {For each changed file, a bullet with the filename in bold and a sub-list of what changed} - **`path/to/file1.go`** - {change description} - {change description} - **`path/to/file2.ts`** - {change description} </details> ### Review Summary **Services affected:** {list} **Risk level:** Low / Medium / High | Category | Finding | Severity | |---|---|---| | {Correctness/Security/Style/...} | {Brief description — file:line} |  | | {Category} | {Brief description — file:line} |  | | ... | ... | ... | ### Checklist - [ ] Tests cover new/changed behavior - [ ] No hardcoded secrets - [ ] Error handling is appropriate - [ ] No breaking API changes (or documented) - [ ] Linting/formatting passes for affected services --- *Automated PR Review* EOF )"
Rules:
<details> block to avoid overwhelming the summary.After the review is complete and all comments are posted, remove the worktree:
bashREPO_ROOT=$(git rev-parse --show-toplevel) WORKTREE_DIR="${REPO_ROOT}-pr-review-${PR_NUMBER}" git -C "$REPO_ROOT" worktree remove "$WORKTREE_DIR" --force
Verify cleanup:
bashgit -C "$REPO_ROOT" worktree list
Print a brief summary to the user:
Review posted on PR #{number}: {title}
- {N} inline comments posted
- {1} summary comment posted
- Worktree cleaned up
PR URL: {url}| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 6,129 | 7,101 | +16% | 1 | 1 | 0% | 992 | 2,933 | +196% | 0 | 0 | — |
case-02 | pass→fail | 2,565 | 6,542 | +155% | 1 | 1 | 0% | 394 | 2,995 | +660% | 0 | 0 | — |
case-03 | pass→fail | 3,783 | 8,184 | +116% | 1 | 1 | 0% | 692 | 3,375 | +388% | 0 | 0 | — |
case-08 | pass→pass | 9,516 | 3,190 | -66% | 1 | 1 | 0% | 1,551 | 3,182 | +105% | 0 | 0 | — |
case-04 | fail→pass | 4,575 | 1,352 | -70% | 1 | 1 | 0% | 663 | 2,773 | +318% | 0 | 0 | — |
case-05 | fail→fail | 8,389 | 6,967 | -17% | 1 | 1 | 0% | 1,479 | 3,820 | +158% | 0 | 0 | — |
case-06 | pass→pass | 12,999 | 4,408 | -66% | 1 | 1 | 0% | 2,214 | 3,340 | +51% | 0 | 0 | — |
case-07 | fail→pass | 8,673 | 2,076 | -76% | 1 | 1 | 0% | 1,397 | 2,919 | +109% | 0 | 0 | — |
case-09 | fail→pass | 10,842 | 1,648 | -85% | 1 | 1 | 0% | 1,840 | 2,874 | +56% | 0 | 0 | — |
case-10 | fail→pass | 4,452 | 2,838 | -36% | 1 | 1 | 0% | 686 | 3,189 | +365% | 0 | 0 | — |
case-11 | pass→pass | 6,523 | 1,579 | -76% | 1 | 1 | 0% | 997 | 2,826 | +183% | 0 | 0 | — |
case-12 | pass→pass | 3,830 | 2,673 | -30% | 1 | 1 | 0% | 606 | 3,066 | +406% | 0 | 0 | — |
case-13 | fail→pass | 12,646 | 3,310 | -74% | 1 | 1 | 0% | 1,970 | 3,173 | +61% | 0 | 0 | — |
case-14 | fail→fail | 3,351 | 1,733 | -48% | 1 | 1 | 0% | 534 | 2,871 | +438% | 0 | 0 | — |
case-15 | pass→pass | 13,968 | 3,839 | -73% | 1 | 1 | 0% | 2,418 | 3,284 | +36% | 0 | 0 | — |
case-16 | fail→pass | 14,151 | 3,125 | -78% | 1 | 1 | 0% | 2,284 | 3,089 | +35% | 0 | 0 | — |
case-17 | fail→fail | 9,966 | 3,043 | -69% | 1 | 1 | 0% | 1,888 | 3,035 | +61% | 0 | 0 | — |
case-18 | fail→pass | 8,440 | 1,420 | -83% | 1 | 1 | 0% | 1,395 | 2,816 | +102% | 0 | 0 | — |
case-19 | fail→pass | 12,055 | 2,680 | -78% | 1 | 1 | 0% | 1,935 | 3,015 | +56% | 0 | 0 | — |
case-20 | fail→pass | 11,979 | 1,860 | -84% | 1 | 1 | 0% | 2,118 | 2,826 | +33% | 0 | 0 | — |
case-21 | pass→pass | 20,674 | 3,072 | -85% | 1 | 1 | 0% | 1,946 | 3,078 | +58% | 0 | 0 | — |
case-22 | fail→pass | 10,010 | 3,407 | -66% | 1 | 1 | 0% | 1,589 | 3,055 | +92% | 0 | 0 | — |
case-23 | fail→fail | 10,439 | 6,770 | -35% | 1 | 1 | 0% | 1,283 | 2,935 | +129% | 0 | 0 | — |
case-24 | fail→fail | 13,635 | 7,260 | -47% | 1 | 1 | 0% | 951 | 3,065 | +222% | 0 | 0 | — |
case-25 | fail→fail | 6,939 | 6,520 | -6% | 1 | 1 | 0% | 589 | 2,924 | +396% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 20 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +28 percentage points is the difference between those two pass rates over the 20 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.