Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Investigate a pull request end-to-end — predict how it could be wrong, run it for real, falsify those predictions with evidence, and report by severity
.claude/skills/nudgebee-test-a-pull-request/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 5 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 162% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 125% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 205% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 217% | 0% |
| case-23 | ✗→✓ | ▲ Improved | 158% | 0% |
Investigate the pull request specified by $ARGUMENTS (a PR number or URL) and decide whether the change actually works — and where it might not.
This is an investigation, not a test run. Behave like a skeptical senior engineer trying to prove the implementation wrong before accepting it, not an automated runner confirming the happy path. The mental model is:
This skill is methodology, not a checklist. Read the change, form hypotheses about how it's wrong, then design the tests that would catch it — executed against the repository's own conventions.
This overlaps in intent with other review-oriented skills (static diff review, code-judging). Its distinct value is going further than any of them: actually bringing the affected service(s) up and exercising the changed path end-to-end, with evidence collected from the running system rather than from reading the diff alone. Prefer those other skills for pure static/diff-level review; use this one when the change has runtime behavior worth proving.
Core principles (apply throughout):
package.json scripts, .github/ workflows, CLAUDE.md/service docs, pyproject.toml, go.mod). Never invent a command.If $ARGUMENTS is empty, stop and ask for a PR number or URL. Usage: /test-pr 123. Do not fall back to the current branch. Confirm gh auth status; if unauthenticated, stop and tell the user.
Fetch metadata and check the head out in a git worktree so the user's tree is untouched:
bashREPO_ROOT=$(git rev-parse --show-toplevel) # Normalize $ARGUMENTS (a PR number OR a full URL) to a clean integer, so the # snippet is directly runnable and the worktree path stays slash/colon-free. PR=$(gh pr view "$ARGUMENTS" --json number -q .number) gh pr view "$PR" --json number,title,body,state,baseRefName,headRefName,additions,deletions,changedFiles,author,url BASE_BRANCH=$(gh pr view "$PR" --json baseRefName -q .baseRefName) WT="${REPO_ROOT}-testpr-${PR}" BRANCH="testpr-${PR}" # Fetch by PR ref, not by head-branch name — the head branch only exists on # `origin` for same-repo PRs; fork PRs 404 there. `pull/$PR/head` resolves # either way (the same primitive `gh pr checkout` uses under the hood). git fetch origin "pull/${PR}/head:${BRANCH}" git worktree add "$WT" "$BRANCH" 2>/dev/null || git worktree add "$WT" "$BRANCH" --force echo "WT=$WT BASE=$BASE_BRANCH"
Run all file/build/test operations inside $WT. Remember $BASE_BRANCH and the merge-base for failure triage.
Fixes #N) — the claimed behavior and any acceptance criteria.gh pr diff $PR). For each hunk, ask what observable behavior it adds/changes/removes.Estimate the blast radius — that, not a fixed rule, sets test depth.
Parallelize large or multi-service PRs. Dispatch independent Task agents concurrently and synthesize their results yourself — e.g. Task A: map architecture (callers/consumers of the changed symbols); Task B: discover the toolchain (Step 4); Task C: read the relevant service CLAUDE.md/docs; Task D: summarize what each changed component does. Keep the hypotheses and testing decisions in the main thread.
Ask outright: "Can this implementation be wrong?" Generate several concrete hypotheses for how it could be incorrect, then design tests intended to disprove them — not to reconfirm the expected path. Cover at least:
Then two staff-engineer moves:
Output a short, ranked list of risks/hypotheses. This list drives Step 5.
For every affected component, find from the repo how to build/typecheck, lint/format (match what .github/workflows/* actually runs), unit/integration test (and how to run a single changed package/test), and — when a live run is warranted — run the service and what it depends on. Prefer the repo's own aggregate command (make validate, npm run lint2 && npm run test) as its definition of "good."
Choose the minimum set of tiers that falsifies your Step-3 hypotheses and exercises the change's observable behavior. Justify inclusion/exclusion in the report.
CLAUDE.md "Required Services" / "Local Development" section, a docker-compose, a dev script). Do not invent your own. Read that service's CLAUDE.md before running it.> Example: for a change to an inference/backend service, the documented topology is often "port-forward the dependencies it calls (its API/services server, a relay, a RAG/vector service, redis, a database) and run the changed service locally pointed at those." Take the actual service list, ports, and commands from the repo, not from this example.
.env loaded by the app's config layer isn't always visible to code reading the OS env directly (os.Getenv/os.environ). If a value is in .env but the service reports it "not set," export it into the launching process's environment, and verify it actually loaded..env / local config as needed — deliberately and reversibly. Adjusting local config to wire services or unlock the changed path is expected. Read the current value first, note the original, make the minimal edit, and restore it during cleanup. Never commit these; never edit config outside the test workspace without saying so.For stateful changes, record the "before" (row counts, current values, output) so the "after" is provable. Run the chosen tiers, capturing concrete evidence per check: command + relevant output, query result, status code, IDs. Keep artifacts (IDs, baselines, log paths) in the scratchpad for resumability.
A failure is only a PR finding if the PR caused it.
$BASE_BRANCH (a second worktree is cheapest). Fails there too → pre-existing (report as context). Passes there → introduced by the PR (real finding).Never report a failure you haven't classified.
If the change emits data/output a human or system consumes (audit record, API payload, UI element, log format), judge whether it's good: coherent, complete, and useful for its consumer; field semantics consistent with sibling cases; nothing mislabeled, double-encoded, dropped, or ambiguous. These "works but the data is wrong/unhelpful" issues pass a green build and still ship a bad feature.
Produce a concise, skimmable, evidence-dense report:
Each finding: what's wrong, why it matters, file:line, and whether it's introduced vs pre-existing/environmental.
Keep every claim tied to evidence. If you couldn't verify something, say so rather than implying a pass. Post to the PR (gh pr comment $PR -F <file>) only if the user asks — it's outward-facing.
Always attempt cleanup, even if the investigation aborts on an error — don't exit on the first fatal failure leaving state behind, and don't force your way past it either. In order:
.env/config value you changed first, from the backup you noted, before touching the worktree.--force (git worktree remove "$WT"). A clean, restored worktree removes cleanly. If Git refuses because something is still dirty, that's a signal something wasn't restored — go find and restore it, don't paper over it with --force. Only reach for --force as a last resort after you've confirmed by hand (git -C "$WT" status) that anything left dirty is disposable (e.g. build artifacts you created), never blind, and say so in the report.Skip this ordering only if the user asked to keep the environment up — then just report what's left running and unrestored.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | fail→fail | 19,424 | 12,012 | -38% | 1 | 1 | 0% | 1,218 | 4,266 | +250% | 0 | 0 | — |
case-05 | fail→fail | 40,737 | 11,506 | -72% | 1 | 1 | 0% | 5,523 | 4,298 | -22% | 0 | 0 | — |
case-01 | pass→pass | 10,967 | 17,349 | +58% | 1 | 1 | 0% | 1,514 | 5,843 | +286% | 0 | 0 | — |
case-02 | fail→fail | 10,869 | 18,927 | +74% | 1 | 1 | 0% | 862 | 4,381 | +408% | 0 | 0 | — |
case-03 | pass→pass | 16,693 | 29,368 | +76% | 1 | 1 | 0% | 1,770 | 5,824 | +229% | 0 | 0 | — |
case-06 | fail→fail | 11,786 | 12,278 | +4% | 1 | 1 | 0% | 303 | 4,236 | +1298% | 0 | 0 | — |
case-07 | pass→pass | 14,641 | 6,700 | -54% | 1 | 1 | 0% | 1,465 | 4,946 | +238% | 0 | 0 | — |
case-08 | fail→fail | 18,146 | 12,599 | -31% | 1 | 1 | 0% | 1,804 | 5,062 | +181% | 0 | 0 | — |
case-09 | pass→pass | 17,232 | 15,951 | -7% | 1 | 1 | 0% | 1,670 | 5,537 | +232% | 0 | 0 | — |
case-10 | pass→pass | 16,710 | 8,529 | -49% | 1 | 1 | 0% | 1,617 | 5,159 | +219% | 0 | 0 | — |
case-11 | pass→pass | 12,422 | 11,964 | -4% | 1 | 1 | 0% | 1,873 | 5,108 | +173% | 0 | 0 | — |
case-12 | fail→pass | 22,497 | 14,332 | -36% | 1 | 1 | 0% | 2,408 | 6,307 | +162% | 0 | 0 | — |
case-13 | pass→pass | 19,844 | 15,255 | -23% | 1 | 1 | 0% | 2,293 | 5,648 | +146% | 0 | 0 | — |
case-14 | pass→pass | 16,098 | 25,653 | +59% | 1 | 1 | 0% | 1,544 | 5,089 | +230% | 0 | 0 | — |
case-15 | pass→fail | 14,827 | 13,340 | -10% | 1 | 1 | 0% | 2,298 | 4,375 | +90% | 0 | 0 | — |
case-16 | pass→fail | 19,000 | 8,286 | -56% | 1 | 1 | 0% | 1,968 | 4,391 | +123% | 0 | 0 | — |
case-17 | pass→pass | 16,502 | 10,240 | -38% | 1 | 1 | 0% | 1,486 | 5,375 | +262% | 0 | 0 | — |
case-18 | fail→pass | 20,873 | 4,616 | -78% | 1 | 1 | 0% | 2,037 | 4,582 | +125% | 0 | 0 | — |
case-19 | pass→pass | 18,412 | 13,195 | -28% | 1 | 1 | 0% | 2,085 | 5,286 | +154% | 0 | 0 | — |
case-20 | pass→pass | 17,452 | 10,437 | -40% | 1 | 1 | 0% | 1,912 | 5,641 | +195% | 0 | 0 | — |
case-21 | fail→pass | 10,522 | 8,513 | -19% | 1 | 1 | 0% | 1,477 | 4,511 | +205% | 0 | 0 | — |
case-22 | fail→pass | 13,817 | 10,593 | -23% | 1 | 1 | 0% | 1,771 | 5,621 | +217% | 0 | 0 | — |
case-23 | fail→pass | 16,443 | 9,647 | -41% | 1 | 1 | 0% | 1,793 | 4,629 | +158% | 0 | 0 | — |
case-24 | pass→pass | 11,285 | 3,832 | -66% | 1 | 1 | 0% | 884 | 4,493 | +408% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 18 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +13 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.