Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Write a clear, structured pull request description from a git diff, branch summary, or commit list. Use when asked to write a PR description, draft a pull request, or document code changes. Produces a description with summary, motivation, changes made, testing steps, and reviewer guidance.
.claude/skills/mohitagw15856-pr-description-writer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 123% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 127% | 0% |
Writes structured, reviewer-friendly pull request descriptions from a diff, commit list, or informal notes. Covers the what, why, and how-to-review so reviewers can start immediately.
Ask for these if not provided:
git log --oneline, or describe the changes in plain English)A clear, imperative-mood title under 72 characters: [type]: [concise description of what changed]
Examples:
feat: add rate limiting to the public APIfix: resolve race condition in session expiryrefactor: extract payment logic into PaymentService2–3 sentences covering:
Bullet list of specific changes — one bullet per logical change, not per file:
If UI change: include before/after screenshots or a screen recording] If API change: include example request/response] If no visual change and no API contract change: omit this section entirely — do not leave it as a placeholder]
Step-by-step instructions a reviewer can follow:
Include any specific commands, test data, or environment flags needed.
Flag anything that warrants extra attention:
This skill ships with support files — use them when they are available:
references/reviewer-empathy.md — PR Descriptions as Review Navigation. Apply it while producing the output; it carries the calibration and judgment calls the method summary above compresses.templates/pr-template.md — a fill-in version of the deliverable with the quality gates inline. Offer it when the user wants to work the document themselves rather than have it generated.Score any output of this skill before handing it over; 32+ is ship-quality.
| Dimension | 0 | 5 | 10 | |---|---|---|---| | Why over what | Description only restates the diff; no motivation given | The what is clear, but the why is thin or generic ("improves the code") | Problem, goal, and approach are explicit; the description adds context the diff cannot convey | | Title & structure | Single unstructured paragraph; title vague, missing a type prefix, or over 72 characters | Structured with headers, but the title or section usage slips (placeholder sections left in) | Valid type prefix, imperative mood, under 72 characters; sections let a reviewer navigate straight to what they need | | Testing reproducibility | No testing steps | Steps exist but assume codebase familiarity or skip edge cases | Someone unfamiliar with the code can reproduce verification — commands, test data, flags, and at least one edge case included | | Risk-calibrated reviewer guidance | High-risk PR with no reviewer notes, or notes that are pure boilerplate | Notes exist but name no specific trade-off, uncertainty, or out-of-scope item | Guidance matches risk: high-risk flags specific concerns and deliberate trade-offs; low-risk keeps notes to one line or omits them |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 13,706 | 21,101 | +54% | 1 | 1 | 0% | 2,553 | 3,991 | +56% | 0 | 0 | — |
case-02 | fail→fail | 18,655 | 18,276 | -2% | 1 | 1 | 0% | 1,973 | 3,429 | +74% | 0 | 0 | — |
case-03 | fail→fail | 13,926 | 13,909 | -0% | 1 | 1 | 0% | 1,667 | 2,984 | +79% | 0 | 0 | — |
case-04 | pass→pass | 8,820 | 7,686 | -13% | 1 | 1 | 0% | 444 | 1,809 | +307% | 0 | 0 | — |
case-05 | fail→fail | 35,349 | 82,801 | +134% | 1 | 1 | 0% | 1,079 | 2,423 | +125% | 0 | 0 | — |
case-06 | pass→fail | 16,881 | 24,608 | +46% | 1 | 1 | 0% | 1,593 | 4,613 | +190% | 0 | 0 | — |
case-07 | fail→pass | 17,874 | 14,027 | -22% | 1 | 1 | 0% | 1,996 | 2,889 | +45% | 0 | 0 | — |
case-08 | pass→fail | 19,168 | 23,912 | +25% | 1 | 1 | 0% | 1,953 | 4,189 | +114% | 0 | 0 | — |
case-09 | fail→fail | 14,698 | 17,277 | +18% | 1 | 1 | 0% | 1,396 | 3,251 | +133% | 0 | 0 | — |
case-10 | fail→pass | 21,872 | 21,862 | -0% | 1 | 1 | 0% | 2,644 | 4,206 | +59% | 0 | 0 | — |
case-11 | fail→pass | 11,941 | 13,771 | +15% | 1 | 1 | 0% | 1,139 | 2,539 | +123% | 0 | 0 | — |
case-12 | pass→pass | 16,397 | 11,534 | -30% | 1 | 1 | 0% | 1,764 | 2,570 | +46% | 0 | 0 | — |
case-13 | fail→pass | 18,022 | 23,504 | +30% | 1 | 1 | 0% | 2,210 | 4,070 | +84% | 0 | 0 | — |
case-14 | pass→pass | 27,528 | 22,583 | -18% | 1 | 1 | 0% | 2,368 | 4,023 | +70% | 0 | 0 | — |
case-15 | fail→pass | 15,713 | 17,065 | +9% | 1 | 1 | 0% | 1,495 | 3,392 | +127% | 0 | 0 | — |
case-16 | fail→pass | 18,946 | 21,969 | +16% | 1 | 1 | 0% | 2,064 | 3,632 | +76% | 0 | 0 | — |
case-17 | fail→fail | 17,298 | 19,722 | +14% | 1 | 1 | 0% | 1,501 | 3,790 | +152% | 0 | 0 | — |
case-18 | pass→pass | 119,479 | 25,120 | -79% | 1 | 1 | 0% | 1,494 | 3,214 | +115% | 0 | 0 | — |
case-19 | pass→pass | 9,979 | 40,209 | +303% | 1 | 1 | 0% | 1,615 | 3,117 | +93% | 0 | 0 | — |
case-20 | fail→fail | 12,369 | 23,003 | +86% | 1 | 1 | 0% | 1,618 | 3,186 | +97% | 0 | 0 | — |
case-21 | pass→pass | 11,762 | 16,948 | +44% | 1 | 1 | 0% | 1,621 | 3,675 | +127% | 0 | 0 | — |
case-22 | pass→pass | 19,248 | 14,819 | -23% | 1 | 1 | 0% | 1,727 | 3,148 | +82% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.