Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Evaluates a PR's title and description for readability — do they clearly and concisely convey what changed and why to a reviewer? Produces findings with concrete proposed rewrites; the caller decides whether to apply them or present them as feedback. Use when finalizing a PR or reviewing PR metadata. For code comments, docstrings, and naming, use reviewing-readability instead.
.claude/skills/streamlit-reviewing-pr-description/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 38% | 0% |
Review a PR's title and description for readability: do they clearly and concisely tell a reviewer what changed and why? Focus on the prose — checking the format (title pattern, required template sections) is a secondary, lighter concern.
This skill only evaluates: it produces findings with concrete proposed rewrites and does not apply them. The caller decides whether to apply the rewrites or present them as feedback.
The reader is a reviewer or teammate skimming the PR to understand what changed and why. They may not know the implementation context, and later readers will find this text via the commit log or changelog. The title and description should stand on their own.
use_container_width" or "The server now rejects oversized uploads" reads more directly than passive or vague phrasing.gh pr view <n> --json title,body).[type] Description within ~63 chars, and the required template sections from .github/pull_request_template.md are present. For the full standards, see creating-pull-requests and wiki/pull-requests.md.For the title and for the description, give the issue and a concrete proposed rewrite.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,481 | 9,440 | +46% | 1 | 1 | 0% | 898 | 1,250 | +39% | 0 | 0 | — |
case-02 | fail→fail | 28,234 | 10,295 | -64% | 1 | 1 | 0% | 2,205 | 2,366 | +7% | 0 | 0 | — |
case-03 | fail→fail | 14,391 | 5,643 | -61% | 1 | 1 | 0% | 793 | 1,041 | +31% | 0 | 0 | — |
case-04 | fail→fail | 17,702 | 9,736 | -45% | 1 | 1 | 0% | 352 | 1,243 | +253% | 0 | 0 | — |
case-05 | fail→fail | 33,873 | 24,521 | -28% | 1 | 1 | 0% | 1,402 | 2,630 | +88% | 0 | 0 | — |
case-06 | fail→fail | 4,635 | 13,903 | +200% | 1 | 1 | 0% | 597 | 3,135 | +425% | 0 | 0 | — |
case-07 | pass→pass | 10,565 | 17,312 | +64% | 1 | 1 | 0% | 1,630 | 1,981 | +22% | 0 | 0 | — |
case-08 | fail→pass | 12,021 | 8,461 | -30% | 1 | 1 | 0% | 1,655 | 2,243 | +36% | 0 | 0 | — |
case-09 | pass→pass | 12,396 | 7,600 | -39% | 1 | 1 | 0% | 1,940 | 2,069 | +7% | 0 | 0 | — |
case-10 | fail→pass | 11,729 | 9,316 | -21% | 1 | 1 | 0% | 1,811 | 2,220 | +23% | 0 | 0 | — |
case-11 | fail→pass | 14,725 | 20,114 | +37% | 1 | 1 | 0% | 2,160 | 2,737 | +27% | 0 | 0 | — |
case-12 | pass→pass | 11,068 | 8,667 | -22% | 1 | 1 | 0% | 1,598 | 2,198 | +38% | 0 | 0 | — |
case-13 | pass→pass | 14,640 | 21,500 | +47% | 1 | 1 | 0% | 1,752 | 1,815 | +4% | 0 | 0 | — |
case-14 | pass→pass | 16,627 | 10,117 | -39% | 1 | 1 | 0% | 1,603 | 2,295 | +43% | 0 | 0 | — |
case-15 | fail→pass | 12,608 | 21,028 | +67% | 1 | 1 | 0% | 1,898 | 2,304 | +21% | 0 | 0 | — |
case-16 | pass→pass | 10,847 | 12,417 | +14% | 1 | 1 | 0% | 1,687 | 2,082 | +23% | 0 | 0 | — |
case-17 | pass→pass | 11,666 | 16,164 | +39% | 1 | 1 | 0% | 1,681 | 1,930 | +15% | 0 | 0 | — |
case-18 | fail→fail | 13,868 | 5,284 | -62% | 1 | 1 | 0% | 909 | 1,011 | +11% | 0 | 0 | — |
case-19 | fail→fail | 4,267 | 6,476 | +52% | 1 | 1 | 0% | 500 | 1,132 | +126% | 0 | 0 | — |
case-20 | fail→pass | 33,651 | 8,592 | -74% | 1 | 1 | 0% | 1,605 | 2,208 | +38% | 0 | 0 | — |
case-21 | pass→pass | 15,135 | 10,809 | -29% | 1 | 1 | 0% | 2,141 | 2,278 | +6% | 0 | 0 | — |
case-22 | pass→pass | 12,114 | 5,246 | -57% | 1 | 1 | 0% | 936 | 1,616 | +73% | 0 | 0 | — |
case-23 | pass→pass | 11,995 | 7,992 | -33% | 1 | 1 | 0% | 1,467 | 1,971 | +34% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 18 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +22 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.