Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Extract and analyze writing improvements from GitHub PR review comments. Use when asked to show review feedback, style changes, or editorial improvements from a GitHub pull request URL. Handles both explicit suggestions and plain text feedback. Produces structured output comparing original phrasing with reviewer suggestions to help refine future writing.
.claude/skills/evalstate-pr-writing-review/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flashlowest | 87% | 15 |
| gemini-3.1-pro-preview | 100% | 1 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 0% | 0% |
Extract editorial feedback from GitHub PRs to learn from review improvements.
gh installedgh session: gh auth status should show you’re logged inrepo).| Tool | Responsibility | | ----------------- | ----------------------------------------------------------------------- | | Python script | API calls, parsing, file tracking across renames, structured extraction | | LLM analysis | Pattern recognition, paragraph comparison, style lesson synthesis |
bash> **All paths are relative to the directory containing this SKILL.md file.** > Before running any script, first `cd` to that directory or use the full path. # Get suggestions and feedback uv run scripts/extract_pr_reviews.py <pr_url> # Get full first→final comparison for deep analysis uv run scripts/extract_pr_reviews.py <pr_url> --diff # Same as above, but cap each FIRST/FINAL dump to 2k chars for LLM prompting uv run scripts/extract_pr_reviews.py <pr_url> --diff --max-file-chars 2000
--diffbashuv run scripts/extract_pr_reviews.py https://github.com/org/repo/pull/123 --diff
This outputs:
suggestion blocks (supports multiple suggestion blocks per comment)> Tip: add --max-file-chars 2000 to keep each FIRST/FINAL dump lightweight, or pair --diff with --no-files if you only need the suggestion/feedback summaries.
With the script output, perform this analysis:
Create a table of mechanical fixes:
| Pattern | Original | Fixed | | -------------- | ------------------ | ------------------ | | Grammar | "Its easier" | "It's easier" | | Filler removal | "using this way" | "this way" | | Capitalization | "Image Generation" | "image generation" |
For each reviewer feedback comment:
Example:
> Feedback: "would be nice to end more enthusiastically" > > First draft: "...it's simple to add new tools to Claude and use them straight away." > > Final: "...Let us know what you find and create in the comments below!" > > Lesson: End blog posts with a call-to-action
Compare FIRST DRAFT to FINAL VERSION section by section:
Group findings into categories:
| Category | Patterns Found | | ------------- | ------------------------------------------------ | | Clarity | Passive→active, shorter sentences, remove filler | | Precision | Vague→specific, "Create"→"Generate" | | Tone | Added enthusiasm, call-to-action endings | | Structure | Added transitions, better section flow | | Grammar | its/it's, subject-verb agreement | | Content | Added links, examples, context |
> 📁 All paths are relative to the directory containing this SKILL.md file.
| Flag | Output | Use Case | | -------------------- | -------------------------------------------------------------------------------- | ----------------------------------------------------------- | | _(none)_ | Suggestions + feedback | Quick review of what reviewers said | | --diff | Adds FIRST/FINAL file dumps to the default output | Deep analysis of how the author responded | | --max-file-chars N | Truncates each FIRST/FINAL block to _N_ chars (appends ...[truncated X chars]) | Keep prompts within LLM token limits | | --no-files | Suppresses FIRST/FINAL dumps even when --diff is set | When you only need explicit suggestions + reviewer feedback | | --json | Raw JSON (includes file_evolutions when --diff without --no-files) | Programmatic processing |
> Input formats: pass either a full PR URL, owner/repo PR_NUMBER, or owner repo PR_NUMBER.
--diff.md/.txt/.rst/.mdx file; add --max-file-chars to truncate each block with a visible ...[truncated X chars] indicatorThe script traces files through renames by:
claudeimages.md → claude-images.md → claude-and-mcp.md)After running the script and performing LLM analysis, produce a summary like:
markdown## Style Lessons from PR #123 ### Mechanical Fixes - Fix grammar: "Its" → "It's" (contraction) - Lowercase generic terms: "Image Generation" → "image generation" - Remove filler: "the output quality of" → "the quality of" ### Reviewer-Driven Changes - **"end more enthusiastically"** → Added call-to-action in conclusion - **"emphasize these are SoTA"** → Changed "latest" to "state-of-the-art" - **"add blurb about MCP Server"** → Added explanatory paragraph ### Structural Improvements - Added transition sentence between sections - Simplified setup instructions (3 sentences → 1) - Added new bullet point for model flexibility
--max-file-chars or pass --no-files to keep outputs prompt-friendly{{currentDate}} {{env}}
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-15 | fail→pass | 12,404 | 7,561 | -39% | 1 | 1 | 0% | 2,196 | 2,995 | +36% | 0 | 0 | — |
case-01 | fail→fail | 19,628 | 5,352 | -73% | 1 | 1 | 0% | 3,451 | 1,939 | -44% | 0 | 0 | — |
case-02 | fail→pass | 15,015 | 6,087 | -59% | 1 | 1 | 0% | 2,492 | 2,782 | +12% | 0 | 0 | — |
case-03 | fail→fail | 6,838 | 5,931 | -13% | 1 | 1 | 0% | 1,176 | 2,127 | +81% | 0 | 0 | — |
case-04 | fail→pass | 16,022 | 2,363 | -85% | 1 | 1 | 0% | 2,609 | 2,020 | -23% | 0 | 0 | — |
case-05 | pass→pass | 8,261 | 1,997 | -76% | 1 | 1 | 0% | 1,291 | 1,970 | +53% | 0 | 0 | — |
case-06 | pass→pass | 7,771 | 2,047 | -74% | 1 | 1 | 0% | 1,336 | 1,992 | +49% | 0 | 0 | — |
case-07 | pass→pass | 8,254 | 4,247 | -49% | 1 | 1 | 0% | 1,515 | 2,476 | +63% | 0 | 0 | — |
case-08 | fail→pass | 9,336 | 2,555 | -73% | 1 | 1 | 0% | 1,864 | 2,089 | +12% | 0 | 0 | — |
case-09 | pass→pass | 11,745 | 3,920 | -67% | 1 | 1 | 0% | 2,300 | 2,331 | +1% | 0 | 0 | — |
case-10 | fail→pass | 15,900 | 4,681 | -71% | 1 | 1 | 0% | 2,485 | 2,490 | +0% | 0 | 0 | — |
case-11 | fail→pass | 10,716 | 4,824 | -55% | 1 | 1 | 0% | 1,839 | 2,572 | +40% | 0 | 0 | — |
case-12 | fail→pass | 14,643 | 12,713 | -13% | 1 | 1 | 0% | 2,555 | 4,020 | +57% | 0 | 0 | — |
case-13 | fail→pass | 12,094 | 2,829 | -77% | 1 | 1 | 0% | 2,142 | 2,126 | -1% | 0 | 0 | — |
case-14 | fail→pass | 8,700 | 2,724 | -69% | 1 | 1 | 0% | 1,527 | 2,144 | +40% | 0 | 0 | — |
case-16 | fail→pass | 9,958 | 2,093 | -79% | 1 | 1 | 0% | 2,001 | 2,113 | +6% | 0 | 0 | — |
case-17 | pass→pass | 10,472 | 3,237 | -69% | 1 | 1 | 0% | 1,980 | 2,182 | +10% | 0 | 0 | — |
case-18 | fail→pass | 14,630 | 5,402 | -63% | 1 | 1 | 0% | 2,631 | 2,638 | +0% | 0 | 0 | — |
case-19 | fail→pass | 9,222 | 2,569 | -72% | 1 | 1 | 0% | 1,401 | 2,104 | +50% | 0 | 0 | — |
case-20 | pass→pass | 14,092 | 6,052 | -57% | 1 | 1 | 0% | 2,925 | 2,780 | -5% | 0 | 0 | — |
case-21 | pass→pass | 10,563 | 7,246 | -31% | 1 | 1 | 0% | 1,769 | 2,438 | +38% | 0 | 0 | — |
case-22 | fail→pass | 9,371 | 4,199 | -55% | 1 | 1 | 0% | 1,915 | 2,440 | +27% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.