Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Converts a markdown PR writeup or code review (one with ```diff fenced blocks and severity-tagged > [!BLOCKER]/[!MAJOR]/[!MINOR]/[!NIT] callouts) into a single-file 2-column HTML review — unified-diff on the left, severity-tagged annotation cards on the right, top jump-nav listing every finding, mandatory named reviewer footer. Triggers when the markdown-html-orchestrator classifies an input as REVIEW, or when invoked directly via /cs:md-review. Refuses without explicit --reviewer (a code review
.claude/skills/alirezarezvani-md-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 8% | 0% |
The code-review converter from Tier 2 of Shihipar's essay ("Code Review and PR Writeups"). Takes a markdown PR writeup with diff blocks + severity callouts and produces a single-file HTML review with a jump-nav, 2-column diff + annotation layout, and a named reviewer footer.
Three stdlib tools pipeline together:
diff_parser.py → annotation_extractor.py → review_html_renderer.py
(md → diff hunks) (md → severity-tagged (hunks + annotations
annotations attached + tokens → 2-col HTML)
to nearest hunk)| Symptom | Action | |---|---| | markdown-html-orchestrator routes input as REVIEW | Invoke this skill | | User runs /cs:md-review <path>.md directly | Invoke this skill | | Input contains diff fenced blocks + > !MAJOR]/> !BLOCKER]/etc. callouts | Invoke this skill | | Input is a long-form spec / report (no diff blocks) | Route to md-document instead | | Input is a slide deck | Route to md-slides instead | | Input < 100 lines | Refuse (Shihipar threshold) | | Design-system not onboarded | Refuse; surface /cs:design-system |
bash# 1. Parse markdown → diff hunks JSON python3 markdown-html/skills/md-review/scripts/diff_parser.py \ --input <path>.md --output hunks.json # 2. Extract severity-tagged annotations, attach to nearest preceding hunk python3 markdown-html/skills/md-review/scripts/annotation_extractor.py \ --input <path>.md --diff-blocks hunks.json --output annotations.json # 3. Render 2-col HTML (--reviewer is mandatory — refuses without) python3 markdown-html/skills/md-review/scripts/review_html_renderer.py \ --diff-blocks hunks.json --annotations annotations.json \ --reviewer "Jane Doe" --title "PR #123: Add retry logic" \ --output review.html
LGTM markers are present and no severity annotations, a success-tinted "LGTM — no findings flagged" bar--reviewer--reviewer is mandatory. A code review must name a human reviewer. Refuses with exit 3 otherwise. Mirrors research-ops's "named owner" discipline.--- a/file + @@ ... @@ blocks means this isn't a code review — refuses with exit 4 and recommends md-document.aria-label + text. WCAG 1.4.1 enforced at the renderer level.--severity-convention "critical,important,suggestion,nit" swaps tier names; position 0 is most severe. Default is BLOCKER / MAJOR / MINOR / NIT (Google Code Review Developer Guide).<title> and header? Recommended: the actual PR / commit title. Canon: docs-as-context-for-readers.LGTM markers ship as the approval bar? Recommended: yes if there are no severity annotations; otherwise the findings take precedence.md-document — that converter renders prose + tables + code + callouts. This one renders diff hunks + margin annotations.md-slides — that converter splits on --- boundaries. This one is a single-page artifact.{default_output_dir}/review-{slug}.html (path resolved by orchestrator's output_path_resolver.py; collision suffix -2, -3, … by default).
references/ for full citations (diff_rendering_canon, severity_coding, pr_annotation_ux)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 27,193 | 6,471 | -76% | 1 | 1 | 0% | 6,245 | 1,975 | -68% | 0 | 0 | — |
case-02 | fail→fail | 24,676 | 5,841 | -76% | 1 | 1 | 0% | 6,229 | 1,809 | -71% | 0 | 0 | — |
case-03 | fail→fail | 27,038 | 13,402 | -50% | 1 | 1 | 0% | 6,226 | 2,583 | -59% | 0 | 0 | — |
case-13 | fail→pass | 12,095 | 3,014 | -75% | 1 | 1 | 0% | 2,058 | 2,057 | -0% | 0 | 0 | — |
case-04 | fail→pass | 14,484 | 5,974 | -59% | 1 | 1 | 0% | 2,850 | 2,603 | -9% | 0 | 0 | — |
case-05 | fail→pass | 12,331 | 6,012 | -51% | 1 | 1 | 0% | 2,312 | 2,604 | +13% | 0 | 0 | — |
case-06 | fail→pass | 10,341 | 2,574 | -75% | 1 | 1 | 0% | 1,896 | 1,969 | +4% | 0 | 0 | — |
case-07 | pass→pass | 9,203 | 2,305 | -75% | 1 | 1 | 0% | 1,608 | 1,876 | +17% | 0 | 0 | — |
case-08 | fail→pass | 12,773 | 4,860 | -62% | 1 | 1 | 0% | 2,214 | 2,387 | +8% | 0 | 0 | — |
case-09 | fail→pass | 6,753 | 3,057 | -55% | 1 | 1 | 0% | 1,106 | 2,087 | +89% | 0 | 0 | — |
case-10 | pass→pass | 8,682 | 2,546 | -71% | 1 | 1 | 0% | 1,446 | 1,912 | +32% | 0 | 0 | — |
case-11 | fail→pass | 12,783 | 8,345 | -35% | 1 | 1 | 0% | 2,440 | 2,813 | +15% | 0 | 0 | — |
case-12 | fail→pass | 5,529 | 1,581 | -71% | 1 | 1 | 0% | 986 | 1,657 | +68% | 0 | 0 | — |
case-14 | fail→pass | 7,758 | 2,138 | -72% | 1 | 1 | 0% | 1,378 | 1,826 | +33% | 0 | 0 | — |
case-15 | fail→pass | 9,204 | 3,130 | -66% | 1 | 1 | 0% | 1,565 | 1,942 | +24% | 0 | 0 | — |
case-16 | fail→pass | 24,623 | 1,850 | -92% | 1 | 1 | 0% | 4,741 | 1,792 | -62% | 0 | 0 | — |
case-17 | fail→pass | 8,533 | 4,711 | -45% | 1 | 1 | 0% | 1,588 | 2,258 | +42% | 0 | 0 | — |
case-18 | fail→pass | 18,339 | 5,916 | -68% | 1 | 1 | 0% | 3,068 | 2,545 | -17% | 0 | 0 | — |
case-19 | fail→fail | 5,984 | 3,591 | -40% | 1 | 1 | 0% | 892 | 1,889 | +112% | 0 | 0 | — |
case-20 | pass→pass | 12,795 | 8,517 | -33% | 1 | 1 | 0% | 2,159 | 3,019 | +40% | 0 | 0 | — |
case-21 | pass→pass | 8,972 | 6,245 | -30% | 1 | 1 | 0% | 1,981 | 2,680 | +35% | 0 | 0 | — |
case-22 | pass→pass | 4,786 | 3,938 | -18% | 1 | 1 | 0% | 874 | 2,217 | +154% | 0 | 0 | — |
case-23 | pass→pass | 12,504 | 6,881 | -45% | 1 | 1 | 0% | 2,754 | 2,916 | +6% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +57 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.