Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Repo-specific review guidance for warp. Only the categories declared overridable by the core review-pr skill may be specialized here.
.claude/skills/warpdotdev-review-pr-local/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 63% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 95% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 76% | 0% |
warpThis skill specializes the core review-pr skill (named in the specializes frontmatter field) and is not functional on its own. Before applying its guidance, confirm the parent skill is installed and resolvable at .agents/skills/review-pr/SKILL.md. If it is missing, install it first by copying the skill directory from the source declared in the specializes_source frontmatter field (warpdotdev/common-skills:.agents/skills/review-pr). Then continue with the guidance below.
This file is a companion to the core review-pr skill. It does not redefine the review output schema, severity labels, safety rules, or evidence rules. It only specializes the override categories the core skill marks as overridable.
.agents/skills/rust-unit-tests/SKILL.md and .agents/skills/gui-integration-test/SKILL.md: flag tests the PR adds that those skills would call out, and new code that should have a test and doesn't. Treat a clear violation of these guidelines as ⚠️ [IMPORTANT], not a nit.AGENTS.md: avoid unnecessary type annotations, prefer imports over long path qualifiers, name context parameters ctx and place them last, remove unused parameters instead of prefixing them with _, and prefer inline format arguments in macros.AGENTS.md — comments carry a maintenance cost, so a new comment should earn its place. Check each added/changed comment individually against every named sub-rule (Minimalist Comments, Strictly "Why" Only, No Line-by-Line Narrations, Clean Docstrings, Single-source of documentation, Don't enumerate function call sites, No "transformation comments") rather than forming one overall impression of the comment's quality. Common issues to flag: comments that restate what the code already says instead of explaining non-obvious why; "transformation" comments that describe the edit rather than the current state (e.g. "this used to ..."); doc comments that narrate a function's internal steps or enumerate its callers; explanations duplicated at a call site or reference that the declaration's doc comment already covers; and existing comments removed as collateral of an otherwise unrelated change. Read the full list in AGENTS.md rather than relying on these examples alone. A comment that is technically accurate, well-written, or explains a subtle/important issue is not exempt from these rules — do not let those qualities substitute for the rule-by-rule check. Treat a confirmed violation as ⚠️ [IMPORTANT], not a nit.log::* / safe_*), review the level choice against .agents/skills/logging-and-error-reporting/SKILL.md: using log::error! for a failure that should be a Sentry issue (only report_error! and panics create issues — log::* at Error/Warn/Info are just breadcrumbs), an inappropriate level for hot paths, and secrets/PII in Info-and-above logs (use the safe_* macros for sensitive detail). For report_error! / report_if_error! calls, run the mandatory audit below instead of relying on a narrative pass._ match arms when an enum can reasonably be matched exhaustively; exhaustive matches are preferred so future variants are surfaced during review.FeatureFlag::YourFlag.is_enabled() over #[cfg(...)] unless the code cannot compile without a compile-time gate.TerminalModel locking when the call stack may already hold the model lock. Prefer passing locked references down the stack and keeping lock scopes short.MouseStateHandle::default() usage during render or event handling. Mouse state handles should be created during construction and then cloned/referenced where needed.This specializes the core skill's Pre-Verdict Audit (error-reporting category). Whenever the diff adds or changes a report_error! or report_if_error! call, this audit is mandatory, no matter how large the diff is — a holistic read-through is not sufficient, and skimming past most of a large migration is exactly how the mass log::error! → report_error! migration merged this form of bug undetected.
Before drafting the body or choosing a verdict: list every report_error! / report_if_error! call the diff adds or changes, one by one with its file:line. For each one, check it against .agents/skills/logging-and-error-reporting/SKILL.md (rules 1–5 and the Anti-patterns block define the exact forms; this list is a lookup index, not a restatement) for:
extra: instead of reported as the payloadanyhow!("{e}") / "{e:?}") instead of preserved via .context() / anyhow::Error::new.context() / extra:Also confirm hot/per-frame or per-message paths use ReportErrorLogMode::OncePerRun where the skill calls for it. Treat a confirmed violation as ⚠️ [IMPORTANT]. The enumerated list is the evidence this audit ran — do not substitute a summary like "spot-checked the report_error! sites."
pr_description.txt and any PR comments available in the workflow context for attached screenshots, GIFs, or videos demonstrating the change end to end., <img ...>, <video ...>), GitHub user-attachment links (e.g. https://github.com/user-attachments/..., https://user-images.githubusercontent.com/...), Loom links, and similar hosted media as valid evidence.Screenshots / Videos section from .github/pull_request_template.md being present but empty does not count as evidence.git diff --check, code-path descriptions, and other textual explanations may supplement visual evidence but do not replace it when visual proof is required.body ## Verdict section to Request changes, even if no other blocking issues were found; the top-level verdict field must be "REJECT" to match. Otherwise, missing visual evidence is not a finding — ordinary tests and checks are sufficient.Request changes.crates/warp_tui or the cell-grid element library at crates/warpui_core/src/elements/tui), acceptable "visual evidence" is a terminal transcript, a render_to_lines / TuiBuffer::to_lines snapshot diff, or a ./script/run-tui capture — NOT a computer_use screenshot or real-display recording (those are for the GUI desktop app). See the tui-verify-change skill. The MouseStateHandle ownership rule still applies to TUI code: the TUI's hover/click elements (TuiHoverable, tui_collapsible) are built on the shared MouseStateHandle and must own the handle outside render (created once, reused) so hover/click state survives rebuilt element trees — so flag inline MouseStateHandle::default() there too. Only the GUI's pixel-based hit-testing specifics are GUI-only.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 7,642 | 9,627 | +26% | 1 | 1 | 0% | 287 | 2,730 | +851% | 0 | 0 | — |
case-04 | pass→pass | 11,529 | 7,556 | -34% | 1 | 1 | 0% | 1,983 | 3,646 | +84% | 0 | 0 | — |
case-05 | pass→pass | 17,231 | 9,470 | -45% | 1 | 1 | 0% | 1,963 | 3,870 | +97% | 0 | 0 | — |
case-06 | pass→pass | 11,685 | 7,503 | -36% | 1 | 1 | 0% | 1,891 | 3,516 | +86% | 0 | 0 | — |
case-07 | fail→pass | 12,121 | 3,064 | -75% | 1 | 1 | 0% | 1,828 | 2,972 | +63% | 0 | 0 | — |
case-08 | fail→pass | 14,536 | 4,147 | -71% | 1 | 1 | 0% | 2,339 | 3,065 | +31% | 0 | 0 | — |
case-09 | pass→pass | 15,530 | 5,989 | -61% | 1 | 1 | 0% | 2,419 | 3,474 | +44% | 0 | 0 | — |
case-10 | fail→pass | 13,346 | 3,691 | -72% | 1 | 1 | 0% | 1,834 | 3,006 | +64% | 0 | 0 | — |
case-11 | fail→pass | 15,464 | 8,134 | -47% | 1 | 1 | 0% | 1,960 | 3,814 | +95% | 0 | 0 | — |
case-12 | pass→pass | 12,692 | 5,647 | -56% | 1 | 1 | 0% | 1,820 | 3,324 | +83% | 0 | 0 | — |
case-13 | fail→pass | 11,879 | 4,315 | -64% | 1 | 1 | 0% | 1,770 | 3,116 | +76% | 0 | 0 | — |
case-14 | fail→pass | 13,534 | 9,255 | -32% | 1 | 1 | 0% | 2,072 | 3,776 | +82% | 0 | 0 | — |
case-15 | pass→pass | 12,291 | 4,821 | -61% | 1 | 1 | 0% | 1,807 | 3,198 | +77% | 0 | 0 | — |
case-16 | fail→pass | 14,708 | 4,699 | -68% | 1 | 1 | 0% | 2,184 | 2,987 | +37% | 0 | 0 | — |
case-17 | pass→pass | 9,962 | 5,270 | -47% | 1 | 1 | 0% | 1,416 | 3,034 | +114% | 0 | 0 | — |
case-18 | fail→pass | 8,307 | 5,728 | -31% | 1 | 1 | 0% | 1,070 | 3,312 | +210% | 0 | 0 | — |
case-19 | fail→pass | 10,626 | 4,461 | -58% | 1 | 1 | 0% | 1,488 | 3,224 | +117% | 0 | 0 | — |
case-20 | pass→pass | 9,443 | 4,056 | -57% | 1 | 1 | 0% | 1,353 | 3,175 | +135% | 0 | 0 | — |
case-21 | pass→pass | 15,603 | 5,765 | -63% | 1 | 1 | 0% | 2,387 | 3,352 | +40% | 0 | 0 | — |
case-22 | fail→pass | 11,629 | 5,968 | -49% | 1 | 1 | 0% | 1,801 | 3,532 | +96% | 0 | 0 | — |
case-23 | pass→pass | 11,785 | 5,759 | -51% | 1 | 1 | 0% | 1,741 | 3,354 | +93% | 0 | 0 | — |
case-01 | fail→fail | 12,407 | 23,306 | +88% | 1 | 1 | 0% | 1,863 | 4,222 | +127% | 0 | 0 | — |
case-02 | fail→fail | 5,711 | 6,228 | +9% | 1 | 1 | 0% | 222 | 2,717 | +1124% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +43 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.