Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Reference guidance for classifying whether an unmarked PR should appear in the changelog and under which category. Used inline by the changelog-draft skill — not dispatched as a separate agent.
.claude/skills/warpdotdev-classify-changelog-pr/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 103% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 269% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 61% | 0% |
This document provides classification rules for PRs that lack explicit CHANGELOG-* markers. The changelog-draft agent follows these rules inline when deciding whether to include an unmarked PR.
fetch_prs.py marker extraction.CHANGELOG-NONE marker (contributor opted out).github/workflows/), test files, or dev toolingCHANGELOG-* markers (handled before this guidance applies)CHANGELOG-TUI and CHANGELOG-OZ markers as authoritative entries. They are independent and may coexist with each other or with regular changelog categories.TUI.TUI when commit_subject contains TUI as a standalone token.crates/warp_tui or crates/warpui_core/src/elements/tui. Shared capabilities such as Agent tool-call or edit-file behavior also impact TUI when Warp Agent CLI users observe the change. If metadata is ambiguous, inspect the commit diff.TUI and its regular category.DOGFOOD_FLAGS or PREVIEW_FLAGS.PREVIEW_FLAGS. Still exclude DOGFOOD_FLAGS-only changes.If a PR mentions a FeatureFlag variant in its diff or title:
RELEASE_FLAGS, PREVIEW_FLAGS, DOGFOOD_FLAGS).RELEASE_FLAGS or enabled by default in app/Cargo.toml, treat it as live.feature_flag in the classification output to the flag name.needs_review: true.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 6,252 | 5,113 | -18% | 1 | 1 | 0% | 1,123 | 1,872 | +67% | 0 | 0 | — |
case-02 | fail→fail | 9,125 | 8,705 | -5% | 1 | 1 | 0% | 1,425 | 2,228 | +56% | 0 | 0 | — |
case-03 | pass→pass | 6,460 | 4,371 | -32% | 1 | 1 | 0% | 995 | 1,459 | +47% | 0 | 0 | — |
case-04 | fail→pass | 8,522 | 3,423 | -60% | 1 | 1 | 0% | 1,329 | 1,447 | +9% | 0 | 0 | — |
case-05 | pass→pass | 7,149 | 2,655 | -63% | 1 | 1 | 0% | 980 | 1,326 | +35% | 0 | 0 | — |
case-06 | pass→pass | 7,566 | 3,926 | -48% | 1 | 1 | 0% | 1,093 | 1,647 | +51% | 0 | 0 | — |
case-07 | pass→pass | 3,419 | 4,707 | +38% | 1 | 1 | 0% | 488 | 1,636 | +235% | 0 | 0 | — |
case-08 | fail→pass | 4,639 | 4,861 | +5% | 1 | 1 | 0% | 724 | 1,470 | +103% | 0 | 0 | — |
case-09 | pass→pass | 5,544 | 5,790 | +4% | 1 | 1 | 0% | 668 | 1,630 | +144% | 0 | 0 | — |
case-10 | pass→pass | 7,004 | 6,696 | -4% | 1 | 1 | 0% | 935 | 1,944 | +108% | 0 | 0 | — |
case-11 | fail→pass | 4,091 | 6,360 | +55% | 1 | 1 | 0% | 536 | 1,978 | +269% | 0 | 0 | — |
case-12 | fail→pass | 9,148 | 7,667 | -16% | 1 | 1 | 0% | 1,315 | 2,086 | +59% | 0 | 0 | — |
case-13 | fail→pass | 6,150 | 8,955 | +46% | 1 | 1 | 0% | 1,023 | 1,649 | +61% | 0 | 0 | — |
case-14 | pass→pass | 6,228 | 3,702 | -41% | 1 | 1 | 0% | 935 | 1,379 | +47% | 0 | 0 | — |
case-15 | pass→pass | 4,395 | 5,365 | +22% | 1 | 1 | 0% | 620 | 1,836 | +196% | 0 | 0 | — |
case-16 | pass→pass | 13,950 | 5,689 | -59% | 1 | 1 | 0% | 1,918 | 1,884 | -2% | 0 | 0 | — |
case-17 | fail→pass | 15,570 | 5,536 | -64% | 1 | 1 | 0% | 2,431 | 1,650 | -32% | 0 | 0 | — |
case-18 | pass→pass | 11,515 | 6,670 | -42% | 1 | 1 | 0% | 1,436 | 2,032 | +42% | 0 | 0 | — |
case-19 | pass→pass | 7,195 | 5,640 | -22% | 1 | 1 | 0% | 1,061 | 1,421 | +34% | 0 | 0 | — |
case-20 | pass→pass | 7,027 | 3,086 | -56% | 1 | 1 | 0% | 889 | 1,411 | +59% | 0 | 0 | — |
case-21 | pass→fail | 16,338 | 13,435 | -18% | 1 | 1 | 0% | 2,474 | 2,960 | +20% | 0 | 0 | — |
case-22 | pass→pass | 14,477 | 10,673 | -26% | 1 | 1 | 0% | 2,533 | 2,757 | +9% | 0 | 0 | — |
case-23 | pass→pass | 19,375 | 21,304 | +10% | 1 | 1 | 0% | 3,693 | 5,034 | +36% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +22 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.