Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create or update GitHub issues (from PR, from draft body, retitle existing). Routes to the canonical flow in `mem:workflow/creating-issues`.
.claude/skills/penpot-create-issue/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -57% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -68% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -45% | 0% |
Entry point for all GitHub issue work. All rules (title derivation, metadata, body templates, Issue Type IDs), all flows, and all gh / GraphQL commands live in mem:workflow/creating-issues (file: .serena/memories/workflow/creating-issues.md). This skill routes to the right flow.
the PR is the implementation. Issue = WHAT, PR = HOW. → memory section Creating Issues from PRs
yet. → memory section Creating Issues from Draft Body
issue; create it first, then link it to its parent. → memory section Adding an Issue as a Sub-issue
→ memory section Retitling an Existing Issue
Everything else (title derivation, metadata policy, body templates, Issue Type IDs, create/verify/cleanup commands) lives in the memory — go to the matching section there.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 10,069 | 7,538 | -25% | 1 | 1 | 0% | 1,396 | 594 | -57% | 0 | 0 | — |
case-02 | fail→fail | 10,568 | 135,834 | +1185% | 1 | 1 | 0% | 1,651 | 591 | -64% | 0 | 0 | — |
case-03 | fail→fail | 7,536 | 14,491 | +92% | 1 | 1 | 0% | 271 | 772 | +185% | 0 | 0 | — |
case-04 | pass→fail | 4,061 | 6,994 | +72% | 1 | 1 | 0% | 476 | 541 | +14% | 0 | 0 | — |
case-05 | pass→fail | 11,934 | 9,980 | -16% | 1 | 1 | 0% | 864 | 547 | -37% | 0 | 0 | — |
case-06 | pass→fail | 10,137 | 8,913 | -12% | 1 | 1 | 0% | 1,459 | 683 | -53% | 0 | 0 | — |
case-07 | pass→fail | 17,401 | 8,360 | -52% | 1 | 1 | 0% | 810 | 702 | -13% | 0 | 0 | — |
case-08 | fail→pass | 10,812 | 8,870 | -18% | 1 | 1 | 0% | 1,728 | 1,625 | -6% | 0 | 0 | — |
case-09 | fail→fail | 10,743 | 22,893 | +113% | 1 | 1 | 0% | 1,634 | 748 | -54% | 0 | 0 | — |
case-10 | fail→fail | 7,929 | 7,442 | -6% | 1 | 1 | 0% | 1,158 | 736 | -36% | 0 | 0 | — |
case-11 | fail→fail | 8,570 | 40,921 | +377% | 1 | 1 | 0% | 448 | 674 | +50% | 0 | 0 | — |
case-12 | fail→fail | 8,819 | 3,585 | -59% | 1 | 1 | 0% | 1,224 | 891 | -27% | 0 | 0 | — |
case-13 | fail→fail | 9,075 | 3,238 | -64% | 1 | 1 | 0% | 1,209 | 680 | -44% | 0 | 0 | — |
case-14 | fail→pass | 13,508 | 5,150 | -62% | 1 | 1 | 0% | 1,857 | 806 | -57% | 0 | 0 | — |
case-15 | fail→pass | 13,621 | 3,206 | -76% | 1 | 1 | 0% | 2,254 | 717 | -68% | 0 | 0 | — |
case-16 | fail→pass | 18,630 | 4,630 | -75% | 1 | 1 | 0% | 1,670 | 799 | -52% | 0 | 0 | — |
case-17 | fail→pass | 11,102 | 4,419 | -60% | 1 | 1 | 0% | 1,637 | 899 | -45% | 0 | 0 | — |
case-18 | fail→pass | 9,337 | 2,972 | -68% | 1 | 1 | 0% | 1,291 | 659 | -49% | 0 | 0 | — |
case-19 | fail→pass | 26,038 | 2,779 | -89% | 1 | 1 | 0% | 1,556 | 651 | -58% | 0 | 0 | — |
case-20 | fail→pass | 7,993 | 4,906 | -39% | 1 | 1 | 0% | 1,045 | 790 | -24% | 0 | 0 | — |
case-21 | pass→pass | 6,087 | 2,598 | -57% | 1 | 1 | 0% | 837 | 584 | -30% | 0 | 0 | — |
case-22 | fail→pass | 17,855 | 3,332 | -81% | 1 | 1 | 0% | 1,048 | 725 | -31% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 12 counted toward the lift figure. The other 10 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 12 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.