Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create high-quality git commits: review/stage intended changes, split into logical commits, and write clear commit messages (including Conventional Commits). Use when the user asks to commit, craft a commit message, stage changes, or split work into multiple commits.
.claude/skills/davila7-commit-work/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-14 | ✓→✓ | = Same ✓ | 5% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 47% | 0% |
Make commits that are easy to review and safe to ship:
1) Inspect the working tree before staging
git statusgit diff (unstaged)git diff --stat2) Decide commit boundaries (split if needed)
3) Stage only what belongs in the next commit
git add -pgit restore --staged -p or git restore --staged <path>4) Review what will actually be committed
git diff --cached5) Describe the staged change in 1-2 sentences (before writing the message)
6) Write the commit message
type(scope): short summarygit commit -vreferences/commit-message-template.md if helpful.7) Run the smallest relevant verification
8) Repeat for the next commit until the working tree is clean
Provide:
git diff --cached, plus any tests run)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | pass→pass | 9,695 | 5,623 | -42% | 1 | 1 | 0% | 1,615 | 1,693 | +5% | 0 | 0 | — |
case-03 | fail→pass | 8,261 | 10,236 | +24% | 1 | 1 | 0% | 1,434 | 1,947 | +36% | 0 | 0 | — |
case-01 | fail→fail | 4,613 | 2,015 | -56% | 1 | 1 | 0% | 267 | 849 | +218% | 0 | 0 | — |
case-02 | fail→pass | 14,072 | 11,289 | -20% | 1 | 1 | 0% | 2,731 | 2,669 | -2% | 0 | 0 | — |
case-04 | pass→pass | 6,884 | 7,448 | +8% | 1 | 1 | 0% | 1,288 | 1,898 | +47% | 0 | 0 | — |
case-05 | pass→pass | 2,837 | 3,296 | +16% | 1 | 1 | 0% | 478 | 1,112 | +133% | 0 | 0 | — |
case-06 | pass→pass | 11,216 | 9,518 | -15% | 1 | 1 | 0% | 1,854 | 2,006 | +8% | 0 | 0 | — |
case-07 | fail→pass | 5,558 | 4,369 | -21% | 1 | 1 | 0% | 1,106 | 1,365 | +23% | 0 | 0 | — |
case-08 | pass→pass | 4,343 | 4,639 | +7% | 1 | 1 | 0% | 801 | 1,448 | +81% | 0 | 0 | — |
case-09 | pass→pass | 5,930 | 3,692 | -38% | 1 | 1 | 0% | 1,053 | 1,180 | +12% | 0 | 0 | — |
case-10 | pass→pass | 2,916 | 2,621 | -10% | 1 | 1 | 0% | 448 | 966 | +116% | 0 | 0 | — |
case-11 | pass→pass | 7,650 | 5,430 | -29% | 1 | 1 | 0% | 1,358 | 1,522 | +12% | 0 | 0 | — |
case-12 | pass→pass | 9,024 | 6,648 | -26% | 1 | 1 | 0% | 1,498 | 1,665 | +11% | 0 | 0 | — |
case-13 | fail→fail | 8,552 | 7,274 | -15% | 1 | 1 | 0% | 1,419 | 1,705 | +20% | 0 | 0 | — |
case-15 | pass→pass | 7,123 | 3,579 | -50% | 1 | 1 | 0% | 1,321 | 1,174 | -11% | 0 | 0 | — |
case-16 | pass→pass | 4,435 | 2,610 | -41% | 1 | 1 | 0% | 761 | 978 | +29% | 0 | 0 | — |
case-17 | pass→pass | 6,645 | 4,483 | -33% | 1 | 1 | 0% | 1,085 | 1,434 | +32% | 0 | 0 | — |
case-18 | pass→pass | 9,639 | 4,935 | -49% | 1 | 1 | 0% | 1,722 | 1,592 | -8% | 0 | 0 | — |
case-19 | pass→pass | 13,420 | 7,179 | -47% | 1 | 1 | 0% | 2,310 | 1,842 | -20% | 0 | 0 | — |
case-20 | pass→pass | 10,553 | 5,297 | -50% | 1 | 1 | 0% | 1,689 | 1,591 | -6% | 0 | 0 | — |
case-21 | pass→pass | 8,499 | 6,507 | -23% | 1 | 1 | 0% | 1,592 | 1,783 | +12% | 0 | 0 | — |
case-22 | fail→fail | 8,550 | 2,979 | -65% | 1 | 1 | 0% | 1,590 | 1,066 | -33% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.