Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Show ponytail's measured impact as a compact scoreboard: less code, less cost, more speed, from the benchmark medians. One-shot display, not a persistent mode, and not a per-repo number. Trigger: /ponytail-gain, "ponytail gain", "what does ponytail save", "show ponytail impact", "ponytail scoreboard".
.claude/skills/dietrichgebert-ponytail-gain/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -59% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -63% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 153% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 105% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -50% | 0% |
Display this scoreboard when invoked. One-shot: do NOT change mode, write flag files, or persist anything.
The figures are the published benchmark medians (5 everyday tasks: email validator, debounce, CSV sum, countdown timer, rate limiter; three models: Haiku, Sonnet, Opus). They are measured, not computed from the current repo. Source: benchmarks/ and the README.
Render plain ASCII bars. The bar length shows the measured range; the label carries the exact figure:
ponytail gain benchmark median · 5 tasks · 3 models
Lines of code no-skill ████████████████████ 100%
ponytail ██▌················· 6–20% ▼ 80–94%
Cost no-skill ████████████████████ 100%
ponytail █████▌·············· 23–53% ▼ 47–77%
Speed ponytail ▸ 3–6× faster
This repo: /ponytail-debt (shortcuts you deferred)
/ponytail-audit (what's still cuttable)These are benchmark medians, not this repo. NEVER print a per-repo savings number ("you saved X lines/tokens here"): the unbuilt version was never written, so there is no real baseline to subtract from in a live repo. The only real per-repo figures come from /ponytail-debt (a counted ledger), and this card points there instead of inventing one.
One-shot display. Edits nothing, changes no mode. "stop ponytail" or "normal mode": revert.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | fail→fail | 3,026 | 2,636 | -13% | 1 | 1 | 0% | 534 | 873 | +63% | 0 | 0 | — |
case-01 | fail→pass | 14,459 | 2,356 | -84% | 1 | 1 | 0% | 2,342 | 955 | -59% | 0 | 0 | — |
case-02 | fail→pass | 12,308 | 2,179 | -82% | 1 | 1 | 0% | 2,428 | 888 | -63% | 0 | 0 | — |
case-03 | fail→pass | 2,306 | 2,156 | -7% | 1 | 1 | 0% | 348 | 881 | +153% | 0 | 0 | — |
case-04 | fail→pass | 3,108 | 3,075 | -1% | 1 | 1 | 0% | 519 | 1,062 | +105% | 0 | 0 | — |
case-05 | fail→pass | 11,525 | 2,824 | -75% | 1 | 1 | 0% | 2,191 | 1,101 | -50% | 0 | 0 | — |
case-06 | fail→pass | 7,690 | 2,894 | -62% | 1 | 1 | 0% | 1,574 | 1,134 | -28% | 0 | 0 | — |
case-07 | fail→pass | 7,635 | 1,707 | -78% | 1 | 1 | 0% | 1,293 | 712 | -45% | 0 | 0 | — |
case-08 | fail→pass | 6,520 | 1,447 | -78% | 1 | 1 | 0% | 1,081 | 644 | -40% | 0 | 0 | — |
case-09 | fail→pass | 9,778 | 2,779 | -72% | 1 | 1 | 0% | 1,870 | 1,018 | -46% | 0 | 0 | — |
case-10 | fail→pass | 11,426 | 3,819 | -67% | 1 | 1 | 0% | 2,135 | 1,271 | -40% | 0 | 0 | — |
case-11 | fail→pass | 5,987 | 3,071 | -49% | 1 | 1 | 0% | 998 | 1,005 | +1% | 0 | 0 | — |
case-12 | pass→fail | 10,939 | 2,485 | -77% | 1 | 1 | 0% | 2,016 | 913 | -55% | 0 | 0 | — |
case-14 | fail→pass | 4,516 | 1,639 | -64% | 1 | 1 | 0% | 769 | 736 | -4% | 0 | 0 | — |
case-15 | fail→pass | 9,218 | 3,037 | -67% | 1 | 1 | 0% | 1,773 | 1,153 | -35% | 0 | 0 | — |
case-16 | fail→pass | 8,612 | 2,837 | -67% | 1 | 1 | 0% | 1,471 | 1,054 | -28% | 0 | 0 | — |
case-17 | fail→pass | 11,735 | 2,478 | -79% | 1 | 1 | 0% | 1,997 | 1,012 | -49% | 0 | 0 | — |
case-18 | fail→pass | 12,117 | 1,780 | -85% | 1 | 1 | 0% | 2,275 | 701 | -69% | 0 | 0 | — |
case-19 | fail→pass | 6,424 | 2,679 | -58% | 1 | 1 | 0% | 1,272 | 1,040 | -18% | 0 | 0 | — |
case-20 | pass→fail | 5,675 | 13,430 | +137% | 1 | 1 | 0% | 1,059 | 3,322 | +214% | 0 | 0 | — |
case-21 | pass→pass | 7,155 | 7,720 | +8% | 1 | 1 | 0% | 1,292 | 1,831 | +42% | 0 | 0 | — |
case-22 | pass→fail | 5,636 | 2,190 | -61% | 1 | 1 | 0% | 1,114 | 764 | -31% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.