Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use only when the user explicitly asks for TDD, a failing test, or a regression test, OR when the bug has an obvious cheap local test target. Skip when the test path is unclear, expensive, integration-heavy, or not requested.
.claude/skills/kunanonj-cursor-plugin-pstack-tdd/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 26 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-22 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-06 | ✓→✗ | ▼ Worse | -49% | 0% |
| case-12 | ✓→✗ | ▼ Worse | 6% | 0% |
| case-18 | ✓→✗ | ▼ Worse | 3% | 0% |
| case-13 | ✓→✓ | = Same ✓ | 17% | 0% |
When fixing a bug with a clear, cheap test path, make the broken behavior executable before changing production code. The goal is a focused regression test that fails before the fix and passes after it.
Do not force a test when it would be impractical. If the available test would require broad harness setup, brittle mocks, slow end-to-end infrastructure, production-only state, vague reproduction steps, or large unrelated fixture churn, skip adding a new test and use the closest useful verification instead.
Do not silently skip the regression step. Before fixing, explicitly explain why a failing test is impossible or not worth the cost, then choose the closest executable regression check available. Examples include a targeted script, manual reproduction command, browser automation, snapshot comparison, log assertion, or focused integration check.
Prefer no new test over a bad test. A bad test is one that mostly tests mocks, encodes current implementation details, depends on timing or unrelated global state, needs expensive infrastructure for a small fix, or would be deleted immediately after proving the fix.
Report the evidence, not just the outcome:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | pass→pass | 9,488 | 7,185 | -24% | 1 | 1 | 0% | 1,602 | 1,877 | +17% | 0 | 0 | — |
case-14 | pass→pass | 7,911 | 5,684 | -28% | 1 | 1 | 0% | 1,525 | 1,728 | +13% | 0 | 0 | — |
case-01 | fail→fail | 2,207 | 4,443 | +101% | 1 | 1 | 0% | 249 | 912 | +266% | 0 | 0 | — |
case-02 | fail→fail | 8,139 | 7,343 | -10% | 1 | 1 | 0% | 276 | 776 | +181% | 0 | 0 | — |
case-03 | fail→fail | 3,924 | 3,450 | -12% | 1 | 1 | 0% | 236 | 833 | +253% | 0 | 0 | — |
case-04 | pass→pass | 15,179 | 13,629 | -10% | 1 | 1 | 0% | 2,730 | 3,288 | +20% | 0 | 0 | — |
case-05 | pass→pass | 12,376 | 11,859 | -4% | 1 | 1 | 0% | 2,002 | 2,257 | +13% | 0 | 0 | — |
case-06 | pass→fail | 8,061 | 3,569 | -56% | 1 | 1 | 0% | 1,591 | 809 | -49% | 0 | 0 | — |
case-07 | pass→pass | 13,037 | 9,386 | -28% | 1 | 1 | 0% | 2,380 | 2,243 | -6% | 0 | 0 | — |
case-08 | fail→fail | 12,669 | 8,134 | -36% | 1 | 1 | 0% | 2,385 | 2,128 | -11% | 0 | 0 | — |
case-09 | pass→pass | 13,240 | 5,446 | -59% | 1 | 1 | 0% | 2,330 | 1,646 | -29% | 0 | 0 | — |
case-10 | fail→fail | 8,290 | 3,972 | -52% | 1 | 1 | 0% | 1,428 | 1,384 | -3% | 0 | 0 | — |
case-11 | pass→pass | 9,881 | 6,674 | -32% | 1 | 1 | 0% | 1,886 | 1,967 | +4% | 0 | 0 | — |
case-12 | pass→fail | 9,868 | 6,596 | -33% | 1 | 1 | 0% | 1,738 | 1,835 | +6% | 0 | 0 | — |
case-15 | pass→pass | 6,760 | 3,505 | -48% | 1 | 1 | 0% | 1,301 | 1,349 | +4% | 0 | 0 | — |
case-16 | pass→pass | 5,189 | 2,702 | -48% | 1 | 1 | 0% | 966 | 1,213 | +26% | 0 | 0 | — |
case-17 | pass→pass | 5,443 | 2,229 | -59% | 1 | 1 | 0% | 939 | 1,053 | +12% | 0 | 0 | — |
case-18 | pass→fail | 12,539 | 9,241 | -26% | 1 | 1 | 0% | 2,235 | 2,296 | +3% | 0 | 0 | — |
case-19 | pass→pass | 6,203 | 4,696 | -24% | 1 | 1 | 0% | 1,151 | 1,486 | +29% | 0 | 0 | — |
case-20 | pass→pass | 8,164 | 5,480 | -33% | 1 | 1 | 0% | 1,466 | 1,685 | +15% | 0 | 0 | — |
case-21 | pass→pass | 5,932 | 2,907 | -51% | 1 | 1 | 0% | 970 | 1,164 | +20% | 0 | 0 | — |
case-22 | fail→pass | 8,956 | 3,669 | -59% | 1 | 1 | 0% | 1,455 | 1,285 | -12% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -33 percentage points is the difference between those two pass rates over the 18 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.