Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Test-Driven Development specialist enforcing write-tests-first methodology. Use PROACTIVELY when writing new features, fixing bugs, or refactoring code. Ensures 80%+ test coverage.
.claude/skills/kunanonj-agent-tdd-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 88% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 28% | 0% |
You are a Test-Driven Development (TDD) specialist who ensures all code is developed test-first with comprehensive coverage.
Write a failing test that describes the expected behavior.
bashnpm test
Only enough code to make the test pass.
Remove duplication, improve names, optimize -- tests must stay green.
bashnpm run test:coverage # Required: 80%+ branches, functions, lines, statements
| Type | What to Test | When | |------|-------------|------| | Unit | Individual functions in isolation | Always | | Integration | API endpoints, database operations | Always | | E2E | Critical user flows (Playwright) | Critical paths |
For detailed mocking patterns and framework-specific examples, see skill: tdd-workflow.
Integrate eval-driven development into TDD flow:
Release-critical paths should target pass^3 stability before merge.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 18,395 | 14,034 | -24% | 1 | 1 | 0% | 2,751 | 3,517 | +28% | 0 | 0 | — |
case-02 | fail→pass | 21,570 | 20,321 | -6% | 1 | 1 | 0% | 3,576 | 5,149 | +44% | 0 | 0 | — |
case-03 | pass→pass | 14,526 | 16,580 | +14% | 1 | 1 | 0% | 2,501 | 3,715 | +49% | 0 | 0 | — |
case-01 | fail→fail | 23,469 | 24,132 | +3% | 1 | 1 | 0% | 4,647 | 5,753 | +24% | 0 | 0 | — |
case-05 | pass→pass | 8,192 | 10,339 | +26% | 1 | 1 | 0% | 1,694 | 2,657 | +57% | 0 | 0 | — |
case-06 | pass→pass | 10,786 | 8,362 | -22% | 1 | 1 | 0% | 2,095 | 2,206 | +5% | 0 | 0 | — |
case-07 | pass→pass | 15,620 | 13,913 | -11% | 1 | 1 | 0% | 2,907 | 3,236 | +11% | 0 | 0 | — |
case-08 | pass→pass | 12,764 | 11,411 | -11% | 1 | 1 | 0% | 2,405 | 2,739 | +14% | 0 | 0 | — |
case-09 | pass→pass | 16,059 | 16,940 | +5% | 1 | 1 | 0% | 2,854 | 4,115 | +44% | 0 | 0 | — |
case-10 | pass→pass | 13,457 | 11,086 | -18% | 1 | 1 | 0% | 2,194 | 3,066 | +40% | 0 | 0 | — |
case-11 | pass→pass | 9,396 | 11,802 | +26% | 1 | 1 | 0% | 1,977 | 3,291 | +66% | 0 | 0 | — |
case-12 | pass→pass | 10,108 | 3,317 | -67% | 1 | 1 | 0% | 1,461 | 1,450 | -1% | 0 | 0 | — |
case-13 | pass→pass | 6,857 | 6,000 | -12% | 1 | 1 | 0% | 1,080 | 1,970 | +82% | 0 | 0 | — |
case-14 | pass→pass | 16,484 | 11,879 | -28% | 1 | 1 | 0% | 2,738 | 3,366 | +23% | 0 | 0 | — |
case-15 | pass→pass | 6,127 | 6,306 | +3% | 1 | 1 | 0% | 1,094 | 1,829 | +67% | 0 | 0 | — |
case-16 | pass→pass | 3,215 | 3,674 | +14% | 1 | 1 | 0% | 674 | 1,579 | +134% | 0 | 0 | — |
case-17 | fail→pass | 4,396 | 2,527 | -43% | 1 | 1 | 0% | 702 | 1,322 | +88% | 0 | 0 | — |
case-18 | fail→pass | 6,166 | 2,997 | -51% | 1 | 1 | 0% | 1,146 | 1,312 | +14% | 0 | 0 | — |
case-19 | fail→pass | 17,220 | 2,849 | -83% | 1 | 1 | 0% | 1,735 | 1,301 | -25% | 0 | 0 | — |
case-20 | pass→pass | 8,322 | 4,028 | -52% | 1 | 1 | 0% | 1,388 | 1,480 | +7% | 0 | 0 | — |
case-21 | pass→pass | 7,139 | 7,296 | +2% | 1 | 1 | 0% | 1,235 | 2,260 | +83% | 0 | 0 | — |
case-22 | pass→pass | 2,771 | 3,446 | +24% | 1 | 1 | 0% | 419 | 1,368 | +226% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.