Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build features with tests-before-code rigor — use for new features needing test coverage
.claude/skills/nyldn-skill-tdd/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-11 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -41% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -6% | 0% |
Read skills/blocks/engineering-method-selection.md from the installed plugin for review admission. Natural-language requests and --peer-review share that policy. Honor host-only requests; risk alone does not authorize paid usage.
Run the red, green, and refactor cycle on the current host. Routine TDD makes zero additional provider dispatches. Use one external reviewer only when the user passes --peer-review, explicitly requests independent review, or an existing risk policy requires it. Explicit debate, council, and multi-model commands retain their own execution contracts.
<HARD-GATE> NO PRODUCTION BEHAVIOR CHANGE WITHOUT AN OBSERVED, EXPECTED FAILING TEST FIRST. </HARD-GATE>
Do not change production behavior until a focused test fails for the expected reason. Existing implementation outside the requested change remains intact.
public boundary cannot isolate its failure mode.
If a test passes before implementation, it is not red evidence. If it errors due to fixture or syntax problems, repair the test until it fails on the missing behavior.
Do not equate similar assertions with duplicate guarantees. Keep separate OS, security, cancellation, and integration boundaries. For every removed test, record old_test, behavior, replacement, mutant, red_observed, baseline_ms, candidate_ms, and reason. The retained test must kill the named mutant at the intended caller boundary.
After one warm-up, measure five isolated runs and report every sample and the median. Review a slowdown only when it exceeds both 20 percent and 100 ms.
Completion requires observed red and green evidence, affected-suite results, and the consolidation ledger when tests were removed. A missing reviewer is reported as incomplete review, never simulated.
If the same test remains red after two implementation attempts, stop and recheck the test boundary, fixture, and expected behavior. The strategy-rotation hook is a signal to try a fundamentally different hypothesis, not another variation of the same patch.
Adapted from DEEPENING in mattpocock/skills at commit 3cca18b368ae95cdbdebbff572ccafa662551015 under the MIT License. See THIRD_PARTY_NOTICES.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 33,224 | 13,208 | -60% | 1 | 1 | 0% | 5,505 | 786 | -86% | 0 | 0 | — |
case-11 | fail→pass | 15,822 | 7,104 | -55% | 1 | 1 | 0% | 1,968 | 947 | -52% | 0 | 0 | — |
case-02 | fail→fail | 29,000 | 15,268 | -47% | 1 | 1 | 0% | 4,427 | 846 | -81% | 0 | 0 | — |
case-03 | fail→fail | 29,989 | 16,216 | -46% | 1 | 1 | 0% | 4,741 | 951 | -80% | 0 | 0 | — |
case-04 | pass→pass | 13,611 | 10,452 | -23% | 1 | 1 | 0% | 1,321 | 1,404 | +6% | 0 | 0 | — |
case-05 | fail→pass | 16,222 | 12,400 | -24% | 1 | 1 | 0% | 1,791 | 1,909 | +7% | 0 | 0 | — |
case-06 | pass→pass | 14,851 | 10,497 | -29% | 1 | 1 | 0% | 1,638 | 1,572 | -4% | 0 | 0 | — |
case-07 | fail→pass | 17,679 | 11,659 | -34% | 1 | 1 | 0% | 1,932 | 1,541 | -20% | 0 | 0 | — |
case-08 | pass→pass | 17,911 | 17,688 | -1% | 1 | 1 | 0% | 2,134 | 2,735 | +28% | 0 | 0 | — |
case-09 | pass→pass | 17,202 | 8,831 | -49% | 1 | 1 | 0% | 1,871 | 1,176 | -37% | 0 | 0 | — |
case-10 | pass→pass | 9,120 | 8,462 | -7% | 1 | 1 | 0% | 674 | 1,273 | +89% | 0 | 0 | — |
case-12 | fail→pass | 16,514 | 8,281 | -50% | 1 | 1 | 0% | 1,842 | 1,085 | -41% | 0 | 0 | — |
case-13 | pass→pass | 12,959 | 11,440 | -12% | 1 | 1 | 0% | 1,393 | 1,515 | +9% | 0 | 0 | — |
case-14 | pass→pass | 15,131 | 9,543 | -37% | 1 | 1 | 0% | 1,575 | 1,400 | -11% | 0 | 0 | — |
case-15 | pass→pass | 15,424 | 11,254 | -27% | 1 | 1 | 0% | 1,632 | 1,590 | -3% | 0 | 0 | — |
case-16 | fail→pass | 19,130 | 13,717 | -28% | 1 | 1 | 0% | 2,189 | 2,052 | -6% | 0 | 0 | — |
case-17 | pass→fail | 16,951 | 28,337 | +67% | 1 | 1 | 0% | 1,851 | 3,920 | +112% | 0 | 0 | — |
case-18 | pass→pass | 15,034 | 10,299 | -31% | 1 | 1 | 0% | 1,727 | 1,387 | -20% | 0 | 0 | — |
case-19 | fail→pass | 15,736 | 11,495 | -27% | 1 | 1 | 0% | 1,708 | 1,644 | -4% | 0 | 0 | — |
case-20 | pass→pass | 15,825 | 10,565 | -33% | 1 | 1 | 0% | 1,775 | 1,552 | -13% | 0 | 0 | — |
case-21 | pass→pass | 14,133 | 9,539 | -33% | 1 | 1 | 0% | 1,406 | 1,374 | -2% | 0 | 0 | — |
case-22 | pass→pass | 19,268 | 8,768 | -54% | 1 | 1 | 0% | 1,967 | 1,264 | -36% | 0 | 0 | — |
case-23 | pass→fail | 16,096 | 21,360 | +33% | 1 | 1 | 0% | 1,823 | 2,011 | +10% | 0 | 0 | — |
case-24 | pass→pass | 13,248 | 33,447 | +152% | 1 | 1 | 0% | 2,732 | 5,180 | +90% | 0 | 0 | — |
case-25 | pass→pass | 21,324 | 36,605 | +72% | 1 | 1 | 0% | 3,148 | 5,343 | +70% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 20 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +16 percentage points is the difference between those two pass rates over the 20 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/17/2026 | +22% |
Other measured skills in the registry, with their headline benchmark lift.