Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Behavior-driven development: Discovery → Formulation → Automation. Use when the user wants to build features or fix bugs via Gherkin scenarios and the BDD lifecycle, or mentions BDD/Gherkin/Executable Specifications. Delegates test implementation to /mattpocock:tdd.
.claude/skills/fradser-bdd/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 112% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 45% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 131% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 101% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 257% | 0% |
BDD is not just about tools; it's a methodology for shared understanding and high-quality implementation. This skill governs the Discovery → Formulation → Automation lifecycle: Gherkin scenarios define the behavior, and the Automation phase (red-green loop) is delegated to /mattpocock:tdd (BDD-driven).
When exploring the codebase, read CONTEXT.md (if it exists) so test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching.
When the user asks for a feature, bug fix, or refactor, apply the following mindset:
The process flows from requirements to code:
/mattpocock:tdd, which governs test quality, seams, mocking, and anti-patterns.See ./references/bdd-best-practices.md for a detailed guide.
Scenarios are your "Executable Specifications".
.feature files or the framework-native executable test format — not as code comments. Once implementation begins, translate the scenarios that will be automated into .feature files or the framework-native executable test format, so they double as living documentation.See ./references/gherkin-guide.md for syntax and storage structure.
The Automation phase (red-green loop) is governed by the /mattpocock:tdd skill (BDD-driven). When the Gherkin scenario is defined and the seam is agreed, invoke /mattpocock:tdd for:
Seams are the public boundaries you test at. See /mattpocock:tdd for the full seam guidance.
Test only at pre-agreed seams. Before writing any test, write down the seams under test and confirm them with the user via the AskUserQuestion tool. No test is written at an unconfirmed seam. Testing everything isn't possible — agreeing the seams up front is how testing effort lands on the critical paths and complex logic instead of every edge case.
Use the AskUserQuestion tool to ask which seams to test, proposing the candidate seams as options (the user can adjust via "Other").
Tests verify behavior through public interfaces, not implementation details. See /mattpocock:tdd for the full guidance, examples, and anti-patterns.
> "No production code is written without a failing test first."
The Red step MUST verify the test fails for the right reason (run the test and read the failure output) before writing any implementation. Skipping or rationalizing this step produces:
Delete it and re-derive it from a failing test — do not keep it "as reference," do not "adapt" it into the test-first version, do not read it while writing the test. Any of those re-introduces the implementation-biased-test failure mode above through the back door: a test written while looking at the code it's meant to constrain will pass on the first try regardless of whether it checks the right thing. Delete means delete.
| Rationalization | Why it fails | |---|---| | "I'll write the test after — same coverage either way" | A test written against working code always passes on the first run. That proves the test doesn't crash, not that it verifies the right behavior. Only a test that failed first, for the stated reason, has been shown capable of catching a regression. | | "I already manually verified it works" | Manual verification is not repeatable and leaves no regression guard. It answers "did this work once," not "will this keep working." | | "This is too simple to need a test" | Simple code changes behavior just as easily as complex code. The Iron Law has no complexity threshold — it has the three named exceptions below and nothing else. | | "I'll be pragmatic, not dogmatic, about BDD-driven TDD" | This is the rationalization, not an alternative to it. Every one of these tables' entries is someone being "pragmatic" about skipping the Red step. | | "I already spent an hour on this, deleting it is wasteful" | Sunk cost. The hour is already spent whether you delete the code or keep it; keeping untested code doesn't recover that hour, it just adds an unverified regression risk on top of it. |
The only legitimate exceptions are named in ./references/bdd-best-practices.md (one-off prototypes, generated code, config files) — and even those should be raised with the user, not silently assumed.
A test-first test encodes "this is what the system is contracted to do." A test-after test encodes "this is what the code I already wrote happens to do" — it will pass even if the code has the wrong behavior, because it was shaped to match that behavior rather than an independent specification. If you catch yourself writing a test against code you can already see, stop, delete the code, and write the test against the behavior instead.
See /mattpocock:tdd (BDD-driven) for the full red-green loop rules, anti-patterns, and test quality guidance. The key points:
/mattpocock:code-review), not the red → green implementation cycle../references/bdd-best-practices.md - BDD methodology: discovery, formulation, automation./references/gherkin-guide.md - Gherkin syntax, storage structure, examples./references/testing-anti-patterns.md - Mocking pitfalls and vacuous-passing tests/mattpocock:tdd - BDD-driven test implementation: seams, mocking, test quality, red-green loop| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 18,882 | 20,125 | +7% | 1 | 1 | 0% | 2,897 | 5,352 | +85% | 0 | 0 | — |
case-01 | fail→fail | 30,965 | 10,043 | -68% | 1 | 1 | 0% | 6,489 | 1,938 | -70% | 0 | 0 | — |
case-02 | fail→fail | 35,138 | 5,807 | -83% | 1 | 1 | 0% | 6,509 | 1,941 | -70% | 0 | 0 | — |
case-03 | fail→fail | 8,517 | 4,704 | -45% | 1 | 1 | 0% | 1,434 | 1,976 | +38% | 0 | 0 | — |
case-04 | fail→pass | 9,809 | 15,728 | +60% | 1 | 1 | 0% | 1,586 | 3,360 | +112% | 0 | 0 | — |
case-05 | fail→pass | 13,445 | 13,626 | +1% | 1 | 1 | 0% | 2,729 | 3,954 | +45% | 0 | 0 | — |
case-06 | fail→pass | 9,904 | 17,537 | +77% | 1 | 1 | 0% | 1,586 | 3,671 | +131% | 0 | 0 | — |
case-07 | pass→pass | 11,423 | 4,770 | -58% | 1 | 1 | 0% | 1,559 | 2,417 | +55% | 0 | 0 | — |
case-08 | fail→fail | 6,952 | 14,587 | +110% | 1 | 1 | 0% | 1,059 | 4,206 | +297% | 0 | 0 | — |
case-09 | fail→fail | 11,411 | 8,618 | -24% | 1 | 1 | 0% | 1,740 | 2,936 | +69% | 0 | 0 | — |
case-10 | pass→pass | 10,499 | 6,151 | -41% | 1 | 1 | 0% | 1,628 | 2,725 | +67% | 0 | 0 | — |
case-11 | pass→pass | 13,074 | 7,454 | -43% | 1 | 1 | 0% | 2,158 | 2,897 | +34% | 0 | 0 | — |
case-12 | fail→pass | 10,088 | 9,402 | -7% | 1 | 1 | 0% | 1,628 | 3,272 | +101% | 0 | 0 | — |
case-13 | fail→pass | 4,445 | 8,758 | +97% | 1 | 1 | 0% | 886 | 3,166 | +257% | 0 | 0 | — |
case-14 | pass→pass | 12,409 | 9,574 | -23% | 1 | 1 | 0% | 1,788 | 3,486 | +95% | 0 | 0 | — |
case-15 | pass→pass | 15,200 | 10,563 | -31% | 1 | 1 | 0% | 2,667 | 3,326 | +25% | 0 | 0 | — |
case-16 | pass→pass | 9,712 | 5,130 | -47% | 1 | 1 | 0% | 1,808 | 2,561 | +42% | 0 | 0 | — |
case-17 | fail→pass | 12,391 | 5,832 | -53% | 1 | 1 | 0% | 2,129 | 2,550 | +20% | 0 | 0 | — |
case-18 | pass→pass | 10,365 | 6,507 | -37% | 1 | 1 | 0% | 1,787 | 2,754 | +54% | 0 | 0 | — |
case-19 | pass→pass | 10,561 | 7,991 | -24% | 1 | 1 | 0% | 1,961 | 3,206 | +63% | 0 | 0 | — |
case-20 | pass→pass | 16,006 | 9,569 | -40% | 1 | 1 | 0% | 2,850 | 3,345 | +17% | 0 | 0 | — |
case-21 | fail→fail | 19,111 | 16,899 | -12% | 1 | 1 | 0% | 3,365 | 4,122 | +22% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.