Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when a Java change needs independent Behavior-Driven Development guidance from trusted behavior facts through concrete examples and observable scenarios, with focused clarification when behavior is pending or ambiguous. This should trigger for requests such as Apply BDD; Facilitate behavior examples; Discover scenarios with Given When Then; Review these examples for shared domain language; Complete BDD discovery as a self-contained interaction. Part of Plinth Toolkit
.claude/skills/jabrena-058-design-bdd/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -4% | 0% |
Guide Java Enterprise teams through Behavior-Driven Development as a collaborative discovery and design practice. This is an interactive SKILL.
What is covered in this Skill?
references/058-design-bdd.md as the complete Gherkin syntax referenceKeep BDD grounded in trusted facts, collaborative discovery, shared language, and externally observable behavior.
references/058-design-bdd.md before applying BDD guidancereferences/058-design-bdd.md as the complete runtime source for Gherkin syntax guidanceRead references/058-design-bdd.md, then confirm the maintainer-provided or maintainer-sanitized sources, actors, desired outcomes, business rules, shared terminology, conflicts, and unresolved questions. Ask focused follow-up questions when pending or ambiguous facts would materially affect the outcome.
Develop supported main, alternative, boundary, and error examples. Connect each example to a confirmed fact or rule and keep unsupported possibilities explicit as unresolved questions.
Express approved examples in shared domain language, using Given/When/Then where useful and observable outcomes instead of incidental implementation detail. Use only the bundled references/058-design-bdd.md for syntax when Gherkin is selected; do not access its external upstream source.
Report trusted facts, shared terminology, approved examples and scenarios, deferred examples, unresolved decisions, source conflicts, and remaining risks. Keep the outcome self-contained so the user may reuse it in a later, separately requested interaction.
For detailed guidance, examples, and constraints, see references/058-design-bdd.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-20 | pass→pass | 11,562 | 9,633 | -17% | 1 | 1 | 0% | 1,719 | 2,129 | +24% | 0 | 0 | — |
case-01 | fail→fail | 19,310 | 16,719 | -13% | 1 | 1 | 0% | 3,341 | 3,727 | +12% | 0 | 0 | — |
case-02 | fail→pass | 14,679 | 15,327 | +4% | 1 | 1 | 0% | 2,368 | 3,074 | +30% | 0 | 0 | — |
case-07 | pass→pass | 13,076 | 13,379 | +2% | 1 | 1 | 0% | 2,088 | 2,850 | +36% | 0 | 0 | — |
case-03 | fail→fail | 19,504 | 4,931 | -75% | 1 | 1 | 0% | 3,153 | 1,167 | -63% | 0 | 0 | — |
case-04 | pass→pass | 13,717 | 22,879 | +67% | 1 | 1 | 0% | 2,601 | 3,946 | +52% | 0 | 0 | — |
case-05 | pass→fail | 15,801 | 23,204 | +47% | 1 | 1 | 0% | 2,765 | 4,213 | +52% | 0 | 0 | — |
case-06 | pass→pass | 9,782 | 22,907 | +134% | 1 | 1 | 0% | 1,934 | 4,674 | +142% | 0 | 0 | — |
case-08 | pass→pass | 14,623 | 9,984 | -32% | 1 | 1 | 0% | 2,105 | 1,899 | -10% | 0 | 0 | — |
case-09 | fail→pass | 13,955 | 11,352 | -19% | 1 | 1 | 0% | 2,043 | 2,551 | +25% | 0 | 0 | — |
case-10 | fail→pass | 10,655 | 4,671 | -56% | 1 | 1 | 0% | 1,685 | 1,574 | -7% | 0 | 0 | — |
case-11 | pass→pass | 12,059 | 11,386 | -6% | 1 | 1 | 0% | 1,743 | 2,459 | +41% | 0 | 0 | — |
case-12 | pass→pass | 13,345 | 13,585 | +2% | 1 | 1 | 0% | 2,131 | 3,019 | +42% | 0 | 0 | — |
case-13 | pass→fail | 17,483 | 5,079 | -71% | 1 | 1 | 0% | 2,557 | 1,059 | -59% | 0 | 0 | — |
case-14 | fail→pass | 19,718 | 10,399 | -47% | 1 | 1 | 0% | 3,810 | 2,268 | -40% | 0 | 0 | — |
case-15 | fail→pass | 14,404 | 10,399 | -28% | 1 | 1 | 0% | 2,337 | 2,246 | -4% | 0 | 0 | — |
case-16 | pass→pass | 10,085 | 7,499 | -26% | 1 | 1 | 0% | 1,597 | 1,918 | +20% | 0 | 0 | — |
case-17 | pass→pass | 12,783 | 14,203 | +11% | 1 | 1 | 0% | 2,222 | 3,057 | +38% | 0 | 0 | — |
case-18 | pass→pass | 18,257 | 10,193 | -44% | 1 | 1 | 0% | 2,648 | 2,258 | -15% | 0 | 0 | — |
case-19 | fail→fail | 10,812 | 5,183 | -52% | 1 | 1 | 0% | 1,571 | 1,096 | -30% | 0 | 0 | — |
case-21 | fail→pass | 10,973 | 11,213 | +2% | 1 | 1 | 0% | 1,665 | 2,320 | +39% | 0 | 0 | — |
case-22 | fail→pass | 13,339 | 10,837 | -19% | 1 | 1 | 0% | 2,013 | 2,429 | +21% | 0 | 0 | — |
case-23 | pass→pass | 16,564 | 10,906 | -34% | 1 | 1 | 0% | 2,327 | 2,390 | +3% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +22 percentage points is the difference between those two pass rates over the 20 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.