Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate automated API and end-to-end tests for implemented features. Use when the user says "create qa automated tests for [feature]"
.claude/skills/bmad-code-org-bmad-qa-generate-e2e-tests/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 265% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -65% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -6% | 0% |
Goal: Generate automated API and E2E tests for implemented code.
Your Role: You are a QA automation engineer. You generate tests ONLY — no code review or story validation (use the bmad-code-review skill for that).
checklist.md) resolve from the skill root.{skill-root} resolves to this skill's installed directory (where customize.toml lives).{project-root}-prefixed paths resolve from the project working directory.{skill-name} resolves to the skill directory's basename.Run: uv run {project-root}/_bmad/scripts/resolve_customization.py --skill {skill-root} --project-root {project-root} --key workflow
If the script is not found, BMad is not set up here. Offer to run the bmad skill's setup, installing bmad first if you do not have it (npx skills add bmad-code-org/BMAD-METHOD --skill bmad), then run the command again.
If it fails for any other reason, resolve the workflow block yourself by reading these three files in base → team → user order and applying the same structural merge rules as the resolver:
{skill-root}/customize.toml — defaults{project-root}/_bmad/custom/{skill-name}.toml — team overrides{project-root}/_bmad/custom/{skill-name}.user.toml — personal overridesAny missing file is skipped. Scalars override, tables deep-merge, arrays of tables keyed by code or id replace matching entries and append new entries, and all other arrays append.
Execute each entry in {workflow.activation_steps_prepend} in order before proceeding.
Treat every entry in {workflow.persistent_facts} as foundational context you carry for the rest of the workflow run. Entries prefixed file: are paths or globs under {project-root} — load the referenced contents as facts. All other entries are facts verbatim.
Run: uv run {project-root}/_bmad/scripts/resolve_config.py --project-root {project-root} --key core.project_name --key modules.bmm.implementation_artifacts
date as system-generated current datetimeGreet the user.
Execute each entry in {workflow.activation_steps_append} in order.
Activation is complete. If activation_steps_prepend or activation_steps_append were non-empty, confirm every entry was executed in order before proceeding. Do not begin the main workflow until all activation steps have been completed.
test_dir = {project-root}/testssource_dir = {project-root}default_output_file = {implementation_artifacts}/tests/test-summary.mdCheck project for existing test framework:
package.json dependencies (playwright, jest, vitest, cypress, etc.)Ask user what to test:
src/components/)For API endpoints/services, generate tests that:
For UI features, generate tests that:
Execute tests to verify they pass (use project's test command).
If failures occur, fix them immediately.
Output markdown summary:
markdown# Test Automation Summary ## Generated Tests ### API Tests - [x] tests/api/endpoint.spec.ts - Endpoint validation ### E2E Tests - [x] tests/e2e/feature.spec.ts - User workflow ## Coverage - API endpoints: 5/10 covered - UI features: 3/8 covered ## Next Steps - Run tests in CI - Add more edge cases as needed
Do:
Avoid:
For Advanced Features:
If the project needs:
> Install Test Architect (TEA) module: <https://bmad-code-org.github.io/bmad-method-test-architecture-enterprise/>
Save summary to: {default_output_file}
Done! Tests generated and verified. Validate against checklist.md.
Run: uv run {project-root}/_bmad/scripts/resolve_customization.py --skill {skill-root} --project-root {project-root} --key workflow.on_complete
If the resolved workflow.on_complete is non-empty, follow it as the final terminal instruction before exiting.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 31,890 | 13,578 | -57% | 1 | 1 | 0% | 168 | 1,836 | +993% | 0 | 0 | — |
case-02 | fail→fail | 10,476 | 26,840 | +156% | 1 | 1 | 0% | 275 | 1,707 | +521% | 0 | 0 | — |
case-03 | fail→fail | 5,519 | 9,104 | +65% | 1 | 1 | 0% | 205 | 1,854 | +804% | 0 | 0 | — |
case-04 | fail→pass | 13,056 | 40,359 | +209% | 1 | 1 | 0% | 1,095 | 3,993 | +265% | 0 | 0 | — |
case-05 | fail→pass | 39,434 | 4,688 | -88% | 1 | 1 | 0% | 6,553 | 2,269 | -65% | 0 | 0 | — |
case-06 | fail→fail | 6,154 | 8,909 | +45% | 1 | 1 | 0% | 995 | 2,027 | +104% | 0 | 0 | — |
case-07 | fail→pass | 9,543 | 3,843 | -60% | 1 | 1 | 0% | 1,436 | 2,095 | +46% | 0 | 0 | — |
case-08 | pass→pass | 7,069 | 5,541 | -22% | 1 | 1 | 0% | 1,190 | 2,382 | +100% | 0 | 0 | — |
case-09 | pass→pass | 14,280 | 3,540 | -75% | 1 | 1 | 0% | 2,260 | 1,929 | -15% | 0 | 0 | — |
case-10 | fail→pass | 11,200 | 4,073 | -64% | 1 | 1 | 0% | 1,706 | 1,961 | +15% | 0 | 0 | — |
case-11 | fail→fail | 10,721 | 2,930 | -73% | 1 | 1 | 0% | 1,459 | 1,818 | +25% | 0 | 0 | — |
case-12 | pass→pass | 10,144 | 8,241 | -19% | 1 | 1 | 0% | 1,608 | 2,698 | +68% | 0 | 0 | — |
case-13 | fail→fail | 13,480 | 7,256 | -46% | 1 | 1 | 0% | 1,775 | 2,496 | +41% | 0 | 0 | — |
case-14 | fail→fail | 15,997 | 12,535 | -22% | 1 | 1 | 0% | 2,582 | 3,286 | +27% | 0 | 0 | — |
case-15 | fail→pass | 37,437 | 6,422 | -83% | 1 | 1 | 0% | 2,308 | 2,161 | -6% | 0 | 0 | — |
case-16 | pass→pass | 14,610 | 7,748 | -47% | 1 | 1 | 0% | 2,156 | 2,639 | +22% | 0 | 0 | — |
case-17 | pass→fail | 19,861 | 7,537 | -62% | 1 | 1 | 0% | 2,219 | 2,493 | +12% | 0 | 0 | — |
case-18 | fail→pass | 10,286 | 4,169 | -59% | 1 | 1 | 0% | 1,464 | 1,971 | +35% | 0 | 0 | — |
case-19 | fail→fail | 21,171 | 19,434 | -8% | 1 | 1 | 0% | 1,208 | 5,505 | +356% | 0 | 0 | — |
case-20 | fail→pass | 9,461 | 3,960 | -58% | 1 | 1 | 0% | 1,116 | 1,872 | +68% | 0 | 0 | — |
case-21 | pass→fail | 11,441 | 4,011 | -65% | 1 | 1 | 0% | 1,621 | 2,010 | +24% | 0 | 0 | — |
case-22 | fail→pass | 7,277 | 4,269 | -41% | 1 | 1 | 0% | 1,023 | 1,992 | +95% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.