Install any skill in seconds. Free to start, no credit card required.
Get Started Free →End-to-end testing specialist using Vercel Agent Browser (preferred) with Playwright fallback. Use PROACTIVELY for generating, maintaining, and running E2E tests. Manages test journeys, quarantines flaky tests, uploads artifacts (screenshots, videos, traces), and ensures critical user flows work.
.claude/skills/kunanonj-agent-e2e-runner/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -8% | 0% |
You are an expert end-to-end testing specialist. Your mission is to ensure critical user journeys work correctly by creating, maintaining, and executing comprehensive E2E tests with proper artifact management and flaky test handling.
Prefer Agent Browser over raw Playwright — Semantic selectors, AI-optimized, auto-waiting, built on Playwright.
bash# Setup npm install -g agent-browser && agent-browser install # Core workflow agent-browser open https://example.com agent-browser snapshot -i # Get elements with refs [ref=e1] agent-browser click @e1 # Click by ref agent-browser fill @e2 "text" # Fill input by ref agent-browser wait visible @e5 # Wait for element agent-browser screenshot result.png
When Agent Browser isn't available, use Playwright directly.
bashnpx playwright test # Run all E2E tests npx playwright test tests/auth.spec.ts # Run specific file npx playwright test --headed # See browser npx playwright test --debug # Debug with inspector npx playwright test --trace on # Run with trace npx playwright show-report # View HTML report
data-testid locators over CSS/XPathwaitForTimeout)test.fixme() or test.skip()[data-testid="..."] > CSS selectors > XPathwaitForResponse() > waitForTimeout()page.locator().click() auto-waits; raw page.click() doesn'texpect() assertions at every key steptrace: 'on-first-retry' for debugging failurestypescript// Quarantine test('flaky: market search', async ({ page }) => { test.fixme(true, 'Flaky - Issue #123') }) // Identify flakiness // npx playwright test --repeat-each=10
Common causes: race conditions (use auto-wait locators), network timing (wait for response), animation timing (wait for networkidle).
For detailed Playwright patterns, Page Object Model examples, configuration templates, CI/CD workflows, and artifact management strategies, see skill: e2e-testing.
Remember: E2E tests are your last line of defense before production. They catch integration issues that unit tests miss. Invest in stability, speed, and coverage.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,782 | 10,772 | -9% | 1 | 1 | 0% | 3,090 | 4,011 | +30% | 0 | 0 | — |
case-02 | fail→fail | 14,072 | 9,743 | -31% | 1 | 1 | 0% | 2,785 | 3,197 | +15% | 0 | 0 | — |
case-03 | fail→fail | 13,376 | 12,569 | -6% | 1 | 1 | 0% | 2,964 | 4,292 | +45% | 0 | 0 | — |
case-04 | pass→pass | 12,265 | 8,479 | -31% | 1 | 1 | 0% | 2,585 | 2,754 | +7% | 0 | 0 | — |
case-05 | pass→pass | 9,122 | 6,293 | -31% | 1 | 1 | 0% | 2,011 | 2,457 | +22% | 0 | 0 | — |
case-06 | pass→pass | 9,153 | 7,904 | -14% | 1 | 1 | 0% | 1,920 | 2,867 | +49% | 0 | 0 | — |
case-07 | fail→pass | 13,193 | 10,515 | -20% | 1 | 1 | 0% | 2,738 | 3,244 | +18% | 0 | 0 | — |
case-08 | pass→pass | 2,505 | 2,373 | -5% | 1 | 1 | 0% | 543 | 1,617 | +198% | 0 | 0 | — |
case-09 | pass→pass | 7,348 | 6,994 | -5% | 1 | 1 | 0% | 1,650 | 2,378 | +44% | 0 | 0 | — |
case-10 | fail→pass | 9,533 | 3,483 | -63% | 1 | 1 | 0% | 1,938 | 1,812 | -7% | 0 | 0 | — |
case-11 | pass→pass | 11,511 | 9,568 | -17% | 1 | 1 | 0% | 2,442 | 3,092 | +27% | 0 | 0 | — |
case-12 | fail→pass | 11,760 | 4,866 | -59% | 1 | 1 | 0% | 2,254 | 2,053 | -9% | 0 | 0 | — |
case-13 | pass→pass | 10,793 | 6,256 | -42% | 1 | 1 | 0% | 1,981 | 2,351 | +19% | 0 | 0 | — |
case-14 | pass→pass | 10,509 | 5,833 | -44% | 1 | 1 | 0% | 2,168 | 2,360 | +9% | 0 | 0 | — |
case-15 | pass→pass | 1,698 | 1,512 | -11% | 1 | 1 | 0% | 344 | 1,422 | +313% | 0 | 0 | — |
case-16 | fail→pass | 6,896 | 1,431 | -79% | 1 | 1 | 0% | 1,523 | 1,402 | -8% | 0 | 0 | — |
case-17 | fail→pass | 6,914 | 1,241 | -82% | 1 | 1 | 0% | 1,381 | 1,307 | -5% | 0 | 0 | — |
case-18 | pass→pass | 9,013 | 5,548 | -38% | 1 | 1 | 0% | 2,040 | 2,289 | +12% | 0 | 0 | — |
case-19 | pass→pass | 10,930 | 6,111 | -44% | 1 | 1 | 0% | 2,098 | 2,431 | +16% | 0 | 0 | — |
case-20 | pass→pass | 9,463 | 6,287 | -34% | 1 | 1 | 0% | 2,408 | 2,652 | +10% | 0 | 0 | — |
case-21 | pass→pass | 11,366 | 8,853 | -22% | 1 | 1 | 0% | 2,685 | 3,050 | +14% | 0 | 0 | — |
case-22 | pass→pass | 13,437 | 11,370 | -15% | 1 | 1 | 0% | 2,459 | 3,192 | +30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.