Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Fix failing or flaky Playwright tests. Use when user says "fix test", "flaky test", "test failing", "debug test", "test broken", "test passes sometimes", or "intermittent failure".
.claude/skills/alirezarezvani-fix/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 103% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 101% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 82% | 0% |
Diagnose and fix a Playwright test that fails or passes intermittently using a systematic taxonomy.
$ARGUMENTS contains:
e2e/login.spec.ts"the checkout test fails in CI but passes locally"Run the test to capture the error:
bashnpx playwright test <file> --reporter=list
If the test passes, it's likely flaky. Run burn-in:
bashnpx playwright test <file> --repeat-each=10 --reporter=list
If it still passes, try with parallel workers:
bashnpx playwright test --fully-parallel --workers=4 --repeat-each=5
Run with full tracing:
bashnpx playwright test <file> --trace=on --retries=0
Read the trace output. Use /debug to analyze trace files if available.
Load flaky-taxonomy.md from this skill directory.
Every failing test falls into one of four categories:
| Category | Symptom | Diagnosis | |---|---|---| | Timing/Async | Fails intermittently everywhere | --repeat-each=20 reproduces locally | | Test Isolation | Fails in suite, passes alone | --workers=1 --grep "test name" passes | | Environment | Fails in CI, passes locally | Compare CI vs local screenshots/traces | | Infrastructure | Random, no pattern | Error references browser internals |
Timing/Async:
waitForTimeout() with web-first assertionsawait to missing Playwright callstoBeVisible() before interacting with elementsTest Isolation:
Environment:
docker locally to match CI environmentInfrastructure:
retries: 2)Run the test 10 times to confirm stability:
bashnpx playwright test <file> --repeat-each=10 --reporter=list
All 10 must pass. If any fail, go back to step 3.
Suggest:
retries: 2 if not alreadytrace: 'on-first-retry' in config| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,916 | 3,975 | +2% | 1 | 1 | 0% | 169 | 960 | +468% | 0 | 0 | — |
case-02 | fail→fail | 4,007 | 4,770 | +19% | 1 | 1 | 0% | 230 | 985 | +328% | 0 | 0 | — |
case-03 | fail→fail | 12,796 | 4,986 | -61% | 1 | 1 | 0% | 2,716 | 1,015 | -63% | 0 | 0 | — |
case-04 | pass→fail | 9,299 | 4,501 | -52% | 1 | 1 | 0% | 1,939 | 944 | -51% | 0 | 0 | — |
case-05 | pass→fail | 5,644 | 4,674 | -17% | 1 | 1 | 0% | 1,138 | 1,004 | -12% | 0 | 0 | — |
case-06 | pass→pass | 9,193 | 12,299 | +34% | 1 | 1 | 0% | 2,120 | 3,222 | +52% | 0 | 0 | — |
case-07 | fail→pass | 3,639 | 3,388 | -7% | 1 | 1 | 0% | 706 | 1,436 | +103% | 0 | 0 | — |
case-08 | fail→fail | 6,325 | 4,818 | -24% | 1 | 1 | 0% | 1,252 | 1,614 | +29% | 0 | 0 | — |
case-09 | fail→pass | 3,282 | 2,733 | -17% | 1 | 1 | 0% | 613 | 1,233 | +101% | 0 | 0 | — |
case-10 | pass→pass | 6,679 | 4,337 | -35% | 1 | 1 | 0% | 1,106 | 1,453 | +31% | 0 | 0 | — |
case-11 | fail→pass | 5,955 | 3,469 | -42% | 1 | 1 | 0% | 1,100 | 1,354 | +23% | 0 | 0 | — |
case-12 | pass→pass | 6,937 | 3,861 | -44% | 1 | 1 | 0% | 1,151 | 1,441 | +25% | 0 | 0 | — |
case-13 | pass→pass | 7,263 | 4,646 | -36% | 1 | 1 | 0% | 1,282 | 1,537 | +20% | 0 | 0 | — |
case-14 | fail→pass | 7,154 | 1,678 | -77% | 1 | 1 | 0% | 1,179 | 1,066 | -10% | 0 | 0 | — |
case-15 | pass→pass | 8,898 | 3,987 | -55% | 1 | 1 | 0% | 1,599 | 1,359 | -15% | 0 | 0 | — |
case-16 | pass→pass | 9,717 | 8,269 | -15% | 1 | 1 | 0% | 1,849 | 2,288 | +24% | 0 | 0 | — |
case-17 | pass→pass | 10,156 | 7,886 | -22% | 1 | 1 | 0% | 1,894 | 2,298 | +21% | 0 | 0 | — |
case-18 | pass→pass | 9,202 | 5,619 | -39% | 1 | 1 | 0% | 1,753 | 1,877 | +7% | 0 | 0 | — |
case-19 | pass→pass | 5,545 | 2,429 | -56% | 1 | 1 | 0% | 906 | 1,172 | +29% | 0 | 0 | — |
case-20 | fail→pass | 3,892 | 3,151 | -19% | 1 | 1 | 0% | 659 | 1,199 | +82% | 0 | 0 | — |
case-21 | fail→fail | 5,620 | 3,205 | -43% | 1 | 1 | 0% | 1,021 | 1,209 | +18% | 0 | 0 | — |
case-22 | fail→pass | 9,493 | 1,112 | -88% | 1 | 1 | 0% | 1,616 | 912 | -44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 17 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.