Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Detect non-deterministic (flaky) tests by reading CI run logs or test result history. Aggregates pass rates per test, identifies intermittent failures, recommends quarantine or fix, and maintains a flaky test registry. Best run during Polish phase or after multiple CI runs.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 84% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 66% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 81% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 106% | 0% |
A flaky test is one that sometimes passes and sometimes fails without any code change. Flaky tests are worse than no tests in some ways — they train the team to ignore red CI runs, masking genuine failures. This skill identifies them, explains likely causes, and recommends whether to quarantine or fix each one.
Output: Updated tests/regression-suite.md quarantine section + optional production/qa/flakiness-report-[date].md
When to run:
/regression-suite identifies quarantined tests that need diagnosisModes:
/test-flakiness [ci-log-path] — analyse a specific CI run log file/test-flakiness scan — scan all available CI logs in .github/ orstandard log output directories
/test-flakiness registry — read existing regression-suite.md quarantinesection and provide remediation guidance for already-known flaky tests
scan if CI logs are accessible, elseregistry
Check for test result artifacts:
bashls -t .github/ 2>/dev/null ls -t test-results/ 2>/dev/null
For Godot projects: GdUnit4 outputs XML results compatible with JUnit format. Check test-results/ for .xml files.
For Unity projects: game-ci test runner outputs NUnit XML to test-results/ by default.
For Unreal projects: automation logs go to Saved/Logs/. Grep for Result: Success and Result: Fail patterns.
If a path argument is provided, read that file directly.
If no logs found: > "No CI log data found. To detect flaky tests, this skill needs test result > history from multiple runs. Options: > 1. Run the test suite at least 3 times and collect the output logs > 2. Check CI pipeline output and save a log to test-results/ > 3. Run /test-flakiness registry to review tests already flagged as flaky > in tests/regression-suite.md"
Stop and ask the user which option to pursue.
For each CI log or result file found, parse:
JUnit XML format (GdUnit4 / Unity):
<testcase name= to get test names<failure or <error to identify failuresclassname and name attributes for full test identifiersPlain text logs:
PASSED / FAILED adjacent to test namesResult: Success / Result: FailTest passed / Test failedBuild a table: test_id → [run1_result, run2_result, run3_result, ...]
A test is flaky if it appears in the result history with both PASS and FAIL outcomes across runs with no code changes between them.
Flakiness thresholds:
genuinely rare failure
For each flaky test, classify the likely cause:
| Cause | Symptoms | Fix direction | |-------|----------|---------------| | Timing / async | Fails after awaiting signals or timers; pass rate correlates with system load | Add explicit await/synchronisation; avoid time-based delays | | Order dependency | Fails when run after specific other tests; passes in isolation | Add proper setup/teardown; ensure test isolation | | Random seed | Fails intermittently with no pattern; involves RNG | Pass explicit seed; don't use randf() in tests | | Resource leak | Fails more often later in a test run | Fix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity) | | External state | Fails when a file, scene, or global exists from a prior test | Isolate test from file system; use in-memory mocks | | Floating point | Fails on comparisons like == 0.5 | Use epsilon comparison (is_equal_approx, Assert.AreApproximately) | | Scene/prefab load race | Fails when scenes are not yet ready | Await one frame after instantiation; use await get_tree().process_frame |
Use Grep to check the test file for timing calls, randf, global state access, or equality comparisons on floats to narrow down the cause.
For each flaky test:
Quarantine (High flakiness): > "Quarantine this test immediately. Disable it in CI by adding > @pytest.mark.skip / [Ignore] / GdUnitSkip annotation. Log it in > tests/regression-suite.md quarantine section. The test is now opt-in only. > Fix the root cause before removing quarantine."
Investigate and fix soon (Moderate): > "This test is intermittently unreliable. Root cause appears to be cause]. > Suggested fix: specific fix based on cause classification]. Do not quarantine > yet — fix the test directly."
Monitor (Low/suspected): > "This test shows suspected flakiness. Collect more run data before > quarantining. Note it as 'suspected' in the regression suite."
## Flakiness Detection Results
**Runs analysed**: [N]
**Tests tracked**: [N]
### Flaky Tests Found
| Test | System | Fail Rate | Likely Cause | Recommendation |
|------|--------|-----------|--------------|----------------|
| [test_name] | [system] | [N]% | Timing | Quarantine + fix async |
| [test_name] | [system] | [N]% | Float comparison | Fix: use epsilon compare |
| [test_name] | [system] | [N]% | Order dependency | Investigate teardown |
### Clean Tests (no flakiness detected)
[N] tests ran across [N] runs with consistent results — no flakiness detected.
### Data Limitations
[Note if fewer than 5 runs were available — fewer runs = less statistical confidence]Ask: "May I update the quarantine section of tests/regression-suite.md with the flaky tests found?"
If yes: use Edit to append entries to the Quarantined Tests table. Never remove existing quarantine entries — only add new ones.
Ask (separately): "May I write a full flakiness report to production/qa/flakiness-report-[date].md?"
The full report includes per-test analysis with cause details and engine-specific fix snippets.
After writing:
disable this test in CI. Re-enable after the root cause is fixed."
change the equality comparison on line N] to use is_equal_approx."
Schedule fix work for the N] quarantined tests before the release gate."
"suspected" not "confirmed"; ask if more run data is available
direction even when recommending quarantine
file require explicit approval. On write: Verdict: COMPLETE — flakiness report written. On decline: Verdict: BLOCKED — user declined write.
actions clearly; do not just silently quarantine without the team knowing
Other measured skills in the registry, with their headline benchmark lift.