Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run end-to-end tests via the CI workflow.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -5% | 0% |
| case-05 | ✓→✗ | ▼ Worse | 57% | 0% |
Run end-to-end tests via the CI workflow.
The user may pass $ARGUMENTS to filter to a specific test case by name (e.g., hello-world, phone-setup). If not provided, infer options from context (see below).
Experimental tests: Always include --experimental by default. The user can pass --no-experimental to exclude them.
Branch detection: Check what branch you're on via git branch --show-current. If you're on a branch other than main, automatically pass -b <branch-name> so the CI run tests against that branch. The user can override with -b <other-branch> explicitly.
Test case filter: Determine which test case to target:
phone-setup), use that as the filter.playwright/cases/phone-setup.md), automatically target that test case.Additional flags from $ARGUMENTS:
--xcode uses the agent-xcode (AXUIElement) runner-d or --detach triggers the run and exits without polling--no-experimental overrides the default and excludes experimental testsbashcd playwright && bun run scripts/agent-ci.ts <options>
Map the resolved options:
-t <case-name>--experimental-b <branch-name>--xcode-dExamples:
/e2e (on main, no context) → bun run scripts/agent-ci.ts --experimental/e2e (on branch feat/phone, after editing phone-setup.md) → bun run scripts/agent-ci.ts --experimental -b feat/phone -t phone-setup/e2e hello-world (on main) → bun run scripts/agent-ci.ts --experimental -t hello-world/e2e --detach (on branch fix/bug) → bun run scripts/agent-ci.ts --experimental -b fix/bug -d/e2e --no-experimental → bun run scripts/agent-ci.tsBefore running, briefly state the resolved options (branch, test case, experimental) so the user can see what was inferred.
Once the command finishes (or is dispatched in detach mode), summarize:
Other measured skills in the registry, with their headline benchmark lift.