Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Guides testing Warp UI features and changes using the computer use tool. Use this skill only when computer-use testing was requested (explicit request or accepted offer) and the computer_use tool is available to the agent. Covers launching Warp and verifying UI behavior.
.claude/skills/warpdotdev-test-warp-ui/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 6% | 0% |
Use the computer_use tool to visually test that Warp looks and behaves as intended after UI changes, when computer-use testing was requested.
Launch Warp from the repository root. The exact command depends on which environment variable holds the API key:
WARP_API_KEY is already set, omit the flag entirely — the --api-key flag is bound to WARP_API_KEY, so Warp reads it automatically:bash cargo run --bin warp
STAGING_USER_WARP_API_KEY instead, pass it explicitly via the flag:bash cargo run --bin warp -- --api-key $STAGING_USER_WARP_API_KEY
Always pass --bin warp explicitly. That target builds the internal (dogfood) channel, which is the only channel that honors --api-key for the GUI app. A plain cargo run builds the OSS channel, which ignores the key and falls back to interactive onboarding.
Authenticating this way starts the app directly without interactive login prompts.
Initial builds may take several minutes; subsequent incremental builds are faster.
After launching, confirm both of the following before testing:
cargo run stderr/terminal output does not contain the substring provided but IGNORED.If that warning appears (or the app is logged out), the wrong binary/channel was launched — stop and relaunch with cargo run --bin warp.
If you just need to verify that a specific UI looks correct, it can be useful to hardcode or mock data so the UI state is immediately reachable without navigating a full flow. This is optional — skip this step when testing end-to-end flows that should work naturally.
Examples of when to hardcode:
Keep mocked changes minimal and focused — only change what's necessary to reach the UI state under test.
Call the computer_use tool with a task description that includes:
cargo run --bin warp when WARP_API_KEY is set in the environment, or cargo run --bin warp -- --api-key $STAGING_USER_WARP_API_KEY when the key is in STAGING_USER_WARP_API_KEY insteadCompare the observations returned by computer_use against your expectations. If the UI doesn't match, investigate and adjust the code or mocks accordingly.
computer_use. The tool cannot fix build errors.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,578 | 11,614 | +154% | 1 | 1 | 0% | 145 | 1,201 | +728% | 0 | 0 | — |
case-02 | fail→fail | 9,653 | 26,766 | +177% | 1 | 1 | 0% | 616 | 2,895 | +370% | 0 | 0 | — |
case-03 | fail→fail | 8,615 | 16,329 | +90% | 1 | 1 | 0% | 486 | 2,646 | +444% | 0 | 0 | — |
case-04 | fail→pass | 12,833 | 4,922 | -62% | 1 | 1 | 0% | 2,103 | 1,727 | -18% | 0 | 0 | — |
case-05 | fail→pass | 19,431 | 4,125 | -79% | 1 | 1 | 0% | 1,721 | 1,496 | -13% | 0 | 0 | — |
case-06 | fail→pass | 10,968 | 3,237 | -70% | 1 | 1 | 0% | 1,740 | 1,270 | -27% | 0 | 0 | — |
case-07 | fail→pass | 11,612 | 4,922 | -58% | 1 | 1 | 0% | 1,627 | 1,683 | +3% | 0 | 0 | — |
case-08 | fail→pass | 11,419 | 6,299 | -45% | 1 | 1 | 0% | 1,526 | 1,614 | +6% | 0 | 0 | — |
case-09 | fail→pass | 13,195 | 6,567 | -50% | 1 | 1 | 0% | 1,878 | 1,987 | +6% | 0 | 0 | — |
case-10 | fail→pass | 10,640 | 5,936 | -44% | 1 | 1 | 0% | 1,401 | 1,309 | -7% | 0 | 0 | — |
case-11 | fail→pass | 14,581 | 4,748 | -67% | 1 | 1 | 0% | 2,076 | 1,577 | -24% | 0 | 0 | — |
case-12 | pass→pass | 18,480 | 2,982 | -84% | 1 | 1 | 0% | 1,323 | 1,234 | -7% | 0 | 0 | — |
case-13 | pass→pass | 29,686 | 6,886 | -77% | 1 | 1 | 0% | 2,048 | 1,683 | -18% | 0 | 0 | — |
case-14 | pass→pass | 20,739 | 14,370 | -31% | 1 | 1 | 0% | 1,699 | 1,993 | +17% | 0 | 0 | — |
case-15 | pass→pass | 28,709 | 5,406 | -81% | 1 | 1 | 0% | 2,312 | 1,811 | -22% | 0 | 0 | — |
case-16 | fail→pass | 17,504 | 5,735 | -67% | 1 | 1 | 0% | 2,526 | 1,628 | -36% | 0 | 0 | — |
case-17 | fail→pass | 9,312 | 2,625 | -72% | 1 | 1 | 0% | 1,235 | 1,182 | -4% | 0 | 0 | — |
case-18 | pass→pass | 30,831 | 5,504 | -82% | 1 | 1 | 0% | 1,864 | 1,459 | -22% | 0 | 0 | — |
case-19 | pass→pass | 14,466 | 4,473 | -69% | 1 | 1 | 0% | 2,079 | 1,413 | -32% | 0 | 0 | — |
case-20 | fail→pass | 17,376 | 4,806 | -72% | 1 | 1 | 0% | 2,767 | 1,597 | -42% | 0 | 0 | — |
case-21 | pass→pass | 22,385 | 30,952 | +38% | 1 | 1 | 0% | 4,003 | 5,752 | +44% | 0 | 0 | — |
case-22 | pass→pass | 20,835 | 20,410 | -2% | 1 | 1 | 0% | 3,818 | 3,351 | -12% | 0 | 0 | — |
case-23 | pass→pass | 14,187 | 12,145 | -14% | 1 | 1 | 0% | 2,184 | 2,895 | +33% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +48 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.