Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Verifies user-facing Warp client changes by spawning a cloud agent with computer use to test Warp. Use only when the user explicitly requested computer-use verification or accepted an offer to run it, and ONLY in non-sandboxed environments and local environments. Triggers a cloud agent that runs the test-warp-ui skill.
.claude/skills/warpdotdev-verify-ui-change-in-cloud/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 87% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -31% | 0% |
Use this workflow to verify a user-facing client change by spawning a cloud agent with computer use capabilities. Do not invoke it automatically after every UI change: launch only when the user explicitly requested computer-use verification or accepted an offer to run it. After completing a user-facing change, you may briefly offer this verification; do not launch until the user says yes. This applies to any change that affects what the user sees or experiences in the running app — not just visual/UI changes, but also startup behavior, config handling, migration flows, and other client-side logic.
The cloud agent runs in a fresh environment that clones the repo. Your changes must be pushed to a branch so the cloud agent can access them.
Before spawning the cloud agent, detect which repository you are running in. Check the Git remote URL to determine the repo:
bashgit remote get-url origin
Verify the remote URL contains warpdotdev/warp. If it does not, warn the user that this skill only supports the warp repository and stop.
The environment ID for the warp Dev Environment is SVhg783GBFQHk1OfdPfFU9.
Use the run_agents tool to spawn a remote cloud agent. A single-child batch (one entry in agent_run_configs) is valid.
summary: a brief declarative explanation, e.g. "Spawning a cloud agent with computer use to verify the UI change."base_prompt: include an instruction to read and follow the test-warp-ui skill, followed by the verification task (see the next section)remote.environment_id: SVhg783GBFQHk1OfdPfFU9remote.computer_use_enabled: trueagent_run_configs: a single entry with name set to a short display name such as "verify-ui-change". The per-agent prompt can be empty since base_prompt covers the task.The test-warp-ui skill is bundled, so the cloud agent has it automatically. Tell the agent to invoke it by name in the base_prompt (e.g. "Read and follow the test-warp-ui skill.").
The prompt should tell the cloud agent:
Example prompts:
I changed the settings dialog header to use a larger font and blue color.
Hardcode the settings dialog to open on launch, then describe the header text,
font size relative to other text, and color.I added a migration that symlinks config from ~/.warp into ~/.warp-preview on first launch.
The migration is gated on Channel::Preview. Before building, hardcode the migration to run
regardless of channel by removing the channel check. Also create a fake ~/.warp directory
with test files. After launching Warp, verify the symlinks were created in ~/.warp-preview.The cloud agent builds Warp with cargo run, which may not match the exact runtime conditions of your change (e.g., different channel, missing feature flags, absent preconditions). When this happens, instruct the cloud agent to temporarily hardcode the code so the build exercises the path you need to test. Common examples:
Be explicit in the prompt about what to hardcode and why — the cloud agent won't infer this on its own.
No extra surfacing step is needed — the Warp client displays the cloud agent run automatically.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,645 | 11,026 | +137% | 1 | 1 | 0% | 682 | 1,126 | +65% | 0 | 0 | — |
case-02 | fail→fail | 8,048 | 7,461 | -7% | 1 | 1 | 0% | 1,217 | 1,341 | +10% | 0 | 0 | — |
case-03 | fail→fail | 4,817 | 19,466 | +304% | 1 | 1 | 0% | 668 | 1,166 | +75% | 0 | 0 | — |
case-04 | fail→pass | 11,071 | 4,012 | -64% | 1 | 1 | 0% | 1,468 | 1,429 | -3% | 0 | 0 | — |
case-05 | fail→pass | 6,945 | 5,533 | -20% | 1 | 1 | 0% | 876 | 1,634 | +87% | 0 | 0 | — |
case-06 | fail→pass | 10,520 | 7,601 | -28% | 1 | 1 | 0% | 1,558 | 1,811 | +16% | 0 | 0 | — |
case-07 | pass→fail | 19,284 | 8,154 | -58% | 1 | 1 | 0% | 1,936 | 1,244 | -36% | 0 | 0 | — |
case-08 | fail→fail | 17,646 | 6,655 | -62% | 1 | 1 | 0% | 1,773 | 2,083 | +17% | 0 | 0 | — |
case-09 | fail→fail | 10,705 | 5,610 | -48% | 1 | 1 | 0% | 1,516 | 1,181 | -22% | 0 | 0 | — |
case-10 | fail→fail | 11,291 | 6,843 | -39% | 1 | 1 | 0% | 1,266 | 1,217 | -4% | 0 | 0 | — |
case-11 | fail→pass | 11,242 | 3,938 | -65% | 1 | 1 | 0% | 1,658 | 1,485 | -10% | 0 | 0 | — |
case-12 | fail→pass | 20,295 | 2,624 | -87% | 1 | 1 | 0% | 1,847 | 1,280 | -31% | 0 | 0 | — |
case-13 | fail→pass | 8,922 | 6,820 | -24% | 1 | 1 | 0% | 1,127 | 1,833 | +63% | 0 | 0 | — |
case-14 | pass→pass | 12,932 | 6,410 | -50% | 1 | 1 | 0% | 1,985 | 1,414 | -29% | 0 | 0 | — |
case-15 | pass→pass | 4,974 | 6,688 | +34% | 1 | 1 | 0% | 583 | 1,425 | +144% | 0 | 0 | — |
case-16 | pass→pass | 8,564 | 3,317 | -61% | 1 | 1 | 0% | 1,203 | 1,248 | +4% | 0 | 0 | — |
case-17 | fail→pass | 13,572 | 8,404 | -38% | 1 | 1 | 0% | 1,895 | 2,061 | +9% | 0 | 0 | — |
case-18 | pass→fail | 4,155 | 6,793 | +63% | 1 | 1 | 0% | 269 | 1,368 | +409% | 0 | 0 | — |
case-19 | pass→pass | 5,514 | 4,521 | -18% | 1 | 1 | 0% | 859 | 1,568 | +83% | 0 | 0 | — |
case-20 | pass→fail | 12,679 | 7,270 | -43% | 1 | 1 | 0% | 1,941 | 1,147 | -41% | 0 | 0 | — |
case-21 | fail→pass | 10,880 | 6,392 | -41% | 1 | 1 | 0% | 1,570 | 1,825 | +16% | 0 | 0 | — |
case-22 | fail→pass | 22,610 | 4,469 | -80% | 1 | 1 | 0% | 2,030 | 1,593 | -22% | 0 | 0 | — |
case-23 | pass→pass | 19,110 | 2,190 | -89% | 1 | 1 | 0% | 3,246 | 1,202 | -63% | 0 | 0 | — |
case-24 | fail→fail | 9,042 | 10,176 | +13% | 1 | 1 | 0% | 1,099 | 2,455 | +123% | 0 | 0 | — |
case-25 | pass→pass | 13,931 | 4,823 | -65% | 1 | 1 | 0% | 1,880 | 1,596 | -15% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 17 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +24 percentage points is the difference between those two pass rates over the 17 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.