Install any skill in seconds. Free to start, no credit card required.
Get Started Free →GUI desktop app only. Writes, runs, and debugs Warp integration tests using the custom Builder/TestStep framework in `crates/integration`. Use when adding a new integration test, fixing a failing integration test, wiring a test into the manual runner or nextest suite, or verifying end-to-end UI and terminal behavior in Warp.
.claude/skills/warpdotdev-gui-integration-test/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 155% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 100% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 78% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 168% | 0% |
Scope — GUI desktop app only. This skill applies to Warp's GUI desktop front-end (the app/ crate on the WarpUI pixel/GPU framework). It does not apply to the headless TUI front-end (crates/warp_tui; cell-grid TuiElement library under crates/warpui_core/src/elements/tui), which has its own components, tests, and change-verification workflow. For TUI work, see the tui-ui-guidelines, tui-testing, and tui-verify-change skills instead.
Use this skill for Rust integration tests in Warp's custom framework under crates/integration/.
These are not ordinary unit tests. They boot a real Warp app instance, give it an isolated test home directory, drive it with synthetic UI and terminal events, and poll assertions until success or timeout.
Integration tests are the most expensive tests in the repo. They boot the app, they are orders of magnitude slower than a unit test, and they are the most likely to go flaky. Write one when the risk you are covering genuinely lives between components:
Do not reach for an integration test when:
rust-unit-tests); it will run in milliseconds and point straight at the failure.tui-testing — this harness does not drive the TUI at all.computer_use or the gui-integration-test-video skill instead of attaching weak assertions to a full app boot.If a behavior is hard to reach from a unit test only because of how the code is structured, prefer fixing the structure over writing a slow test around it.
The core pieces are:
crates/integration/src/bin/integration.rsBuilder factories.crates/integration/tests/common/mod.rscargo test and cargo nextest.PATH, RUST_*, WARP_*, WARPUI_*, WGPU_*, display-related vars).crates/integration/src/test.rspub use their functions so the runner can see them.crates/integration/tests/integration/ui_tests.rscrates/integration/tests/integration/shell_integration_tests.rscrates/integration/src/builder.rscrates/warpui_core/src/integration/driver.rson_finish.crates/warpui_core/src/integration/step.rsTestStep, input/event APIs, assertion polling, step-to-step data passing, and screenshot/recording hooks.app/src/integration_testing/crates/integration/tests/integration/*.rs calls run_integration_test("test_name").integration binary with the test name.crates/integration/src/bin/integration.rs looks up the name in register_tests(), builds the Builder, and turns it into a TestDriver.Builder::build(...) creates an isolated temp directory, points HOME at it, writes minimal rc files, and initializes file-backed user preferences.TestStep in order:PreconditionFailed, the binary exits with the rerun code and the outer harness retries the whole test.on_finish and export artifacts/runtime tags.This means integration tests should be written for a hermetic environment. Do not rely on the developer's real shell dotfiles, home directory contents, or persisted Warp settings.
Add the actual test function in a module under crates/integration/src/test/.
Use these heuristics:
crates/integration/tests/integration/ui_tests.rs if it is primarily a UI/app behavior test.crates/integration/tests/integration/shell_integration_tests.rs if it needs to run against every shell, or depends on a specific shell/set of shells.Being present in crates/integration/src/test/*.rs is not enough. For a test to run under cargo nextest, it also needs to be listed in one of the macro files in crates/integration/tests/integration/.
When adding a new integration test, do all of the following:
pub fn test_name() -> Builder in a module under crates/integration/src/test/.crates/integration/src/test.rs.pub use the new module's exports from crates/integration/src/test.rs.register_test!(test_name); in crates/integration/src/bin/integration.rs.test_name to either:crates/integration/tests/integration/ui_tests.rs, orcrates/integration/tests/integration/shell_integration_tests.rs#[ignore] when the task explicitly calls for manual-only coverage or there is a concrete, documented reason it cannot run reliably in CI.The normal shape is:
rustuse crate::Builder; use warp::integration_testing::step::new_step_with_default_assertions; use warp::integration_testing::terminal::{ clear_blocklist_to_remove_bootstrapped_blocks, execute_command_for_single_terminal_in_tab, wait_until_bootstrapped_single_pane_for_tab, util::ExpectedExitStatus, }; pub fn test_example() -> Builder { Builder::new() .with_step(wait_until_bootstrapped_single_pane_for_tab(0)) .with_step(clear_blocklist_to_remove_bootstrapped_blocks()) .with_step(execute_command_for_single_terminal_in_tab( 0, "echo hello".to_string(), ExpectedExitStatus::Success, "hello".to_string(), )) .with_step( new_step_with_default_assertions("Assert some UI state") .add_named_assertion("specific assertion name", |app, window_id| { // inspect app state and return AssertionOutcome warpui::integration::AssertionOutcome::Success }), ) }
Prefer a small number of focused steps with descriptive names over a huge monolithic test.
Builder::new()Start here almost every time.
Warp's wrapper automatically gives you:
HOMEWARPUI_USE_REAL_DISPLAY_IN_INTEGRATION_TESTS is presentwith_setup(...)Use this for filesystem or environment setup before the app runs.
Common patterns:
utils.set_env("NAME", Some(value))utils.test_dir()Prefer this over reaching into the real filesystem.
with_user_defaults(...)Use this to set persisted Warp preferences before the test starts.
This is the right tool for settings backed by user preferences rather than environment variables.
set_should_run_test(...)Use this to gate tests on shell/platform/runtime capabilities when the test genuinely cannot run everywhere.
with_on_finish(...)Use this for final verification or artifact inspection that should happen after all steps complete, such as checking that screenshots or recordings were written.
with_real_display()Use this explicitly when the test needs a real display for frame capture or visual workflows. Video/screenshot tests should normally be manual or ignored in CI unless there is a stable real-display path.
TestStep guidanceTestStep is the unit of execution. Each step can have:
Prefer:
wait_until_bootstrapped_single_pane_for_tab(0)new_step_with_default_assertions("...")new_step_with_default_assertions_for_pane("...", tab, pane)The default step helpers already assert:
These are good baseline invariants for most UI interactions.
Use high-level helpers from app/src/integration_testing/ whenever possible:
Drop to raw with_event(...), with_event_fn(...), or saved-position mouse events only when there is no suitable helper.
Prefer add_named_assertion(...) over unnamed assertions. Named assertions make failure output and runtime tags much easier to interpret.
Assertions are polled until success or timeout. Lean on that model instead of hardcoding sleeps.
Good pattern:
Avoid brittle timing assumptions.
If a later step needs data from an earlier one, use:
add_named_assertion_with_data_from_prior_step(...)StepDataMapThis is useful for saving measured positions, counts, IDs, or other values from prior frames.
set_retries(...) can help for a legitimately retryable step, but do not use it to hide deterministic failures. Prefer making the step more robust first.
PreconditionFailed for genuinely environmental flakesIf the environment reaches a state where the rest of the test is invalid, return AssertionOutcome::PreconditionFailed(...) instead of failing hard. The outer harness can rerun the entire test up to 10 times. The existing bootstrap helper is a good model for this.
Use this deliberately. The rerun mechanism exists for conditions the test genuinely cannot control, such as bootstrap racing or shell startup timing. It is not a way to turn an intermittently failing test green. A real bug that reproduces one run in five will pass under rerun and ship to users.
Before reaching for PreconditionFailed, confirm the failure is actually environmental by looking at the failure rate and the failure mode:
bashfor i in {0..50}; do RUST_BACKTRACE=full cargo run -p integration --bin integration -- test_name || break done
If the same assertion fails in different ways, or fails while the environment looks fine, it is a product or test bug. Fix it rather than absorbing it into a rerun.
The scope of a test follows the scope of what you boot and drive, so cover one user journey per test and let separate tests cover separate journeys. When a flow is long, prefer several shorter tests that each verify one hop over a single test that walks the whole path. A long test is slower, harder to diagnose, and fails for many unrelated reasons.
The framework can see internal state, which makes it tempting to assert on it. Anchor each test on the user-observable outcome:
Internal state assertions are still useful, but they should support the visible behavior rather than replace it. Assertions that mirror internal call sequences become change detectors: they break on every refactor and catch no bugs.
An integration test spans processes, so a stack trace tells the next engineer almost nothing. The failure output has to carry the context instead:
"terminal shows command output", not "check blocks".Assume the person reading the failure has never seen this test and does not own the code it covers.
Builder::new() already gives you an isolated HOME, generated rc files, and file-backed preferences. Preserve that. Set up everything you depend on inside the test via with_setup(...) and with_user_defaults(...), and never rely on the developer's dotfiles, real settings, network access, or state left behind by another test. A test that assumes a resource is already in the right state will fail for the wrong reason on someone else's machine or in CI.
Larger tests span components, so ownership is ambiguous by default and unowned tests rot. If you add one, you are the person who fixes it when it breaks. Put it in the module matching the feature area it covers so the next person can find the right owner.
For most terminal-facing tests, the first real step should be:
wait_until_bootstrapped_single_pane_for_tab(0)Do not start asserting on terminal UI before bootstrap completes.
If the test relies on saved positions like block_index:0, clear the block list after bootstrap:
clear_blocklist_to_remove_bootstrapped_blocks()Otherwise the first user-generated block index depends on bootstrap output and the active shell.
Prefer helpers like:
execute_command_for_single_terminal_in_tab(...)execute_echo(...)execute_echo_str(...)execute_long_running_command(...)These helpers already handle a lot of correctness and output validation.
Use this first while authoring:
bashcargo run -p integration --bin integration -- test_name
This is the fastest way to iterate on a specific test because it bypasses the outer Rust test wrapper and runs the named test directly.
Once it is wired into one of the tests/integration/*.rs macro lists, run it with nextest:
bashcargo nextest run --no-fail-fast --workspace test_name
For screenshot/video or other real-display flows:
bashWARPUI_USE_REAL_DISPLAY_IN_INTEGRATION_TESTS=1 cargo run -p integration --bin integration -- test_name
Or with nextest:
bashWARPUI_USE_REAL_DISPLAY_IN_INTEGRATION_TESTS=1 cargo nextest run --no-fail-fast --workspace test_name
bashRUST_BACKTRACE=1 cargo run -p integration --bin integration -- test_name
This is useful when running locally and you want to inspect the failed UI state:
bashWARPUI_PAUSE_INTEGRATION_TEST_ON_FAILURE=1 cargo run -p integration --bin integration -- test_name
Useful for understanding exactly what the test is doing:
bashWARPUI_PAUSE_INTEGRATION_TEST_AT_EVERY_STEP=1 cargo run -p integration --bin integration -- test_name
If the task is specifically about recording a test, collecting screenshots, or validating overlay/video artifacts, also use the gui-integration-test-video skill (located at .warp/skills/gui-integration-test-video/SKILL.md).
utils.set_env(...) affects runtime environment lookups such as std::env::var(...).
It does not affect compile-time lookups like option_env!(...). If the product code uses option_env!, changing the env var inside the test will not change that behavior without rebuilding.
Before considering a new integration test done, verify all of the following:
crates/integration/src/test/.crates/integration/src/test.rs.crates/integration/src/bin/integration.rs.src/test/*.rs and forgetting the nextest macro list.PreconditionFailed to paper over a deterministic bug.When asked to add or fix an integration test:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 16,956 | 14,379 | -15% | 1 | 1 | 0% | 2,938 | 7,506 | +155% | 0 | 0 | — |
case-02 | fail→pass | 22,395 | 11,278 | -50% | 1 | 1 | 0% | 3,446 | 6,877 | +100% | 0 | 0 | — |
case-03 | fail→pass | 19,315 | 4,904 | -75% | 1 | 1 | 0% | 3,552 | 5,670 | +60% | 0 | 0 | — |
case-04 | fail→pass | 19,898 | 7,143 | -64% | 1 | 1 | 0% | 3,347 | 5,965 | +78% | 0 | 0 | — |
case-05 | pass→pass | 14,815 | 7,552 | -49% | 1 | 1 | 0% | 2,279 | 5,955 | +161% | 0 | 0 | — |
case-06 | pass→pass | 16,086 | 5,805 | -64% | 1 | 1 | 0% | 2,263 | 5,587 | +147% | 0 | 0 | — |
case-07 | fail→pass | 15,771 | 8,451 | -46% | 1 | 1 | 0% | 2,355 | 6,315 | +168% | 0 | 0 | — |
case-08 | fail→pass | 15,205 | 9,781 | -36% | 1 | 1 | 0% | 2,340 | 6,182 | +164% | 0 | 0 | — |
case-09 | fail→pass | 14,884 | 3,958 | -73% | 1 | 1 | 0% | 2,246 | 5,477 | +144% | 0 | 0 | — |
case-10 | fail→pass | 14,403 | 5,548 | -61% | 1 | 1 | 0% | 2,220 | 5,715 | +157% | 0 | 0 | — |
case-11 | fail→pass | 15,062 | 7,482 | -50% | 1 | 1 | 0% | 2,362 | 5,714 | +142% | 0 | 0 | — |
case-12 | pass→pass | 11,556 | 6,608 | -43% | 1 | 1 | 0% | 1,868 | 5,799 | +210% | 0 | 0 | — |
case-13 | fail→pass | 11,783 | 3,437 | -71% | 1 | 1 | 0% | 1,586 | 5,305 | +234% | 0 | 0 | — |
case-14 | fail→fail | 11,301 | 4,020 | -64% | 1 | 1 | 0% | 1,783 | 5,484 | +208% | 0 | 0 | — |
case-15 | fail→pass | 14,982 | 8,932 | -40% | 1 | 1 | 0% | 2,465 | 6,271 | +154% | 0 | 0 | — |
case-16 | fail→fail | 12,995 | 6,362 | -51% | 1 | 1 | 0% | 2,078 | 5,603 | +170% | 0 | 0 | — |
case-17 | pass→pass | 13,662 | 7,299 | -47% | 1 | 1 | 0% | 1,924 | 6,119 | +218% | 0 | 0 | — |
case-18 | pass→pass | 14,833 | 6,391 | -57% | 1 | 1 | 0% | 2,228 | 5,787 | +160% | 0 | 0 | — |
case-19 | fail→fail | 19,350 | 3,500 | -82% | 1 | 1 | 0% | 2,902 | 5,303 | +83% | 0 | 0 | — |
case-20 | fail→pass | 13,139 | 3,518 | -73% | 1 | 1 | 0% | 2,370 | 5,263 | +122% | 0 | 0 | — |
case-21 | fail→fail | 21,279 | 3,247 | -85% | 1 | 1 | 0% | 3,070 | 5,243 | +71% | 0 | 0 | — |
case-22 | pass→pass | 15,149 | 8,096 | -47% | 1 | 1 | 0% | 2,111 | 5,953 | +182% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.