Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Orchestrates cost-conscious cloud verification across available operating-system and architecture runners after cheaper checks pass. Use whenever a code change, bug fix, build, packaging flow, native dependency, UI behavior, filesystem behavior, or test has material cross-platform implications, even if the user only asks to verify or test the change. Consult domain verification skills for what to test; use this skill to choose and access the smallest relevant set of platforms.
.claude/skills/warpdotdev-cross-platform-cloud-verification/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 320% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 609% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 128% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 167% | 0% |
Verify a change on the smallest useful set of cloud platforms. This skill owns runner discovery, platform selection, child orchestration, and result aggregation. It does not define the product-specific verification procedure.
Cross-platform execution consumes remote compute and credits, so use it as the last verification layer rather than as an exploratory first step.
Finish the cheaper feedback loops first:
verification.
by cloud agents. Do not test the default branch when the intended change is only local.
Do not commit or push changes merely to satisfy this workflow unless the user has authorized that action. If the change is not remotely reachable, report the precondition and ask for the minimum action needed.
Skip cloud fan-out when local verification already shows the change is broken. Do not use remote platforms as a substitute for diagnosing an ordinary local failure.
Inspect the available skill descriptions and read every skill that materially defines how to verify the affected surface. Examples include GUI computer-use verification, TUI live verification, integration testing, unit testing, CI diagnosis, packaging, or repository-specific validation.
Extract from those skills:
Build one shared verification procedure from that guidance. Tell each child which verification skill to read when it is available in the child environment, and include the essential procedure directly so the run remains actionable if the skill is unavailable there.
This skill decides where to run that procedure, not what the procedure should be.
List runners immediately before selecting platforms:
shoz-dev runner list --output-format json
Use the returned runner metadata rather than hard-coding names or IDs. Relevant fields include uid, name, os, arch, compute capacity, image or macOS version, and setup commands.
If runner discovery fails because of authentication, permissions, or service availability, report the blocker instead of inventing a runner matrix.
Determine which dimensions the change can plausibly affect from the diff, repository configuration, reported bug, and verification guidance.
Treat the change as OS-sensitive when it touches or depends on items such as:
Treat the change as architecture-sensitive only with evidence such as:
Do not infer architecture sensitivity merely because both x86-64 and AArch64 runners exist.
Filter discovered runners to those that can execute the verification procedure, then apply these defaults:
architecture-sensitive.
When several runners cover the same platform, prefer in order:
For architecture-insensitive Linux changes, prefer the repository's primary Linux architecture; if the repository gives no signal, use x86-64 as the single representative.
Before launching, record a concise selection rationale and explicitly list relevant platforms omitted because no runner is available. Missing runners do not block useful verification on available relevant platforms.
State each omission once. After the matrix is decided, do not keep repeating that an irrelevant or redundant platform was not selected in child prompts, result rows, evidence, and the conclusion.
Every child prompt should contain:
Trust orchestration to route the child to the requested runner. Do not make every child re-confirm its OS and architecture. Ask for platform introspection only when the verification depends on an exact OS version or capability, or when there is evidence of a routing problem.
Ask each child to return:
textPlatform: <runner name; OS; architecture; version if relevant> Commit: <tested SHA> Status: <passed | failed | blocked> Checks: - <command or interaction>: <result> Evidence: - <artifact, screenshot, log, or concise observation> Deviations: - <difference from the requested procedure, or none>
Children may make temporary, uncommitted setup adjustments required by the verification skill, but they must report them and must not push source changes.
run_agentsUse run_agents so child IDs, messages, lifecycle events, and artifacts remain part of the parent orchestration flow. Use the same repository environment for every selected platform and set remote.runner_id to the discovered runner UID.
textsummary: Verifying the change on <OS/architecture>. base_prompt: <shared verification-only instructions> remote: environment_id: <repository environment ID> runner_id: <selected runner UID> computer_use_enabled: <true only when the verification procedure needs it> agent_run_configs: - name: <short platform-specific name> prompt: <platform-specific procedure and expected evidence>
If no suitable repository environment is already known, inspect oz-dev environment list and choose one that checks out the target repository. Do not silently create or mutate an environment.
remote.runner_id is run-wide, so runners with different UIDs require separate run_agents calls. Use one single-child batch per selected runner, or group multiple independently useful children only when they share the same runner and run-wide configuration. Distinct runner IDs are a legitimate reason for separate batches; do not place children for different platforms in one batch.
Omit model_id and remote.harness unless the user requested them. Attach relevant verification skills when the children need them. If an approved orchestration config is active, ensure its resolved runner matches the selected runner because config resolution takes precedence over call fields.
Capture each trusted agent_id from the launched result. Coordinate through pushed child messages and lifecycle events; do not poll oz-dev run get or list_messages_from_agents. Read notified messages, intervene only for a failure or actionable block, and use wait_for_events when no other work can proceed. Allow at most one retry for a clearly transient infrastructure failure. A product failure is evidence, not a reason to spend more credits repeating the same check.
Wait for every launched child, then report:
textCross-platform verification: <passed | failed | incomplete> Local gate: <checks completed before cloud launch> Results: - <OS/arch; runner>: <why selected> - <passed/failed/blocked> Unverified: - <relevant OS/arch>: unverified because no suitable runner was available Evidence: - <platform>: <commands, observations, and artifact/run links> Conclusion: - <what the results establish and what remains unverified>
Use passed only when every selected platform passed and no required platform is unavailable. Use failed when any verification check failed. Use incomplete when runs were blocked or a relevant platform lacked a runner.
Keep dry-run proposals and final summaries compact. Include the local gate, the selected matrix, one concise omission rationale when it is material, the verification procedure, and verdict rules. Do not repeat setup mechanics or platform exclusions in multiple sections. A dry-run proposal should usually fit in roughly 300-500 words; expand only when the domain verification procedure genuinely requires more detail.
independently necessary.
to obtain a passing result.
fix locally, rerun cheap checks, and only then decide whether another cloud verification pass is justified.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 21,361 | 10,524 | -51% | 1 | 1 | 0% | 433 | 2,469 | +470% | 0 | 0 | — |
case-02 | fail→fail | 7,500 | 20,601 | +175% | 1 | 1 | 0% | 229 | 3,000 | +1210% | 0 | 0 | — |
case-03 | fail→fail | 8,214 | 12,215 | +49% | 1 | 1 | 0% | 314 | 2,885 | +819% | 0 | 0 | — |
case-04 | pass→pass | 34,310 | 26,979 | -21% | 1 | 1 | 0% | 5,452 | 7,014 | +29% | 0 | 0 | — |
case-05 | pass→fail | 17,085 | 7,152 | -58% | 1 | 1 | 0% | 2,404 | 2,479 | +3% | 0 | 0 | — |
case-06 | pass→pass | 9,691 | 16,073 | +66% | 1 | 1 | 0% | 1,080 | 4,265 | +295% | 0 | 0 | — |
case-07 | fail→pass | 6,237 | 10,380 | +66% | 1 | 1 | 0% | 888 | 3,727 | +320% | 0 | 0 | — |
case-08 | fail→pass | 4,739 | 17,427 | +268% | 1 | 1 | 0% | 688 | 4,880 | +609% | 0 | 0 | — |
case-09 | fail→pass | 41,794 | 12,011 | -71% | 1 | 1 | 0% | 1,756 | 3,996 | +128% | 0 | 0 | — |
case-10 | pass→pass | 6,873 | 10,998 | +60% | 1 | 1 | 0% | 996 | 4,237 | +325% | 0 | 0 | — |
case-11 | pass→pass | 19,270 | 5,172 | -73% | 1 | 1 | 0% | 1,669 | 2,877 | +72% | 0 | 0 | — |
case-12 | pass→pass | 9,932 | 3,700 | -63% | 1 | 1 | 0% | 1,355 | 2,691 | +99% | 0 | 0 | — |
case-13 | pass→pass | 14,068 | 4,122 | -71% | 1 | 1 | 0% | 1,976 | 2,636 | +33% | 0 | 0 | — |
case-14 | pass→pass | 12,715 | 6,129 | -52% | 1 | 1 | 0% | 1,934 | 2,896 | +50% | 0 | 0 | — |
case-15 | fail→pass | 11,785 | 4,064 | -66% | 1 | 1 | 0% | 1,713 | 2,614 | +53% | 0 | 0 | — |
case-16 | fail→pass | 7,179 | 4,513 | -37% | 1 | 1 | 0% | 1,029 | 2,745 | +167% | 0 | 0 | — |
case-17 | pass→pass | 13,014 | 4,893 | -62% | 1 | 1 | 0% | 1,844 | 2,866 | +55% | 0 | 0 | — |
case-18 | fail→pass | 10,742 | 4,109 | -62% | 1 | 1 | 0% | 1,333 | 2,614 | +96% | 0 | 0 | — |
case-19 | pass→pass | 11,112 | 6,050 | -46% | 1 | 1 | 0% | 1,572 | 2,695 | +71% | 0 | 0 | — |
case-20 | pass→pass | 8,119 | 3,920 | -52% | 1 | 1 | 0% | 1,139 | 2,573 | +126% | 0 | 0 | — |
case-21 | fail→pass | 19,707 | 6,842 | -65% | 1 | 1 | 0% | 2,697 | 3,123 | +16% | 0 | 0 | — |
case-22 | pass→pass | 14,436 | 3,491 | -76% | 1 | 1 | 0% | 2,143 | 2,549 | +19% | 0 | 0 | — |
case-23 | fail→pass | 15,751 | 3,410 | -78% | 1 | 1 | 0% | 2,599 | 2,587 | -0% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 18 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +30 percentage points is the difference between those two pass rates over the 18 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.