Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Reproduce CUA-Harness experiments on WeaveBench from a GitHub checkout. Use when the user wants an AI coding agent to set up dependencies, download WeaveBench assets, prepare the 120G VM, configure Qwen/Anthropic-compatible APIs, run smoke tests, launch full or subset evaluations, inspect logs, or summarize scores for this repository.
.claude/skills/amap-ml-weavebench-cua-reproduce/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 1410% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 140% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 180% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 19% | 0% |
Use this skill to guide an agent through a complete, reproducible WeaveBench CUA-Harness run. The expected repository layout is:
textWeaveBench-harness/ WeaveBench/ cua-harness/ skills/weavebench-cua-reproduce/
Do not invent alternate launch commands. Prefer the bundled helper script and the project scripts under WeaveBench/scripts/.
Supported automation:
cua_harness_claudecode.Not automated:
If the user requests unsupported automation, explain the boundary and provide the closest supported local Docker/KVM path.
Before doing setup work on a new machine, run or mentally follow:
bash./skills/weavebench-cua-reproduce/scripts/reproduce.sh intake
Confirm:
Ubuntu_120G.qcow2.doctor, smoke, subset, or full 114-task evaluation.WeaveBench/ and cua-harness/.references/configuration.md when you need exact defaults, environment variables, or paths.references/assets.md before downloading or validating WeaveBench tasks, runtime assets, judge templates, or VM files.references/verify.md before reporting whether setup or a run is actually verified.references/troubleshooting.md when setup, VM, API, image proxy, warmup, or judge errors occur.WeaveBench/.Use:
bash./skills/weavebench-cua-reproduce/scripts/reproduce.sh <command>
Commands:
textinstall Install local WeaveBench package and OpenClaw if needed. download Download task files, Claude Code runtime, judge template, and VM. vm120g Create cache/vm/Ubuntu_120G.qcow2 from the official Ubuntu.qcow2. doctor Run the read-only environment checker. intake Print setup questions for a new machine. status Print a non-destructive setup status report. smoke Run a one-task, one-round smoke test. full Launch the full 114-task evaluation in tmux. stats Summarize scores for a result directory. plan Print the minimal manual command sequence.
For a fresh machine, the usual sequence is:
bash./skills/weavebench-cua-reproduce/scripts/reproduce.sh install ./skills/weavebench-cua-reproduce/scripts/reproduce.sh download ./skills/weavebench-cua-reproduce/scripts/reproduce.sh vm120g ./skills/weavebench-cua-reproduce/scripts/reproduce.sh status ./skills/weavebench-cua-reproduce/scripts/reproduce.sh doctor ./skills/weavebench-cua-reproduce/scripts/reproduce.sh smoke ./skills/weavebench-cua-reproduce/scripts/reproduce.sh full
Before doctor, smoke, or full, ensure these are set:
bashexport WEAVEBENCH_LITELLM_KEY="YOUR_API_KEY" export WEAVEBENCH_LITELLM_BASE_URL="https://YOUR_ANTHROPIC_COMPATIBLE_ENDPOINT/v1"
If the provider cannot accept large base64 image payloads, enable image URL proxy:
bashexport WEAVEBENCH_IMAGE_PROXY=1 export WEAVEBENCH_IMAGE_PROXY_UPLOAD_URL_TPL="https://YOUR_IMAGE_UPLOAD_ENDPOINT/{id}" export WEAVEBENCH_IMAGE_PROXY_SHOW_URL_TPL="https://YOUR_PUBLIC_IMAGE_URL/{id}.png" export WEAVEBENCH_IMAGE_PROXY_UPLOAD_MODE="raw"
If the provider can directly accept large base64 screenshots:
bashexport WEAVEBENCH_IMAGE_PROXY=0
WeaveBench/cache/vm/Ubuntu_120G.qcow2.WEAVEBENCH_GROW_ROOTFS=1 unless the user explicitly disables it.qwen3.7-plus, cua_harness_claudecode, GUI mode, 25 CUA rounds, and 5 concurrent VM environments for full evaluation unless the user asks for a subset.smoke before full on a new machine.cache/, results/, logs/, Docker volumes, or judge workspaces unless the user explicitly asks.configured and verified, configured but not smoke-tested, not configured, or blocked awaiting user action.Run one domain:
bashexport WEAVEBENCH_DOMAINS="WEB" ./skills/weavebench-cua-reproduce/scripts/reproduce.sh full
Run one task:
bashexport WEAVEBENCH_DOMAINS="WEB" export WEAVEBENCH_TASK_FILTER="WEB_task_10_lighthouse" export WEAVEBENCH_NUM_ENVS=1 ./skills/weavebench-cua-reproduce/scripts/reproduce.sh full
Summarize a run:
bash./skills/weavebench-cua-reproduce/scripts/reproduce.sh stats \ WeaveBench/results/<run_name>/gui/qwen3.7-plus
End setup or launch tasks with this shape:
textEnvironment: configured and verified / configured but not smoke-tested / not configured / blocked awaiting user action Assets: ... 120G VM: ... Model API: ... Image proxy: ... Judge: ... Smoke test: passed / failed / skipped Full run: launched in tmux <name> / not launched Provenance: run_provenance.json written / not applicable Remaining blockers: ...
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | fail→pass | 26,801 | 10,347 | -61% | 1 | 1 | 0% | 4,056 | 2,580 | -36% | 0 | 0 | — |
case-01 | fail→fail | 12,476 | 9,330 | -25% | 1 | 1 | 0% | 492 | 1,972 | +301% | 0 | 0 | — |
case-02 | fail→fail | 8,319 | 9,785 | +18% | 1 | 1 | 0% | 233 | 1,997 | +757% | 0 | 0 | — |
case-03 | pass→fail | 10,719 | 15,618 | +46% | 1 | 1 | 0% | 2,023 | 2,121 | +5% | 0 | 0 | — |
case-04 | fail→pass | 15,011 | 12,033 | -20% | 1 | 1 | 0% | 206 | 3,111 | +1410% | 0 | 0 | — |
case-05 | fail→fail | 4,635 | 18,095 | +290% | 1 | 1 | 0% | 705 | 2,172 | +208% | 0 | 0 | — |
case-07 | fail→fail | 7,895 | 15,959 | +102% | 1 | 1 | 0% | 1,335 | 3,019 | +126% | 0 | 0 | — |
case-08 | fail→pass | 10,628 | 3,669 | -65% | 1 | 1 | 0% | 997 | 2,394 | +140% | 0 | 0 | — |
case-09 | fail→fail | 8,993 | 10,336 | +15% | 1 | 1 | 0% | 1,418 | 2,791 | +97% | 0 | 0 | — |
case-10 | fail→pass | 17,926 | 9,933 | -45% | 1 | 1 | 0% | 954 | 2,674 | +180% | 0 | 0 | — |
case-11 | fail→pass | 17,769 | 8,790 | -51% | 1 | 1 | 0% | 2,032 | 2,423 | +19% | 0 | 0 | — |
case-12 | fail→pass | 15,959 | 16,617 | +4% | 1 | 1 | 0% | 2,043 | 2,478 | +21% | 0 | 0 | — |
case-13 | fail→pass | 29,975 | 9,807 | -67% | 1 | 1 | 0% | 1,858 | 2,633 | +42% | 0 | 0 | — |
case-14 | fail→fail | 19,846 | 28,166 | +42% | 1 | 1 | 0% | 2,486 | 4,890 | +97% | 0 | 0 | — |
case-15 | fail→pass | 15,246 | 8,997 | -41% | 1 | 1 | 0% | 1,816 | 3,451 | +90% | 0 | 0 | — |
case-16 | fail→pass | 5,903 | 8,156 | +38% | 1 | 1 | 0% | 1,006 | 2,361 | +135% | 0 | 0 | — |
case-17 | fail→pass | 16,979 | 11,545 | -32% | 1 | 1 | 0% | 1,882 | 3,041 | +62% | 0 | 0 | — |
case-18 | fail→pass | 15,000 | 12,809 | -15% | 1 | 1 | 0% | 1,774 | 3,013 | +70% | 0 | 0 | — |
case-19 | pass→fail | 11,401 | 14,211 | +25% | 1 | 1 | 0% | 1,018 | 4,323 | +325% | 0 | 0 | — |
case-20 | fail→fail | 20,422 | 13,669 | -33% | 1 | 1 | 0% | 2,124 | 3,191 | +50% | 0 | 0 | — |
case-21 | fail→pass | 15,763 | 10,418 | -34% | 1 | 1 | 0% | 2,011 | 2,711 | +35% | 0 | 0 | — |
case-22 | fail→pass | 17,407 | 15,030 | -14% | 1 | 1 | 0% | 2,367 | 3,538 | +49% | 0 | 0 | — |
case-23 | pass→fail | 13,824 | 10,808 | -22% | 1 | 1 | 0% | 1,157 | 2,386 | +106% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 17 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +43 percentage points is the difference between those two pass rates over the 17 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.