Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Learn a reliable browser workflow by iterating on a real web task, recording strategy, and proposing a reusable skill.
.claude/skills/cowork-os-autobrowse/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 445% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 101% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 166% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 81% | 0% |
Autobrowse turns expensive browser exploration into durable operational memory. It runs a real browser task, studies the trace and diagnostics, iterates on the strategy, then graduates the reliable path into a reviewable skill artifact.
strategy.md, iterations.md, draft-skill.md, and usually an approval-gated skill_proposal.The user should not need to provide structured fields. Treat a plain request plus an optional link as enough input. Infer the objective from the request, infer the target site from any URL/domain in the request, default to 3 iterations, and default to creating a proposal.
browser_console, browser_network, browser_storage, browser_snapshot, and browser_evaluate when available. Use browser_trace_start and browser_trace_stop as supplemental diagnostics when the runtime exposes a readable trace summary. Capture only redacted, relevant evidence.strategy.md before the next iteration. The next attempt must read it first.draft-skill.md. If safe and useful, create a skill_proposal so the user can approve the new skill.A graduated skill must include:
Use skill_proposal with action create by default. Use draft-only when the workflow is too fragile, too sensitive, or still missing validation evidence.
An Autobrowse run is not complete until strategy.md, iterations.md, and draft-skill.md exist in the run directory. If proposal creation fails, record the exact failure in iterations.md and keep the draft skill reviewable.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 20,859 | 23,789 | +14% | 1 | 1 | 0% | 3,750 | 3,994 | +7% | 0 | 0 | — |
case-02 | fail→fail | 22,274 | 21,070 | -5% | 1 | 1 | 0% | 3,805 | 2,891 | -24% | 0 | 0 | — |
case-12 | pass→pass | 10,225 | 8,321 | -19% | 1 | 1 | 0% | 770 | 1,496 | +94% | 0 | 0 | — |
case-03 | fail→fail | 19,473 | 10,897 | -44% | 1 | 1 | 0% | 3,044 | 1,881 | -38% | 0 | 0 | — |
case-04 | pass→pass | 6,686 | 6,581 | -2% | 1 | 1 | 0% | 1,234 | 2,169 | +76% | 0 | 0 | — |
case-05 | pass→pass | 15,173 | 9,100 | -40% | 1 | 1 | 0% | 2,870 | 2,458 | -14% | 0 | 0 | — |
case-06 | fail→pass | 3,123 | 8,423 | +170% | 1 | 1 | 0% | 407 | 2,219 | +445% | 0 | 0 | — |
case-07 | fail→fail | 9,708 | 32,279 | +232% | 1 | 1 | 0% | 1,647 | 5,472 | +232% | 0 | 0 | — |
case-08 | fail→pass | 14,605 | 24,635 | +69% | 1 | 1 | 0% | 2,458 | 4,947 | +101% | 0 | 0 | — |
case-09 | fail→pass | 9,663 | 20,542 | +113% | 1 | 1 | 0% | 1,612 | 4,286 | +166% | 0 | 0 | — |
case-10 | fail→pass | 12,790 | 22,236 | +74% | 1 | 1 | 0% | 2,018 | 3,652 | +81% | 0 | 0 | — |
case-11 | fail→fail | 10,851 | 21,411 | +97% | 1 | 1 | 0% | 633 | 2,466 | +290% | 0 | 0 | — |
case-13 | fail→pass | 16,780 | 16,889 | +1% | 1 | 1 | 0% | 2,259 | 3,560 | +58% | 0 | 0 | — |
case-14 | fail→fail | 14,100 | 18,974 | +35% | 1 | 1 | 0% | 2,274 | 3,928 | +73% | 0 | 0 | — |
case-15 | pass→pass | 15,338 | 11,605 | -24% | 1 | 1 | 0% | 2,520 | 2,869 | +14% | 0 | 0 | — |
case-16 | fail→fail | 16,971 | 6,419 | -62% | 1 | 1 | 0% | 2,639 | 1,103 | -58% | 0 | 0 | — |
case-17 | pass→pass | 6,232 | 13,299 | +113% | 1 | 1 | 0% | 962 | 3,093 | +222% | 0 | 0 | — |
case-18 | fail→pass | 12,829 | 14,167 | +10% | 1 | 1 | 0% | 2,140 | 3,519 | +64% | 0 | 0 | — |
case-19 | fail→pass | 9,774 | 16,543 | +69% | 1 | 1 | 0% | 1,333 | 3,722 | +179% | 0 | 0 | — |
case-20 | fail→fail | 14,465 | 8,897 | -38% | 1 | 1 | 0% | 1,341 | 1,098 | -18% | 0 | 0 | — |
case-21 | fail→fail | 11,587 | 8,778 | -24% | 1 | 1 | 0% | 1,845 | 1,152 | -38% | 0 | 0 | — |
case-22 | pass→pass | 11,289 | 16,968 | +50% | 1 | 1 | 0% | 1,701 | 3,632 | +114% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.