Install any skill in seconds. Free to start, no credit card required.
Get Started Free →AI-powered browser automation — navigate sites, fill forms, extract structured data, log in with stored credentials, and build reusable multi-step workflows using natural language. Install: pip install skyvern && skyvern setup
.claude/skills/davepoon-skyvern/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 112% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -30% | 0% |
Control a real browser with natural language. Skyvern uses Vision LLMs and computer vision instead of brittle XPath/DOM selectors, so automations survive UI changes.
Cloud (recommended):
bashpip install skyvern skyvern setup # interactive client selection
Or add the MCP server directly to your Claude Code config:
json{ "mcpServers": { "skyvern": { "type": "streamable-http", "url": "https://api.skyvern.com/mcp/", "headers": { "x-api-key": "YOUR_SKYVERN_API_KEY" } } } }
Get your API key at app.skyvern.com.
"Navigate to example.com and extract all product prices"
"Log into my account and download the latest invoice"
"Fill out the shipping form and click Submit"
"Take a screenshot of the current page"
"Build a workflow that runs this every Monday"| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,434 | 5,559 | -51% | 1 | 1 | 0% | 2,025 | 1,535 | -24% | 0 | 0 | — |
case-02 | fail→pass | 4,240 | 5,652 | +33% | 1 | 1 | 0% | 716 | 1,516 | +112% | 0 | 0 | — |
case-03 | fail→pass | 14,246 | 9,517 | -33% | 1 | 1 | 0% | 2,839 | 2,223 | -22% | 0 | 0 | — |
case-04 | pass→pass | 6,534 | 6,987 | +7% | 1 | 1 | 0% | 1,490 | 1,993 | +34% | 0 | 0 | — |
case-05 | pass→pass | 8,171 | 6,414 | -22% | 1 | 1 | 0% | 1,437 | 1,579 | +10% | 0 | 0 | — |
case-06 | pass→pass | 10,695 | 8,034 | -25% | 1 | 1 | 0% | 2,273 | 2,256 | -1% | 0 | 0 | — |
case-07 | pass→pass | 15,644 | 5,202 | -67% | 1 | 1 | 0% | 2,591 | 1,484 | -43% | 0 | 0 | — |
case-08 | fail→pass | 5,917 | 1,687 | -71% | 1 | 1 | 0% | 1,114 | 745 | -33% | 0 | 0 | — |
case-09 | fail→pass | 11,873 | 5,449 | -54% | 1 | 1 | 0% | 2,045 | 1,429 | -30% | 0 | 0 | — |
case-10 | fail→pass | 11,363 | 1,466 | -87% | 1 | 1 | 0% | 2,201 | 674 | -69% | 0 | 0 | — |
case-11 | fail→pass | 9,508 | 1,377 | -86% | 1 | 1 | 0% | 1,674 | 691 | -59% | 0 | 0 | — |
case-12 | pass→pass | 8,010 | 2,883 | -64% | 1 | 1 | 0% | 1,431 | 784 | -45% | 0 | 0 | — |
case-13 | fail→pass | 9,436 | 3,192 | -66% | 1 | 1 | 0% | 1,563 | 968 | -38% | 0 | 0 | — |
case-14 | pass→pass | 12,102 | 11,147 | -8% | 1 | 1 | 0% | 2,145 | 2,505 | +17% | 0 | 0 | — |
case-15 | fail→pass | 10,838 | 4,658 | -57% | 1 | 1 | 0% | 1,997 | 1,222 | -39% | 0 | 0 | — |
case-16 | pass→pass | 1,985 | 1,158 | -42% | 1 | 1 | 0% | 289 | 625 | +116% | 0 | 0 | — |
case-17 | pass→pass | 12,361 | 8,532 | -31% | 1 | 1 | 0% | 2,243 | 2,009 | -10% | 0 | 0 | — |
case-18 | pass→pass | 3,847 | 1,417 | -63% | 1 | 1 | 0% | 715 | 674 | -6% | 0 | 0 | — |
case-19 | pass→pass | 2,000 | 1,497 | -25% | 1 | 1 | 0% | 249 | 705 | +183% | 0 | 0 | — |
case-20 | fail→pass | 5,089 | 1,624 | -68% | 1 | 1 | 0% | 981 | 681 | -31% | 0 | 0 | — |
case-21 | pass→pass | 10,829 | 3,734 | -66% | 1 | 1 | 0% | 1,831 | 1,073 | -41% | 0 | 0 | — |
case-22 | pass→pass | 1,879 | 1,336 | -29% | 1 | 1 | 0% | 220 | 638 | +190% | 0 | 0 | — |
case-23 | pass→pass | 13,477 | 7,303 | -46% | 1 | 1 | 0% | 2,472 | 1,879 | -24% | 0 | 0 | — |
case-24 | pass→pass | 5,127 | 2,408 | -53% | 1 | 1 | 0% | 810 | 794 | -2% | 0 | 0 | — |
case-25 | fail→pass | 6,372 | 1,026 | -84% | 1 | 1 | 0% | 1,125 | 632 | -44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +44 percentage points is the difference between those two pass rates over the 25 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.