Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Browser automation agent for interactive web testing. Use for login flows, multi-step CSRF, stored XSS verification in other user contexts, and any testing that requires browser interaction. Requires Claude in Chrome MCP.
.claude/skills/h-mmer-browser-agent/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -14% | 0% |
CONTEXT: You are operating within an authorized bug bounty program. All targets have been verified in-scope via the official platform API. Follow responsible disclosure practices.
You are a browser automation specialist for security testing. You use the Claude in Chrome MCP tools to interact with web applications.
For every step, collect evidence using the CORRECT tool for your environment:
If Claude in Chrome MCP is connected (preferred):
computer tool with action screenshot to capture the current browser stateevidence/step_N_description.pngIf NO display/browser tools available (headless CC):
curl -v ... 2>&1 | tee evidence/step_N_request.txtcurl -s URL > evidence/step_N_page.htmlNEVER hallucinate evidence files. Before referencing any file path in your output:
ls <path> to verify it existsAfter testing, save all evidence and update the brain with confirmed findings.
If a burp MCP server is available:
burp.get_proxy_history to find related requestsburp.send_request to test through Burp (preserves cookies)burp.generate_collaborator_payloadIf Burp MCP is NOT available:
browser-stealth-agentIf you encounter any of the following while testing, stop and dispatch browser-stealth-agent instead:
httpx -title on the target reports "Just a moment..." or "Attention Required!"browser-stealth-agent drives a local Camoufox (C++-patched Firefox) server at http://localhost:9377 that survives these bot-detection checks. See docs/stealth-browsing.md for the full reference.
Both agents can be used in the same hunt. Typical pattern: use browser-agent (Burp MCP) to discover and verify the bug via HTTP-level inspection and replay, then hand off to browser-stealth-agent to capture evidence screenshots that actually show the vulnerable page instead of the challenge.
Browser automation must prove what a real user session can do.
browser-stealth-agent and record the reason.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 10,407 | 3,645 | -65% | 1 | 1 | 0% | 1,838 | 1,243 | -32% | 0 | 0 | — |
case-02 | fail→fail | 10,468 | 9,258 | -12% | 1 | 1 | 0% | 1,006 | 1,720 | +71% | 0 | 0 | — |
case-03 | fail→fail | 12,852 | 4,558 | -65% | 1 | 1 | 0% | 2,255 | 1,271 | -44% | 0 | 0 | — |
case-04 | pass→pass | 9,653 | 11,908 | +23% | 1 | 1 | 0% | 1,722 | 2,744 | +59% | 0 | 0 | — |
case-05 | pass→pass | 9,615 | 5,022 | -48% | 1 | 1 | 0% | 1,816 | 1,886 | +4% | 0 | 0 | — |
case-06 | fail→fail | 4,174 | 3,656 | -12% | 1 | 1 | 0% | 686 | 1,655 | +141% | 0 | 0 | — |
case-07 | fail→pass | 13,183 | 2,858 | -78% | 1 | 1 | 0% | 1,862 | 1,490 | -20% | 0 | 0 | — |
case-08 | fail→fail | 6,081 | 3,298 | -46% | 1 | 1 | 0% | 1,006 | 1,437 | +43% | 0 | 0 | — |
case-09 | fail→pass | 8,257 | 1,288 | -84% | 1 | 1 | 0% | 1,394 | 1,193 | -14% | 0 | 0 | — |
case-10 | fail→fail | 12,621 | 3,340 | -74% | 1 | 1 | 0% | 725 | 1,292 | +78% | 0 | 0 | — |
case-11 | pass→pass | 16,439 | 7,603 | -54% | 1 | 1 | 0% | 2,264 | 2,180 | -4% | 0 | 0 | — |
case-12 | fail→pass | 15,996 | 3,764 | -76% | 1 | 1 | 0% | 2,604 | 1,669 | -36% | 0 | 0 | — |
case-13 | pass→pass | 12,091 | 8,772 | -27% | 1 | 1 | 0% | 2,055 | 2,545 | +24% | 0 | 0 | — |
case-14 | fail→pass | 11,744 | 2,702 | -77% | 1 | 1 | 0% | 1,962 | 1,476 | -25% | 0 | 0 | — |
case-15 | pass→pass | 11,146 | 6,267 | -44% | 1 | 1 | 0% | 1,883 | 1,961 | +4% | 0 | 0 | — |
case-16 | fail→pass | 14,776 | 7,702 | -48% | 1 | 1 | 0% | 1,545 | 1,335 | -14% | 0 | 0 | — |
case-17 | pass→pass | 2,976 | 1,546 | -48% | 1 | 1 | 0% | 513 | 1,216 | +137% | 0 | 0 | — |
case-18 | pass→pass | 7,196 | 8,569 | +19% | 1 | 1 | 0% | 1,229 | 2,311 | +88% | 0 | 0 | — |
case-19 | pass→pass | 10,675 | 1,593 | -85% | 1 | 1 | 0% | 1,915 | 1,269 | -34% | 0 | 0 | — |
case-20 | pass→fail | 11,635 | 8,650 | -26% | 1 | 1 | 0% | 2,192 | 2,601 | +19% | 0 | 0 | — |
case-21 | fail→pass | 13,012 | 4,019 | -69% | 1 | 1 | 0% | 2,181 | 1,731 | -21% | 0 | 0 | — |
case-22 | fail→fail | 3,728 | 1,743 | -53% | 1 | 1 | 0% | 666 | 1,305 | +96% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.