Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Automate computer tasks using Anthropic's computer use API — Claude controls mouse, keyboard, and screen. Use when automating GUI workflows, browser automation without CSS selectors, desktop app testing, or any task that requires seeing and interacting with a graphical interface.
.claude/skills/terminalskills-claude-computer-use/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-23 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-15 | ✓→✗ | ▼ Worse | 170% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 93% | 0% |
Claude's computer use capability lets the model see your screen (via screenshots) and control mouse and keyboard. Unlike Playwright or Selenium, it requires no selectors — Claude navigates visually, the same way a human would. It's ideal for legacy software, complex multi-app workflows, and tasks that are hard to automate programmatically.
⚠️ Always run in a sandboxed environment (Docker, VM, or dedicated machine). Never run on a machine with access to sensitive accounts or production systems.
bashpip install anthropic pillow pyautogui # Linux: also install scrot or gnome-screenshot for screenshots apt-get install -y scrot
Use Docker for safety:
dockerfileFROM ubuntu:22.04 RUN apt-get update && apt-get install -y \ python3 python3-pip \ xvfb x11vnc \ scrot \ chromium-browser \ && pip install anthropic pillow pyautogui ENV DISPLAY=:1 CMD ["Xvfb", ":1", "-screen", "0", "1280x800x24"]
bashdocker build -t computer-use-sandbox . docker run -d --name sandbox \ -e ANTHROPIC_API_KEY=$ANTHROPIC_API_KEY \ -p 5900:5900 \ # VNC to monitor what's happening computer-use-sandbox
The computer use beta provides three tools:
pythonTOOLS = [ { "type": "computer_20241022", # Screen + mouse + keyboard "name": "computer", "display_width_px": 1280, "display_height_px": 800, "display_number": 1 }, { "type": "bash_20241022", # Run shell commands "name": "bash" }, { "type": "text_editor_20241022", # View and edit files "name": "str_replace_editor" } ]
pythonimport subprocess import base64 from pathlib import Path def take_screenshot() -> str: """Take a screenshot and return as base64 PNG.""" path = "/tmp/screenshot.png" subprocess.run(["scrot", path], check=True) return base64.standard_b64encode(Path(path).read_bytes()).decode("utf-8") def get_screenshot_block() -> dict: """Return an image content block for the API.""" return { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": take_screenshot() } }
pythonimport pyautogui import time def execute_computer_action(action: dict) -> str: """Execute a computer action from Claude's tool use response.""" action_type = action["action"] if action_type == "screenshot": return take_screenshot() elif action_type == "mouse_move": pyautogui.moveTo(action["coordinate"][0], action["coordinate"][1]) return "Mouse moved" elif action_type == "left_click": pyautogui.click(action["coordinate"][0], action["coordinate"][1]) time.sleep(0.5) return "Clicked" elif action_type == "double_click": pyautogui.doubleClick(action["coordinate"][0], action["coordinate"][1]) time.sleep(0.5) return "Double-clicked" elif action_type == "right_click": pyautogui.rightClick(action["coordinate"][0], action["coordinate"][1]) return "Right-clicked" elif action_type == "type": pyautogui.write(action["text"], interval=0.02) return "Typed text" elif action_type == "key": pyautogui.hotkey(*action["key"].replace("+", " ").split()) return f"Pressed {action['key']}" elif action_type == "scroll": direction = action.get("direction", "down") clicks = action.get("clicks", 3) pyautogui.scroll(-clicks if direction == "down" else clicks) return f"Scrolled {direction}" else: return f"Unknown action: {action_type}" def execute_bash(command: str) -> str: """Execute a bash command and return output.""" result = subprocess.run( command, shell=True, capture_output=True, text=True, timeout=30 ) return result.stdout + result.stderr
pythonimport anthropic client = anthropic.Anthropic() def run_computer_use_agent(task: str, max_steps: int = 20) -> str: """ Run Claude computer use to complete a task. Args: task: Natural language description of what to do max_steps: Safety limit on number of actions Returns: Claude's final response describing what was done """ # Start with screenshot so Claude sees current state messages = [ { "role": "user", "content": [ {"type": "text", "text": task}, get_screenshot_block() ] } ] for step in range(max_steps): print(f"\n[Step {step + 1}/{max_steps}]") response = client.beta.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=4096, tools=TOOLS, messages=messages, betas=["computer-use-2024-10-22"] ) # Add Claude's response to history messages.append({"role": "assistant", "content": response.content}) # Check if Claude is done if response.stop_reason == "end_turn": final_text = next( (b.text for b in response.content if hasattr(b, "text")), "" ) print(f"\n✅ Task complete: {final_text}") return final_text # Process tool uses if response.stop_reason == "tool_use": tool_results = [] for block in response.content: if block.type != "tool_use": continue print(f" → Tool: {block.name}, Action: {block.input.get('action', block.name)}") if block.name == "computer": result = execute_computer_action(block.input) # Always include a fresh screenshot after actions tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": [ {"type": "text", "text": result}, get_screenshot_block() ] }) elif block.name == "bash": output = execute_bash(block.input["command"]) print(f" ← Output: {output[:200]}") tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": output }) elif block.name == "str_replace_editor": # Handle file viewing/editing tool_results.append({ "type": "tool_result", "tool_use_id": block.id, "content": "File operation completed" }) messages.append({"role": "user", "content": tool_results}) return "Max steps reached without completion"
For sensitive tasks, add approval checkpoints:
pythonSENSITIVE_ACTIONS = ["left_click", "type", "key"] REQUIRE_APPROVAL_FOR = ["submit", "delete", "purchase", "send"] def execute_with_approval(action: dict, task_context: str) -> str: """Pause and ask for human approval on potentially dangerous actions.""" action_type = action.get("action", "") text = action.get("text", "") # Check if this looks like a high-risk action is_risky = any(word in task_context.lower() for word in REQUIRE_APPROVAL_FOR) if is_risky and action_type in SENSITIVE_ACTIONS: print(f"\n⚠️ APPROVAL REQUIRED") print(f" Action: {action_type}") print(f" Details: {action}") approval = input(" Approve? (y/n): ") if approval.lower() != "y": raise RuntimeError("Action denied by user") return execute_computer_action(action)
python# Fill out a web form result = run_computer_use_agent( "Open Chrome, go to https://forms.example.com/application, " "fill in Name='John Smith', Email='john@smith.com', and submit the form." ) # Extract data from a desktop app result = run_computer_use_agent( "Open the Excel file at /home/user/data.xlsx, " "copy all values from column B rows 2-50, and save them to /tmp/extracted.txt" ) # Automate a repetitive workflow result = run_computer_use_agent( "In the open CRM application, find all contacts with status 'Follow Up', " "change their status to 'Active', and export the list to CSV." )
max_steps to prevent runaway loopsmax_steps (10–30 depending on task complexity)time.sleep) after clicks to let UI render before next screenshot| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 8,415 | 11,701 | +39% | 1 | 1 | 0% | 1,818 | 5,027 | +177% | 0 | 0 | — |
case-04 | pass→pass | 11,863 | 9,086 | -23% | 1 | 1 | 0% | 2,287 | 4,416 | +93% | 0 | 0 | — |
case-01 | pass→pass | 22,890 | 14,007 | -39% | 1 | 1 | 0% | 5,140 | 5,788 | +13% | 0 | 0 | — |
case-02 | fail→fail | 12,730 | 9,651 | -24% | 1 | 1 | 0% | 2,826 | 4,457 | +58% | 0 | 0 | — |
case-05 | pass→pass | 12,701 | 11,202 | -12% | 1 | 1 | 0% | 2,610 | 4,927 | +89% | 0 | 0 | — |
case-06 | pass→pass | 6,798 | 9,249 | +36% | 1 | 1 | 0% | 1,374 | 4,491 | +227% | 0 | 0 | — |
case-07 | fail→pass | 10,643 | 7,753 | -27% | 1 | 1 | 0% | 1,968 | 4,081 | +107% | 0 | 0 | — |
case-08 | fail→fail | 10,639 | 7,964 | -25% | 1 | 1 | 0% | 2,161 | 4,178 | +93% | 0 | 0 | — |
case-09 | pass→pass | 13,108 | 12,193 | -7% | 1 | 1 | 0% | 2,324 | 4,652 | +100% | 0 | 0 | — |
case-10 | pass→pass | 12,982 | 11,083 | -15% | 1 | 1 | 0% | 2,507 | 5,072 | +102% | 0 | 0 | — |
case-11 | pass→pass | 7,081 | 3,820 | -46% | 1 | 1 | 0% | 1,259 | 3,297 | +162% | 0 | 0 | — |
case-12 | pass→pass | 11,683 | 6,663 | -43% | 1 | 1 | 0% | 2,013 | 3,881 | +93% | 0 | 0 | — |
case-13 | pass→pass | 8,282 | 6,427 | -22% | 1 | 1 | 0% | 1,587 | 3,769 | +137% | 0 | 0 | — |
case-14 | pass→pass | 16,251 | 9,401 | -42% | 1 | 1 | 0% | 2,120 | 4,331 | +104% | 0 | 0 | — |
case-15 | pass→fail | 6,391 | 1,870 | -71% | 1 | 1 | 0% | 1,055 | 2,852 | +170% | 0 | 0 | — |
case-16 | pass→pass | 6,053 | 2,234 | -63% | 1 | 1 | 0% | 1,137 | 2,986 | +163% | 0 | 0 | — |
case-17 | pass→pass | 16,121 | 23,690 | +47% | 1 | 1 | 0% | 1,874 | 3,994 | +113% | 0 | 0 | — |
case-18 | pass→pass | 8,796 | 3,507 | -60% | 1 | 1 | 0% | 1,617 | 3,194 | +98% | 0 | 0 | — |
case-19 | pass→pass | 13,049 | 11,061 | -15% | 1 | 1 | 0% | 2,363 | 4,716 | +100% | 0 | 0 | — |
case-20 | fail→fail | 7,632 | 2,658 | -65% | 1 | 1 | 0% | 1,279 | 3,062 | +139% | 0 | 0 | — |
case-21 | pass→pass | 5,097 | 2,044 | -60% | 1 | 1 | 0% | 1,047 | 2,999 | +186% | 0 | 0 | — |
case-22 | fail→pass | 17,135 | 9,199 | -46% | 1 | 1 | 0% | 2,980 | 4,216 | +41% | 0 | 0 | — |
case-23 | fail→pass | 11,373 | 2,995 | -74% | 1 | 1 | 0% | 2,144 | 3,137 | +46% | 0 | 0 | — |
case-24 | pass→pass | 3,146 | 2,233 | -29% | 1 | 1 | 0% | 496 | 2,998 | +504% | 0 | 0 | — |
case-25 | pass→pass | 10,970 | 7,162 | -35% | 1 | 1 | 0% | 1,699 | 3,759 | +121% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +8 percentage points is the difference between those two pass rates over the 25 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.