Claude Integrates Computer Use into SDK: No More Manual Click Mapping Loops Required
Anthropic has built the computer use and browser use toolset into its Python and TypeScript SDKs. The SDK handles the loop internally, and the API directly tells developers where Claude wants to click. All developers need to do is implement a driver. This update lowers the entry barrier, but increases security complexity.
Anthropic has directly integrated the computer use and browser use toolset into its Python and TypeScript SDKs.
Previously, to let Claude control a computer or browser, you had to write the loop yourself: receive the action intent output by the model, map "click this coordinate" or "enter this text" to commands for your automation tool, then execute them. Now the SDK runs this loop on its own. The API directly tells you where Claude wants to click and what it wants to input, and the SDK handles forwarding actions to the driver layer. You only need to implement a driver.
The official team provides a [quickstart](https://github.com/anthropics/claude-quickstarts/tree/main/computer-toolset) for a VNC driver. It's just a few hundred lines of code and runs on an ephemeral desktop. The full documentation is available [here](https://platform.claude.com/docs/en/agents-and-tools/tool-use/browser-use-sdk). The SDK's abstract base class handles tool schemas, loop management, and disabling unimplemented methods. You only need to implement methods like `screenshot`, `left_click`, and `type`, then pass your instance to the tool runner.
This change lowers the barrier to building an agent from "writing the loop" to "writing the driver". Writing the loop isn't hard, but handling message formatting, retry logic, and failure recovery is all repetitive work. As one developer put it: "Everyone is writing the same loop over and over again."
Third parties have moved quickly to adapt. Browser Use, Browserbase, E2B, and Daytona all provide compatible drivers. Browserbase's Stagehand even released an integration that lets you spin up a browser instance in just a few lines of code and pass it to Claude's tool runner.

However, when letting an external model control a system-level browser driver, the engineering bottleneck has shifted from state mapping to sandbox isolation. The official documentation devotes a lot of space to security: URL policies, egress rules, disabling the `javascript:` and `file:` protocols, and isolating browser hosts. Content on the screen guides the model's next step, and an unhandled UI popup can trap the agent in an infinite retry loop, burning through a huge amount of tokens. Some developers have commented that the current engineering bottleneck is sandbox isolation, not state mapping.
Community reactions are split. Some are relieved: "No more writing click and key mapping manually." Some joked: "Claude now clicks 'Accept All Cookies' with the same conviction a 14-year-old has when installing a Minecraft mod." Others asked "How does the SDK know when the loop ends" — the tool runner exits once the model stops calling tools. Some also pointed out the lack of Windows support: "Still no Windows."
Another notable detail: some people noticed that the model in the demo video is Haiku 5.5, a newly released cheap small model. Even affordable small models can now use computer use. As someone put it: "Cheap agents are going to be clicking everywhere soon."
Giving an agent control of your browser sounds convenient, until it accidentally places an order for you. Fortunately, the SDK provides a `confirm` callback that lets you block it before it performs sensitive operations.
发布时间: 2026-10-08 20:52