Glossary

Computer-use agent

A computer-use agent is an AI agent that operates software the way a person does, by looking at the screen and using a mouse and keyboard instead of calling APIs. It takes a screenshot, decides what to click or type, does it, then looks again. Any application with a screen becomes reachable, including ones that were never built to be automated.

The appeal is reach. Back offices are full of software with no usable API, like old billing consoles, carrier portals and internal admin tools. A computer-use agent can work those screens directly. The catch is that it's slower, less reliable and harder to secure than an agent calling structured tools, so in customer service it's best kept as a fallback.

How a computer-use agent works

The agent runs a perceive-and-act loop. First, the surrounding application captures a screenshot and passes it to the model. OpenAI's Computer-Using Agent works from raw pixels, though some systems also read a page's accessibility data to find buttons more easily.

Next, the model looks at the current screen, the task and what it has already tried, then picks an action. In Anthropic's computer use tool, targets are given as pixel coordinates on the screenshot. The application, not the model, then carries out the click, keystroke or scroll.

Finally, the agent takes a fresh screenshot to check that the action did what it meant to. A click that landed on the wrong row only gets caught here. The cycle repeats until the task is done, judged impossible, or a step limit is hit.

Why computer use is slow and error-prone

Working through pixels is harder than calling a function. OpenAI reported that its agent completed 38.1% of tasks on OSWorld, a benchmark of full desktop tasks, against a human score of 72.4%. Pop-ups, slow page loads, small layout changes and look-alike buttons all cause errors a structured call would never hit.

It's slow, too. Researchers behind the OSWorld-Human benchmark found agents taking tens of minutes on tasks a person finishes in a few. In one case, changing the line spacing of two paragraphs took an agent 12 minutes.

Every step is a full model call with an image attached. That makes computer use expensive as well as slow.

Common risks with computer-use agents

A computer-use agent reads everything on the screen, and anything it reads can try to give it orders. Anthropic warns that its model may follow instructions found in webpages or images even when they conflict with the user's. That's prompt injection arriving as pixels, which is harder to filter than plain text.

The risk is real for support work. A hostile line in a ticket or an email preview could redirect an agent that's logged into an admin console. Because the agent holds whatever permissions that login has, scoping the account narrowly matters as much as any model-level defense.

Anthropic's suggested precautions are mostly about isolation. It recommends a dedicated virtual machine or container with minimal privileges, keeping login details and sensitive data out of the model's reach, and limiting internet access to an allowlist of domains. It also advises having a person confirm actions with real consequences, such as financial transactions.

Computer use versus tool calling

Tool calling gives the model a defined function with set inputs and outputs. The call either works or returns a clear error, it runs fast, and permissions can be set per function. Computer use gives the model a picture and a pointer.

For any system with an API, the structured path wins on speed, cost and auditability. Computer use earns its place where that path doesn't exist. A legacy order screen with no integration, or a partner portal that only offers a web login, are reasonable candidates, especially for low-volume tasks where a slow automated step still beats a manual one.

What computer-use agents mean for customer service

Most support teams won't run a computer-use agent as their main agent. It's more often one capability inside a larger system, and the last one reached for. The agent harness runs the loop, enforces step and time limits, and records every screenshot and action for review.

A sensible setup looks like this. The agent looks up an order through an API, switches to a computer-use subtask only to update one field in a legacy console, then confirms the change with a separate read before replying to the customer. The task is narrow, the screen is known, and a person is ready to take over if the check fails.

Before adopting it, ops teams should ask a few plain questions. Which systems truly lack an API? What can the logged-in account change? Who reviews the recordings when something goes wrong?

For a deeper dive, download Decagon's guide to agentic AI for customer experience.

Deliver the concierge experiences your customers deserve

Get a demo