Glossary

Graduated autonomy

Graduated autonomy is the practice of widening what an AI agent can do without human approval in steps, each one earned by proving reliable at the step below. The agent's capability doesn't change between steps. What changes is how much it's trusted to do alone.

It replaces a yes-or-no decision with a ladder that goes both ways. That suits customer service, where the same agent might be perfectly safe answering policy questions and risky issuing refunds. Teams can put an agent to work early on low-risk tasks, watch how it behaves in production, and widen its authority only where the evidence supports it.

How graduated autonomy works

Most setups use four rungs. At suggest only, the agent drafts a reply or proposes an action and a person carries it out. Every proposal gets a human verdict, so this rung is as much about collecting data as about getting work done.

Next is act with approval. The agent executes, but only after someone signs off. Throughput is capped by reviewer speed, so teams try not to stay here longer than they need to.

Then comes act within limits: refunds under a cap, edits to certain fields, messages from approved templates. Those limits have to be enforced in code, credentials or allow-lists. A limit written into a prompt is a suggestion, not a control.

At the top, the agent handles a whole category of work end to end, and people review samples and audits instead of single actions. AWS's reference architecture for graduated autonomy follows the same shape, moving from read-only probation up to full access with after-the-fact audit only.

What evidence earns a promotion

Promotion criteria should be written down before the agent starts. Time served and enthusiasm don't count. Offline evals against a golden set show the agent handles the expected cases, including edge cases and requests it should refuse.

Coverage matters more than volume. Hundreds of approvals of near-identical requests prove the same thing over and over and say little about the cases that cause incidents.

Production signals carry the most weight. The human override rate tracks how often reviewers reject or change what the agent proposed. A low, steady rate across varied traffic is the clearest sign an action is ready for less oversight, and audits of completed actions catch what reviewers missed.

AWS scores each agent over a rolling window of 50 actions and treats safety as a separate floor that strong scores elsewhere can't offset. It also requires a margin above a tier's floor before promotion, a technique called hysteresis, so an agent sitting at the boundary doesn't flip back and forth.

Why risk tiers should be set per action

Autonomy works best when it's assigned per action, not per agent. A password reset for a verified customer is low risk and easy to undo. A refund above a set threshold moves money that's awkward to recover. The same agent can sit at the top rung for one and at the approval rung for the other.

Data Science Dojo's traffic-light model sorts actions into green, yellow and red using three questions. Can it be undone? What's the downside of a wrong answer? Does a regulation require a person to authorize it?

Green actions run without review. Yellow actions run inside guardrails like budget limits, or drop back to propose-only mode. Red actions need a person to authorize them. When an action is hard to classify, the advice is to start one tier more cautious than seems necessary and let the evidence move it.

Common risks in graduated autonomy

Demotion is the part teams most often neglect. Models get updated, prompts change, upstream APIs shift and the mix of incoming requests drifts. Any of these can weaken an agent that earned its rung months ago.

Demotion rules should be written alongside promotion rules and fire automatically on signals like a drop in eval scores, a rising override rate or an unfamiliar pattern of tool calls. AWS demotes an agent immediately when safety falls below its floor. Promotion is a considered decision. Demotion should work like a circuit breaker.

The approval rung has its own trap. Reviewers who see a long run of correct proposals start approving without reading, and the override rate falls for the wrong reason. Sampling approved actions for an independent audit, and rotating reviewers, keeps that signal honest.

Any change to the agent itself should reset the evidence, the same way a canary deployment treats a new version as unproven until it performs.

What graduated autonomy means for customer service

For support teams, the biggest benefit is that there's no all-or-nothing launch. An agent earns authority one action type at a time, and each step leaves a record that can be reviewed later.

Making it work takes three things: a list of actions with a tier for each, a named owner who can demote an agent without calling a meeting, and limits enforced in the systems that actually carry out the actions. The same setup that grants autonomy can take it back the moment performance slips.

Introducing Simulations | Decagon Dialogues '25

For a deeper dive, download Decagon's guide to agentic AI for customer experience.

Deliver the concierge experiences your customers deserve

Get a demo