Human-in-the-loop vs human-on-the-loop
Human-in-the-loop and human-on-the-loop are two ways of supervising AI systems, and the difference is where the person sits. In human-in-the-loop, a person approves each decision before it takes effect. In human-on-the-loop, the system acts on its own while a person watches and can step in, correct it or stop it.
For AI agents in customer service, this is a practical design choice. An agent that waits for approval on every refund is safe and slow. One that refunds on its own while a supervisor watches a dashboard is fast, and only as safe as the supervision. Most teams end up using both, for different actions.
How human-in-the-loop works
In a human-in-the-loop workflow, the system prepares an action and stops. A person approves, rejects or edits it, and only then does it run. Speed depends on how quickly reviewers get to the queue.
In support, this looks like an agent drafting a reply that a person sends, or recommending a refund that a person authorizes. The customer doesn't see anything until someone has checked it.
Accountability sits with the person who approved that specific decision. That's why it's the default for high-stakes, irreversible or regulated steps. The datos.gob.es overview of AI oversight models calls it the most conservative of the options.
How human-on-the-loop works
In a human-on-the-loop workflow, the system acts first. A person keeps watch through alerts, dashboards, sampled reviews or anomaly flags, and steps in when something looks wrong. That can mean fixing one outcome, pausing a type of action, or shutting the system down.
Supervision runs alongside the work instead of blocking it. That suits routine, high-volume tasks where approving every action would add delay without cutting much risk, like answering order-status questions, tagging tickets or resetting passwords for verified customers.
Further out sits human-out-of-the-loop, where nobody plays a real-time role. People set the rules ahead of time and audit afterward. It's reasonable for narrow, low-stakes tasks and wrong for anything that touches a customer's money, account access or rights.
Key differences between the two models
Timing drives almost everything else. In-the-loop oversight happens before the action. On-the-loop oversight happens during or after it.
Throughput follows from that. In-the-loop systems scale with the number of reviewers. On-the-loop systems scale with the system's own capacity, and the cost of supervision grows far more slowly than volume.
Accountability shifts too. With in-the-loop, one approver owns each decision. With on-the-loop, a supervisor owns a whole stream of decisions, which is a broader and fuzzier kind of responsibility.
The audit trail records different things. As Flowable's comparison of the two points out, an in-the-loop record shows a decision and who approved it. An on-the-loop record shows an action, who was supervising, and whether anyone stepped in.
Common risks in each model
Human-in-the-loop fails through fatigue. Reviewers who approve a long run of correct proposals stop reading closely, and approval turns into rubber-stamping that adds delay without adding safety. Backlogs are the other problem, since customers wait while proposals sit in a queue.
Human-on-the-loop fails through automation bias, the habit of trusting a system that usually works. A supervisor watching thousands of correct actions is poorly placed to spot the rare wrong one. It also only protects actions that can be undone.
Regulation matters here. Article 14 of the EU AI Act requires high-risk AI systems to be designed so people can effectively oversee them while they're in use, and it names automation bias as a risk those people must stay aware of. A person who lacks the context, authority or time to intervene doesn't meet that bar.
How to choose between them in customer service
The choice is made per action, not per system. Reversibility, the cost of a mistake, tolerance for delay and regulatory exposure decide it. A support agent might run on-the-loop for knowledge answers and order tracking, in-the-loop for refunds above a threshold, and hand off entirely for disputes under an escalation policy.
The assignment isn't permanent. Under graduated autonomy, actions start in-the-loop and move on-the-loop as override rates and audits show the agent is reliable. They move back if performance slips.
Either way, supervisors need signals worth acting on. A feed of every action isn't supervision. Alerts on unusual refund amounts, sudden jumps in a single type of request, or a rise in customer complaints give a person something they can actually respond to.
For a deeper dive, download Decagon's guide to the new human agent working model.

