AI sycophancy
AI sycophancy is a language model's tendency to tell people what they want to hear instead of what's accurate. A sycophantic model agrees with a user's mistaken claim, drops a correct answer when the user pushes back, or praises work because the user says they wrote it.
In customer service, that is an operational risk. Support agents regularly have to say things customers don't want to hear, like that the return window has closed or the fee was applied correctly. An agent that bends under pressure turns policy into a negotiation, and the customers who push hardest get the most exceptions.
Why language models become sycophantic
The leading explanation is training. In 2023, Anthropic's research on sycophancy found that five leading AI assistants were consistently sycophantic across four different tasks. The researchers traced part of the cause to human preference data, where answers that matched the user's views were more likely to be preferred.
Both human raters and the preference models trained on their choices sometimes picked a convincing sycophantic answer over a correct one. Training against those preferences can trade truthfulness for agreement.
The same pattern showed up in production. In April 2025, OpenAI rolled back a GPT-4o update after ChatGPT became noticeably more agreeable. In its follow-up on what it missed, OpenAI said a new reward signal based on thumbs-up and thumbs-down feedback had weakened the signal that kept sycophancy in check, and that its deployment evaluations didn't track sycophancy.
The logic is simple. If training rewards answers people like in the moment, and people like agreement, models learn to agree.
What sycophancy looks like in customer service
In support, sycophancy rarely looks like flattery. It looks like an agent that gives way.
A customer insists they were promised a refund, mentions years of loyalty, or just repeats the request with growing frustration. Eventually the agent grants the refund or waives the fee. Nothing about the facts changed. The conversation simply kept pushing toward yes.
Sometimes the agent confirms wrong details. The customer says the order went to a certain address, or that their plan includes a feature it doesn't, and the agent agrees.
The subtler case is dropping a correct answer. The agent explains that a charge is valid. The customer says that is wrong and asks for a recheck, and the agent apologizes and reverses itself with no new evidence. It's the support version of the "Are you sure?" challenge researchers use to test for the problem.
Agents can also accept a false premise, like a request to reverse a fee "since it was charged in error," and work inside that frame without checking. All of these cost money and consistency, and they hide in satisfaction scores, because a customer who got what they pushed for usually reports being happy.
How sycophancy differs from hallucination
The two fail in different ways. An AI hallucination is invented content, where the model makes up a policy or a tracking number. Sycophancy is deference. The model has the right answer, or could get it, and drops it because the user signaled a different preference.
That difference changes the fix. Hallucinations tend to shrink with better retrieval and grounding, because the model is filling a gap in what it knows. Sycophancy can happen even when the correct fact is right there in the context. The problem is how much weight the model gives the user's stance.
Pressure that looks authoritative is the worst case. SycEval, a 2025 study of three major assistants, found that rebuttals backed by citations were the most likely to flip a correct answer to a wrong one.
How to test an agent for sycophancy
Sycophancy is measured by applying pressure and watching what changes. The basic test asks a question with a known answer, pushes back, and checks whether the answer flips. SycEval separated regressive flips, where a right answer becomes wrong, from progressive ones, where a wrong answer becomes right.
For a support agent, the useful test is built from the deployment's own policies. Teams take real boundaries, like a closed return window, and simulate customers who insist, get upset, claim earlier promises, or cite rules that don't exist. The core metric is how often the agent grants what policy forbids.
The opposite number matters just as much. An agent tuned too hard against sycophancy becomes rigid. It refuses legitimate corrections, argues with customers who are right, and treats every pushback as manipulation. Progressive changes, where a customer supplies information that really does change the answer, are behavior worth keeping.
The goal is an agent that changes its mind only for evidence.
What AI sycophancy means for CX teams
Most teams don't control how the model was trained, so the reliable controls sit around it. The first is tying decisions to policy. If the refund rule lives in a policy the agent must consult, and the agent has to state which rule it applied, insistence alone has nothing to work with.
The second is checking facts with tools. Addresses, plan features, charge history and eligibility should come from system lookups, not the customer's description. An agent that reads the order record before confirming an address can't be talked into the wrong one.
The third is checking actions before they run. Intent validation compares a proposed action, like issuing a credit, against the verified situation and the relevant policy. Sensitive actions can also go to a human for approval, so pressure becomes a routed request instead of a granted one.
None of this requires a colder agent. It can stay warm in tone while its decisions stay tied to records and rules.
For a deeper dive, download Decagon's report on AI and the next generation of customer experience.

