Glossary

Voice agent observability

Voice agent observability is the practice of instrumenting an AI voice agent so any call can be replayed step by step. That covers the audio that came in, when the system decided the caller had stopped talking, what the speech recognizer heard, what the model decided, what the agent said back, and how long each step took. It's a branch of AI observability built for live audio instead of text.

The need comes from how voice agents fail. A text agent's mistakes usually show up in the transcript. A voice agent can produce a perfectly reasonable transcript and still deliver a bad call, because the real problem was a noisy line, a long pause before every reply, or an agent that kept talking over the caller.

How voice agent observability works

The core unit is the trace, a record of one call laid out as a tree. There's a root for the call, a branch for each turn, and a timed span inside each turn for every stage of the pipeline.

A typical voice agent needs five layers covered. Telephony and audio data, like packet loss, jitter and audio levels, separates a carrier problem from an agent problem. Turn-taking data records each voice activity, endpointing and interruption decision with a timestamp.

Recognition data includes the transcript, confidence scores and the partial transcripts as they streamed in. Low confidence on something like an account number often comes right before a mistake. Sampled audio can also be checked against reference transcripts to track word error rate over time.

The last two layers are reasoning and speech output. Reasoning covers model calls, time to first token and tool calls with their results. Speech output covers the text sent for synthesis, time to first audio, and whether playback finished or got cut off.

Because every span has timestamps, per-stage latency comes straight from the trace. That matters because voice latency is a sum. Twilio's guide to voice agent latency measures the mouth-to-ear turn gap, from the moment the caller stops speaking to the moment the reply reaches their ear, and counts at least ten network trips behind a single response.

Why text tracing misses voice failures

Standard LLM tracing starts at the model call. Much of what goes wrong in a voice agent happens before the first token or after the last one.

Audio problems are invisible in text. If the line is noisy, recognition gets worse and the transcript just shows the wrong words. Without audio metrics on the same trace, it looks like the model misunderstood, and someone spends a day rewriting the prompt.

Gaps between stages are hidden too. The model can answer quickly while the caller still waits, because endpointing held the turn open or a slow tool call blocked speech. Twilio points to end-of-turn detection as often the slowest part, and warns that a sudden jump in tail latency near the silence timeout is a sign it's misfiring.

Turn-taking errors leave almost no mark in text. When the agent mistakes a cough for an interruption, it stops, waits and starts again. The transcript shows a slightly repetitive agent. An audio-aware trace shows an interruption event with no matching speech, which is the real cause.

How traces connect to call outcomes

Stage-level data gets useful once it's joined to what happened on the call. For a support team the outcomes are concrete. The issue was resolved, the call went to a human, the caller hung up mid-conversation, or they called back an hour later.

With outcomes attached, teams can ask better questions. Do calls with low recognition confidence early on transfer more often? Does abandonment cluster after long pauses?

Common trade-offs in instrumenting voice agents

Privacy comes first. Full recordings and transcripts are the most useful data, and they routinely contain payment details, health information and the caller's voice itself, which some jurisdictions treat as biometric data. Teams have to decide what to redact and how long raw audio is kept.

Overhead is next. Instrumentation in the live path adds delay where a few hundred milliseconds is noticeable, so spans should be sent asynchronously.

Then there's volume. Frame-level audio metrics and streaming transcripts produce far more data than a chat. Sampling cuts cost but works against the main use case, which is digging into the one bad call a customer complained about. A common compromise is a light trace for every call and full audio for a subset.

Attribution across vendors is the last. A voice stack often mixes a telephony provider, a recognizer, a model provider and a speech service, each with its own logs and clock. Without a shared call ID and synced timestamps, the latency numbers won't add up.

What voice agent observability means for support teams

Good tracing changes quality review. Instead of sampling calls at random, reviewers can pull the ones where a stage misbehaved, like the longest turn gaps on calls that ended in a transfer, and play each back next to its trace.

Shared formats keep this portable. The OpenTelemetry GenAI semantic conventions name standard operations for invoking an agent, calling a model and executing a tool. The agent span specification is still marked as in development and doesn't cover telephony or speech, so voice teams usually add their own attributes for audio quality and turn-taking.

The practical test is simple. Given any call a customer complains about, an engineer should be able to hear it, see every stage's decision and timing on one timeline, and find the stage that caused the problem in minutes.

The three pillars of effective voice AI in CX | Jesse Zhang | Decagon Dialogues '25

‍

For a deeper dive, download Decagon's guide to the 10 principles of a production-grade voice AI agent.

Deliver the concierge experiences your customers deserve

Get a demo