Glossary

Speculative decoding

Speculative decoding is a way to make a language model write faster without changing what it writes. A small, fast drafter guesses the next several tokens, and the full model checks all of those guesses in one pass, keeping the ones it agrees with. The output follows the same probabilities the full model would have produced on its own.

It targets one specific slowdown. Models generate one token per step, and each step spends most of its time moving model weights through memory rather than computing. Checking several tokens takes about as long as generating one, so every correct guess is close to free.

In customer service, that means replies stream faster once they've started, which matters most for long answers and for voice agents that need text ready before the speech catches up.

How speculative decoding works

Each round starts with a draft. The drafter looks at the conversation so far and proposes a short run of tokens. The full model, called the target model, then reads the context plus all of the drafted tokens in a single pass, which gives it its own prediction at every drafted position.

Then it checks the draft from left to right. If the target rates a drafted token at least as likely as the drafter did, the token stays. If not, it is kept with a probability based on the gap between the two. At the first rejection, the target picks a replacement and the rest of the draft is thrown away.

The worst case is mild. Even if every guess is wrong, the round still produces one correct token from the full model. The original speculative decoding paper by Leviathan and colleagues reported two to three times faster generation on T5-XXL, with identical outputs and no retraining.

Why the speedup varies

The number that matters most is the acceptance rate, the share of drafted tokens the target keeps. When it's high, each check produces a long run of text. When it's low, the system pays for drafting and gets little back.

Acceptance depends on how predictable the text is. Boilerplate, structured output and text that repeats the input, such as order numbers or policy language, are easy to guess. Open-ended reasoning isn't.

The author of prompt lookup decoding reported an average 2.4x speedup on summarization and document-based question answering. The gain was much smaller on the first turn of a chat, when there was little earlier text to copy from.

Common types of drafters

The checking rule stays the same across versions. What changes is where the guesses come from.

The classic setup uses a separate small model that shares the target's vocabulary. It needs no changes to the main model, but both have to fit in GPU memory, and the small one has to behave enough like the big one to guess well. Methods like Medusa and EAGLE skip the second model and attach small prediction heads to the target itself, which means training those heads.

Some models are trained to predict several future tokens at once, and those extra layers can double as a built-in drafter. The simplest option, prompt lookup, uses no model at all. It finds the latest matching phrase in the context and proposes whatever followed it last time, which works well when a reply echoes the input.

Limits and trade-offs

Speculative decoding only speeds up the writing phase. The model still has to read the whole prompt before the first token appears, so time to first token doesn't change. Short replies, where that first wait is most of the total, see little benefit.

It also costs memory. A separate drafter takes GPU space that could otherwise hold model weights or KV cache, and drafted tokens take up cache space even when they're rejected. The vLLM documentation on speculative decoding says its version can improve inter-token latency in memory-bound inference, but warns that it will not reduce latency for every set of prompts or sampling settings.

Load matters too. As more requests are batched together, a GPU spends more of its time on real computation and has less idle capacity for checking drafts. The benefit shrinks with it.

What speculative decoding means for customer service

Most CX teams never switch this on themselves. Model providers and self-hosting teams enable it in the inference server, and from the outside it only shows up as faster streaming. Since the output stays the same, it doesn't trade away answer quality the way moving to a smaller model would.

The useful habit is to split latency into two numbers. One is the wait before the first token, which depends on prompt length, queueing and caching. The other is how fast the rest of the reply arrives, and that's the one speculative decoding improves.

A chat agent sending a few sentences mostly cares about the first number. An AI voice agent reading out a long return policy cares about both, and so does an agent filling in tool arguments that repeat account details already in the conversation. Measuring the two separately is the only way to know whether this technique will help.

For a deeper dive, download Decagon's report on AI and the next generation of customer experience.

Deliver the concierge experiences your customers deserve

Get a demo