Glossary

KV cache

A KV cache (key-value cache) is the working memory a language model keeps while it writes a response. Each attention layer turns every token it reads into a key and a value, and the cache holds on to them so the model doesn't have to reprocess the whole conversation each time it adds a word.

That shortcut is why long conversations are affordable to serve, and also why they're expensive. A support conversation can run for dozens of turns, with tool results and policy text piling up along the way. All of it sits in GPU memory while the request runs, which affects how fast the agent replies and how many customers one server can handle at once.

How a KV cache works

A model generates text in two phases. During prefill, it reads the whole prompt at once (instructions, conversation history, retrieved documents) and writes a key and value for every token into the cache. That work largely sets the time to first token, the pause before anything appears.

During decode, the model produces one token at a time. Each new token looks back at everything stored in the cache, and only its own key and value get added. Without the cache, every step would redo the prefill work for the whole sequence so far.

Decode steps can't run in parallel, since each one depends on the last. They spend most of their time reading stored state rather than doing math. That's why memory, more than raw compute, tends to be the limit.

Why the KV cache gets expensive

The cache grows with every token. The paper that introduced vLLM and PagedAttention worked through a 13-billion-parameter model where each token took about 800 KB of cache, so a single 2,048-token sequence could need up to 1.6 GB.

Two things push that number up. Longer conversations hold more tokens, so a chat that has gathered a long transcript and a few large tool outputs costs far more to keep alive than a fresh one. More users at once means more caches, since every request carries its own.

Model weights stay the same size no matter how busy the server is. The cache doesn't. When it runs out of room, the server has to queue new requests, pause running ones, or recompute what it threw away. Customers feel all three as lag.

How serving systems manage the cache

Older serving systems reserved one large block of memory per request, sized for the longest possible response. The vLLM authors found that only 20.4% to 38.2% of that memory held real token data. Their fix, PagedAttention, borrows from how operating systems handle memory.

The cache is split into small blocks that are handed out only as tokens arrive, and the authors reported two to four times higher throughput at the same latency.

Other techniques save memory in different ways. Prefix caching lets requests that start with the same text, like a shared system prompt, reuse the same cache blocks. Quantization stores keys and values at lower precision so more of them fit.

When space still runs short, servers evict or offload older blocks. The cost is recomputing them later, or some loss of accuracy if tokens are dropped for good.

KV cache versus prompt caching

The two are easy to mix up. The KV cache exists inside every request whether anyone configures it or not, and it's normally discarded when the response ends. Prompt caching is a provider feature that keeps the cache for a prompt's opening section alive afterward, so the next request starting with the same tokens can skip that prefill work.

The rules reflect how scarce that memory is. Anthropic's prompt caching documentation says a cache hit needs an identical prompt prefix, and entries last five minutes by default, refreshed each time they're used, with a one-hour option at extra cost. Change one token near the top, like a timestamp, and everything after it misses.

Prompt caching cuts the wait for the first token on repeated prompts. It does nothing to speed up the words that follow.

What the KV cache means for customer service

Teams that call hosted models never touch the cache directly. They still feel it in every latency and cost number. How a prompt is ordered decides whether caching can help, and how long a conversation runs decides how much memory each session holds.

A multi-turn conversation resends a growing transcript on every turn. It benefits most when stable content like instructions and tool definitions comes first and changing content comes last, so the shared opening stays cacheable. Summarizing old turns keeps the context window, and the cache with it, from growing without limit.

Voice is where this shows up first. A caller notices a pause after they stop talking much faster than a chat user notices a slow reply, and long prompts mean longer prefill.

For an ops team, the practical checks are simple. Is the static part of the prompt actually static? Are old turns being compacted? Is response time measured separately for the first token and for the rest of the reply?

For a deeper dive, download Decagon's report on AI and the next generation of customer experience.

Deliver the concierge experiences your customers deserve

Get a demo