Glossary

Durable execution

Durable execution is a way of running software so its progress is saved as it goes. If the server crashes or gets redeployed, or the workflow has to wait days for a reply, it picks up exactly where it left off. Temporal's documentation describes workflows that may run anywhere from a second to a year and still resume after a crash with their state intact.

It matters for AI agents because real support tasks rarely finish in one request. A refund might need a manager's sign-off. A damage claim might wait on a photo from the customer, or on a carrier updating a shipment. Those pauses can last hours or days, and the servers running the agent will restart several times in between.

How durable execution works

The common design keeps an event history, an ordered log of everything the workflow has done. Each time a step starts, finishes, fails or receives a message, the runtime adds an event. That log, not the memory of any single server, is the workflow's real state.

When a server dies mid-task, another one reads the log and reruns the workflow code from the top. This is called replay. Whenever the code reaches a step the log says is already done, the runtime hands back the recorded result instead of running it again, and execution continues from the point of failure.

Anything that touches the outside world, like calling an API, charging a card or calling a model, runs as its own step with retries and timeouts set in configuration.

Why ordinary agent loops lose work

A typical agent prototype keeps everything in memory. A request comes in, the model reasons, tools get called, and the conversation lives inside one process until the reply goes out. That's fine for tasks that finish in seconds. It breaks when the process disappears halfway through.

Microsoft's guidance on Durable Task for AI agents lists the ordinary causes, including restarts, deployments, scale-in events and transient failures. Without recovery, the agent starts over and pays again for model calls it already made.

Side effects are the bigger worry. An agent that crashed after issuing a refund, but before recording that it did, may issue the refund again on restart. Teams that build recovery by hand end up with state tables, checkpoint columns and cleanup jobs. Durable execution moves that plumbing into the runtime.

Durable execution for long-running agents

Microsoft describes two shapes of durable agent. In one, code sets the order of steps and calls the model at certain points. In the other, the model decides which tools to call and when the task is done. Both benefit, because each finished model call and tool result is saved, so a crash doesn't repeat completed work.

Waiting is the other half. A durable workflow can sleep, or wait for an outside event such as an approval, without holding a server while it waits. Inngest's engineering blog gives the example of a workflow that waits up to seven days for an approval to arrive by webhook, then carries on.

Common risks with durable execution

Replay comes with a rule. Workflow code has to make the same decisions every time it runs against the same history, so it can't read the clock, pick random numbers or call outside services directly. Those go into recorded steps instead.

Language models make this rule matter more. A model call can give a different answer each time, so it must be treated as a recorded step, never as part of the logic that gets replayed. It's also delicate to change workflow code while cases are still open, since an old run replayed through new code can take a path its history doesn't match.

Retries need care too. A step that timed out may have actually succeeded before it was retried. Refunds, emails and account changes need idempotency keys or a check-before-write at the receiving system, or a retry can repeat a real action. All of this adds infrastructure that short, stateless tasks don't need.

What durable execution means for customer service

In support, the pattern fits any case that spans time. An agentic workflow for a damaged item can collect details, open a return and then pause until the warehouse confirms receipt before refunding. A large billing adjustment can wait for human-in-the-loop approval without using compute. When the customer writes back two days later, the case resumes with its full context.

The event history doubles as an AI audit trail. It records every step, result, retry and wait in a case, which makes reviewing a disputed refund far easier than piecing it together from scattered logs.

For teams planning agent orchestration across several agents, a good first question is which parts of the work have to survive a restart. Anything that spends money, changes an account or waits on a person usually does.

For a deeper dive, download Decagon's guide to agentic AI for customer experience.

Deliver the concierge experiences your customers deserve

Get a demo