Skip to main content

How we work

Running an AI agent in production

A live agent needs about an hour a month, and it needs that hour to have a name attached. The failure mode is not a crash — it is drift: the business changes, the agent does not, and nobody notices for a quarter. Three numbers and a monthly review prevent most of it.

By Gorden WübbeAutomation, agents and search visibilityUpdated 10 August 2026

Agents do not crash. They drift.

The expected failure is technical: something breaks, an alarm goes off, someone fixes it. That happens rarely, and when it does it is obvious and quick.

The actual failure is quieter. The business changes and the agent does not. Prices go up in March and it quotes February’s. A product is withdrawn and it keeps offering it. A colleague leaves and escalations still route to their inbox.

Nothing broke. Every check is green. And a customer is being told something that stopped being true eleven weeks ago.

The three numbers

A dashboard with twenty metrics gets opened twice and then never again. Three get looked at:

  • Escalation rate. What share of cases the agent hands to a person. Expect it to fall over the first weeks as edge cases become rules, then flatten. A rise afterwards is the first symptom of drift.
  • Resolution rate. What share it completed end to end. Read alongside escalation rather than on its own — a high resolution rate with rising complaints means it is resolving things wrongly.
  • The list of what it could not handle. Not a number, a list. This is the most useful output the agent produces, because it is a continuously updated description of what to fix next.

If the escalation rate never flattens, the process was less defined than everyone believed at the start — which is information worth having, whatever you do with it.

The monthly hour

Once settled, keeping an agent healthy is roughly an hour a month:

  • Read the escalations from the last month. Ten minutes, and it is where the surprises live.
  • Check whether anything the agent states has changed — prices, products, people, hours.
  • Turn recurring escalations into rules, so the same case does not keep arriving.
  • Glance at the three numbers as a trend, not a snapshot.

The hour matters less than the name. An agent that is everybody’s responsibility is nobody’s, and that is the condition under which drift goes unnoticed for a quarter.

The owner is the process owner, not IT

The person who knows that pricing changed is in sales, not in IT. Routing agent updates through a technical queue means the update competes with tickets and loses. Give the agent to whoever owns the process and give them a way to change what it says without filing a request.

Models change underneath you

Providers update models, and behaviour can shift without any change on your side. Usually it is imperceptible. Occasionally an agent becomes more cautious and escalates more, or more confident and escalates less.

This is the practical reason to watch the escalation rate as a trend. An unexplained movement is visible in the numbers before anyone complains, which is the difference between noticing in a week and noticing in a quarter.

What running it costs

Model usage, hosting and monitoring. For one process at the volumes a company of 10 to 100 people generates, a monthly figure in the low hundreds rather than thousands. The variable is conversation volume; the rest is close to fixed.

We put the expected range in writing before the pilot starts. A build price without a running cost is only half an answer — and it is the half that turns a good decision into an awkward conversation three months later.

Knowing when to switch it off

An agent that has stopped earning its keep should be turned off rather than tolerated. The signals are consistent: the escalation rate climbing back to where it started, the team routing around it, the monthly review being skipped twice in a row.

That last one is the reliable early indicator. When nobody can be bothered to read the escalations, the agent has already lost the organisation — and no amount of technical health changes that.

What people ask about operations

How much attention does a live agent need?
About an hour a month once it has settled, and considerably more in the first two weeks. What matters is not the quantity but that it has a name attached. An agent with no owner degrades quietly and gets blamed for an organisational gap.
What does "drift" mean for an AI agent?
The business changes and the agent does not. Prices go up and it quotes the old ones. A product is discontinued and it keeps offering it. Nothing broke technically, which is why nobody notices for a quarter. Drift is the most common cause of an agent being switched off.
What should we actually monitor?
Escalation rate, resolution rate, and the list of cases the agent could not handle. Three numbers get looked at every week; a dashboard with twenty does not. The list of failures is the most useful of the three because it tells you what to fix next.
What happens when a model provider changes something?
Behaviour can shift without any change on your side. That is why the escalation rate is worth watching as a trend rather than a snapshot: an unexplained rise is usually the first visible symptom, and it shows up before anyone complains.
Who should own the agent internally?
Whoever owns the process, not IT. The person who knows that the pricing changed or that a product was withdrawn is the one who needs to update the agent, and routing that through a technical queue is how the update does not happen.

Read next

Already running something that is drifting?

Thirty minutes on what it does and what it has stopped doing well is usually enough to say whether it is worth fixing.

Book a scoping call