AI in engineering companies · Agents

Running agents in production

Getting an agent to work in a demonstration is easy. Getting it to work reliably, every day, for months, while the data changes, the models change and the people using it change, is an engineering discipline of its own. It is also where most agent projects quietly fail.

This is the last article in this series on agents. It is about what it takes to run them in production: how to know they are working, what changes underneath them, what to do when something goes wrong, and when to use one agent or several.


You cannot run what you cannot measure

Build an evaluation set before you build the agent. A collection of real tasks, with known good outcomes, drawn from your own work: past fault investigations with their confirmed causes, past change impact assessments, past test suites. This is the agent equivalent of the fifty test questions I recommended for search in Your company already knows the answer.

Run it every time anything changes. A new model version, a new prompt, a new tool, a change to the data. Agents are sensitive to all of them, and the evaluation set is how you find out before your users do.

Test with the real model. This is a lesson from my own development work that I think is widely under-appreciated. Tests that use simulated, pre-written AI responses are useful for checking the plumbing around an agent. They prove nothing about how a real model will behave, particularly when it is given misleading or hostile input. In the platform I am developing, I have had to be careful to distinguish tests that check the surrounding code from evidence about the model's actual behaviour. Both are needed. Only one tells you whether the agent works.

Include the difficult cases. Missing evidence, contradictory records, misleading input, tasks that should end in "I don't know". An agent that only ever sees clean cases in testing will meet the messy ones in production.


Watch it while it runs

Monitor outcomes, not just uptime. An agent can be running perfectly and producing poor results. Track the measures that matter: how often its output is accepted unchanged, corrected or rejected, as described in AI drafts, engineers decide, and the cost per accepted result, from The cost and infrastructure of agents.

Watch for drift. Acceptance rates that slowly fall, runs that slowly get longer, costs that slowly rise. These are early signs that something has changed: the data, the model, or the way people are using it.

Keep the full record. Every step, every tool call, every input and output, as described in Guardrails: autonomy is earned. When something goes wrong, the record is the only way to understand why.


The model will change underneath you

This is the operational risk most specific to AI.

If you use a hosted model, the provider will update it, retire it and replace it on its schedule, not yours. A new version may be better overall and worse at your particular task. Behaviour you relied on may change without notice.

Pin versions where you can. Use a specific model version, and change it deliberately.

Treat a model change as a release. Run the evaluation set, compare the results with the current version, and roll out only when you are satisfied, with the ability to go back.

Keep the model replaceable. Design agents so the model behind them can be swapped without rebuilding everything around it. In my own development, AI access sits behind a defined boundary with interchangeable providers, so a model can be changed or compared without touching the rest of the system. That also makes it possible to move steps between cloud and local models as costs and requirements change, as discussed in Where should your AI live?.


When things go wrong

Agents will fail. The question is whether the failure is contained, noticed and understood.

Design for failure. Every tool call can fail, time out or return something unexpected. The agent, and the system around it, should handle each case explicitly: retry where it is safe, stop and report where it is not.

Be careful with retries. Retrying a read is harmless. Retrying an action that may already have happened can cause real damage. Actions with consequences should be designed so that repeating them is safe, or checked before they are repeated.

Have an incident process. Who is told, who can stop the agent, how its actions are reviewed and, where necessary, reversed, and how the lesson is fed back into the evaluation set so it cannot happen silently again.

Degrade gracefully. If an agent is unavailable, the work should fall back to the manual process, not stop. I have spent much of my career designing redundant systems, including the dual-node environmental monitoring system I built, and the principle is the same: assume parts will fail, and make sure the whole keeps working.


One agent or several?

It is fashionable to build systems of many agents that talk to each other. Sometimes that is the right design. Often it is not.

Several agents make sense when roles are genuinely different and benefit from different tools, permissions or models. The five agents in Agents at work in engineering are specialised for good reasons: an investigation agent that only reads should not share permissions with one that drafts notifications. A second agent reviewing the first agent's work, using a different model, is also valuable, for the same reason I use cross-model checking in my own development.

One agent is better when the task is coherent and the handoffs would add nothing but complexity. Every handoff between agents is a place where context is lost, errors compound and costs rise. Each additional agent is another component to test, monitor and debug.

My rule is to start with one well-bounded agent, and split it only when there is a clear reason: different permissions, different expertise, or independent checking.


Four things worth taking seriously

For engineering leaders: before an agent goes into production, ask to see its evaluation set and its results with the real model. If there is none, it is not ready.

For technical teams: treat every model change as a release, with testing, comparison and the ability to roll back.

For operations: agents need monitoring, incident handling and fallback plans like any other production system. Plan them before they are needed.

For everyone: the difference between a promising demonstration and a dependable tool is almost entirely in the operational discipline around it.


This is the last article in the series on agents. The next series turns to security: what changes when AI is connected to your data and systems, and how to keep it safe. It begins with The attack surface has moved.

I would be interested to hear how your organisation tests AI systems before relying on them, and whether that testing uses the real model.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.