AI in engineering companies · Agents
The cost and infrastructure of agents
The first time an organisation runs agents seriously, the bill is often a surprise. A task that cost pennies as a single question can cost many times more when an agent works through it. And the first time someone suggests running agents on their own hardware to save money, they often discover that the machine that runs one agent well cannot run twenty.
This article is about the economics and infrastructure of agents: why they cost what they do, how to keep that under control, and when running your own makes sense. It builds on Where should your AI live?, in the five levels series, which covered the general trade-offs between cloud, UK-hosted and local AI.
Why agents cost more than chat
A single question to an AI model is read once and answered once. An agent works in a loop, and at every turn of the loop it reads everything relevant again: its instructions, the task, the tools available, everything it has found so far, and the results of the last step. Then it reasons about what to do next.
That has three consequences.
Reading dominates. For agent work, most of the computing goes on reading, not writing. A coding agent that works through a codebase reads the same files repeatedly as it goes. A fault investigation agent rereads its accumulated evidence at every step.
Cost grows with the length of the task. Each additional step costs more than the last, because there is more accumulated context to reread.
Failures are expensive. An agent that is stuck keeps looping, and every loop costs money. Without limits, a single confused run can cost more than a hundred successful ones.
None of this makes agents uneconomic. It means the economics need designing, in the same way as any other engineering system.
Keeping cost under control
Use workflows where you can. As I argued in Workflows before agents, a workflow uses AI once per step. For work whose steps are known, it is far cheaper than an agent.
Set budgets on every run. Limits on steps, time, retries and spending are a cost control as well as a safety control, as described in Guardrails: autonomy is earned. In the platform I am developing, every AI call goes through a metered gateway that checks the budget before it runs and reconciles the actual usage afterwards. That discipline is worth having from the first day.
Choose the model for the step. Not every step needs the most capable model. Routine classification, extraction and summarising can use smaller, cheaper models, keeping the frontier models for the reasoning that needs them.
Give agents less to read. Well-designed tools return what is needed, not everything available. An agent handed a precise summary of a unit's history reasons faster and more cheaply than one handed the raw records.
Measure cost per result. The number that matters is not cost per call. It is cost per accepted result: per investigation completed, per test suite approved, per report signed off. That number tells you whether an agent is worth running.
When running your own makes sense
Running open models on your own hardware changes the economics. Once the hardware is paid for, each additional task costs very little. For high-volume, sustained agent work, that can be much cheaper than paying per use, provided the hardware is kept busy.
There are also reasons beyond cost: keeping sensitive material in the building, working without a reliable connection, and having a model that does not change underneath you. I covered those in Where should your AI live?.
But there are two things about local hardware that are often misunderstood.
Memory decides what fits. Bandwidth decides how fast it runs. The amount of memory in a machine determines which models it can hold. The speed at which that memory can be read determines how quickly the model produces its answer. They are different specifications, and buying for one while ignoring the other is an expensive mistake. For agent work, which is dominated by reading long contexts, the speed at which the machine processes its input matters a great deal, and it varies widely between hardware platforms.
A workstation is not a serving platform. A powerful workstation can be excellent for one engineer, or for a handful of agents working in turn. It is not built to serve many users or many agents at once. When several requests arrive together, they compete for the same memory bandwidth and often simply queue. Serving AI to a whole organisation, or running many agents in parallel, is a different engineering problem, usually solved with dedicated server hardware and software designed for concurrent requests.
This matters because the right platform for experimenting and developing is often not the right platform for running in production. If agents are going to be part of how an organisation works, the development environment should resemble the environment they will run in.
A sensible path
For most engineering companies, I would suggest:
- Start in the cloud. Develop and prove agents using hosted models, with UK residency and no data retention where the material requires it. It is the fastest way to learn what works.
- Measure. Record cost per accepted result, volume and which steps use the most.
- Move the volume. When specific high-volume steps are proven and stable, consider moving them to local or UK-hosted models, where the cost per task falls sharply.
- Keep the frontier for the hard parts. Route the difficult reasoning to the most capable models, as described in the section on blending in Where should your AI live?.
- Plan for serving separately. If agents become part of daily operations across the organisation, treat running them as infrastructure, with the capacity, resilience and monitoring that implies.
Four things worth taking seriously
For finance and engineering leaders: budget for agents per result, not per call, and expect the first real bills to teach you something.
For technical teams: put a budget check in front of every AI call from the start. It is far easier than adding one after a surprise.
For anyone considering local hardware: size it for the work you will actually run, understand the difference between memory and bandwidth, and do not confuse a workstation with a server.
For everyone: agents are an engineering system with running costs, like any other. Design the economics as deliberately as the functionality.
I would be interested to hear whether your organisation has had its first surprising AI bill yet, and what it taught you.
Further reading in this series
- Where should your AI live? Cloud, UK-hosted or in the building (the five levels series)
- Workflows before agents
- Guardrails: autonomy is earned
- Running agents in production
- Agents with keys (security series)
- What tokens are, and why they matter (coding agents series)