AI in engineering companies · Coding agents
Managing a token budget
In What tokens are, and why they matter, I explained that tokens are the unit in which AI does its work, and that coding agents consume them in large quantities, mostly by rereading context at every step. This article is about managing that consumption: keeping costs predictable without giving up the productivity that makes coding agents worth using.
The aim is not to spend as little as possible. An engineer's time is far more expensive than tokens, and an agent that saves a day of work is cheap at almost any realistic price. The aim is to spend deliberately, on work that benefits, and to stop waste.
Where the tokens go
Before managing a budget, it helps to know where it is being spent. In my experience with coding agents, the biggest consumers are:
- long sessions, where accumulated history is reread at every step
- reading too much, such as whole directories or large files when only part is relevant
- verbose tool output, such as complete build logs or test output fed back into the context
- agents that are stuck, repeating attempts that do not work
- the wrong model, using the most capable and expensive model for routine steps
- repeated work, such as re-explaining the project at the start of every session
Most of these are avoidable, and none requires giving up anything useful.
Practical techniques
Scope the task tightly. The most effective technique is also the simplest. A small, well-defined task with a clear definition of done finishes in fewer steps, with less context, and with better results. This is the same discipline I described in Building production software with AI coding agents, and it saves money as well as improving quality.
Point the agent at what matters. Tell it which files and modules are relevant. An agent that has to search a large codebase to find its way spends tokens exploring.
Keep a project instruction file. Most coding agent tools support a file of standing instructions for the project: architecture, conventions, security rules, where things are. Written once and kept concise, it saves re-explaining the project in every session and, because it is the same every time, it benefits from caching.
Start fresh for new tasks. A long session carries the history of everything done so far. When a task is complete, start a new session for the next one rather than continuing to accumulate context. Where a long task is unavoidable, use the tool's ability to summarise and compact the history.
Trim tool output. Feed back the relevant part of a build log or test run, not all of it. Well-configured tools do this automatically. Badly configured ones can fill the context with thousands of lines of noise.
Use caching. Keep the stable parts of the input, such as instructions and project context, at the beginning and unchanged, so the provider's caching can discount them. Many agent tools do this well by default. Check that yours does.
Choose the model per task. Use smaller, cheaper models, or local ones, for searching, summarising, simple edits and first drafts, and keep the most capable models for design review, difficult debugging and security-sensitive work. This is the subject of Local, cloud and frontier: getting the most from your budget.
Set limits. Cap the number of steps, the time and the spend for each task, so an agent that is stuck stops and asks rather than looping. As I described in Guardrails: autonomy is earned, this is a safety control as well as a cost control.
Stop and think when it struggles. If an agent has tried the same thing several times, more attempts rarely help. Usually the task is unclear, the context is wrong or the problem needs a person. Stopping early is the cheapest option.
Budgeting for a team
Measure the right thing. Cost per call tells you little. Cost per completed task, per merged change or per accepted result tells you whether agents are worth what they cost. As I argued in The cost and infrastructure of agents, that is the number to track.
Give visibility. Engineers who can see what their sessions cost quickly learn which habits are expensive. Most waste is unintentional.
Set budgets, not bans. A monthly allowance per engineer or per team, with the ability to request more for work that justifies it, encourages sensible use without discouraging valuable work.
Understand the pricing model. Some tools charge per token, some offer subscriptions with usage limits, and some combine the two. The right choice depends on how heavily your teams use agents. Heavy users on per-token pricing can find subscriptions far cheaper, and light users the reverse.
Watch for the surprise. A single misconfigured automated agent, running unattended in a build pipeline, can consume more than a whole team. Automated uses need hard limits and alerts.
A sense of proportion
It is easy to become preoccupied with token costs. Keep them in proportion. An engineer costs the organisation hundreds of pounds a day. If a coding agent saves hours on a task, the tokens it consumed are almost always worth it.
The real waste is not an agent that costs a few pounds to complete useful work. It is an agent left looping for hours on a badly defined task, a team using the most expensive model for everything by default, or an automated process with no limits at all. Those are what budget management should catch.
Four things worth taking seriously
For engineering leaders: measure cost per completed task, not total spend, and set budgets that encourage good habits rather than discourage use.
For engineers: small tasks, focused context, fresh sessions and the right model for each step. Those four habits remove most waste.
For teams running agents automatically: set hard limits and alerts on every automated use.
For everyone: tokens are cheap compared with engineering time. Spend them deliberately, not carelessly.
I would be interested to hear how your organisation budgets for coding agents, and whether the first bills matched expectations.