AI in engineering companies · Coding agents
Local, cloud and frontier: getting the most from your budget
Most engineers using coding agents use one model for everything, usually the most capable one they have access to. It is the simplest approach, and for an individual it is often reasonable. For an organisation, it is expensive, it sends more code outside the building than necessary, and it ignores the fact that most coding work does not need the most capable model.
This article describes a tiered approach: using local models, mid-range cloud models and frontier models each for the work they suit. It applies the principles from Where should your AI live?, in the first series, specifically to coding.
Three tiers
Local models. Open models running on your own hardware, whether a capable workstation or a shared server. Once the hardware is paid for, the cost per task is close to zero, and nothing leaves the building. The best open coding models are now genuinely capable, though still behind the frontier on the hardest problems.
Cloud models, mid-range. The smaller and faster models offered by the major providers. Much cheaper per token than the frontier models, quick to respond, and entirely adequate for a large share of routine coding work.
Frontier models. The most capable models available, in the cloud. The best at complex reasoning, holding a large system in mind, subtle debugging and security review. Also the most expensive, and the ones most often used by default for work that does not need them.
Matching work to tier
The useful question for each piece of work is: what is the least capable model that will do this well?
Suited to local models:
- code completion and small edits while typing
- searching and indexing the codebase, and building the search index itself
- explaining what a piece of code does
- generating routine tests and documentation
- summarising logs, diffs and tool output
- any work on code that should not leave the building
Suited to mid-range cloud models:
- implementing well-specified, contained features
- routine refactoring and migration steps
- first drafts of larger pieces of documentation
- most of the individual steps in a well-scoped agent task
Worth the frontier:
- reviewing an architecture or a significant design decision
- difficult debugging, where the cause is not obvious
- security review of sensitive code
- tasks that need the whole system held in mind at once
- a second opinion when the cheaper model is struggling
In practice, the frontier share of the work is much smaller than its share of the typical bill.
Escalation, not choice
The most effective pattern is not to choose a tier in advance but to escalate. Start with the cheapest model that might succeed. If it struggles, repeats itself or produces work that fails review, move the task up a tier. Most tasks finish at the bottom. The difficult ones reach the top, where the capability is worth its price.
Agent tools increasingly support routing different steps to different models automatically. Where they do not, engineers can learn to do it by habit: a quick local model for exploration and routine work, a frontier model called in deliberately for the hard parts.
Cross-checking across tiers
One practice I rely on works particularly well across tiers. Having a different model review work produced by the first catches errors that either would miss alone, because different models fail in different ways. A frontier model reviewing the output of a cheaper model gives much of the quality of using the frontier throughout, at a fraction of the cost. For security-sensitive work, I have a different model review changes as a matter of routine, as described in The security of AI-written code.
The hardware question
Running models locally needs hardware, and it is easy to buy the wrong thing. As I described in The cost and infrastructure of agents, memory capacity decides which models fit, and memory bandwidth decides how fast they run. For coding agents, which read long contexts at every step, the speed at which the machine processes its input matters a great deal.
A capable workstation can serve one engineer well for local completion, search and routine agent work. A team needs shared hardware designed for many requests at once. Start by measuring what your engineers actually use, then decide which work is worth moving local. Buying hardware first and finding uses later is the expensive way round.
Privacy is part of the calculation
Choosing a tier is not only about cost. Every request to a cloud model sends code outside the organisation. For most code, with the right provider terms, that is acceptable. For the crown jewels, such as proprietary firmware, core algorithms and security-sensitive modules, it may not be. Local models make it possible to use AI on that code without it leaving the building, at the cost of some capability. That trade-off is the subject of the final article, Keeping your code local and private.
A sensible path
- Start with cloud models, choosing the mid-range tier as the default rather than the frontier.
- Measure where tokens go and which tasks need the frontier, as described in Managing a token budget.
- Route deliberately: frontier for the hard parts, mid-range for routine work, cross-checking where it matters.
- Add local models for high-volume routine work and for code that should stay in the building, once you know the volume justifies the hardware.
- Keep the arrangement flexible. Models change quickly. Build so that any tier can be swapped without rework.
Four things worth taking seriously
For engineering leaders: the frontier model should be a deliberate choice for difficult work, not the default for everything.
For engineers: start cheap and escalate. Use a second model to check the work that matters.
For anyone buying hardware: measure first, and buy for the work you will actually move local.
For everyone: the best result for the budget comes from using each tier for what it does best.
I would be interested to hear whether your teams use one model for everything, and what share of their work genuinely needs the most capable one.