AI in engineering companies · Coding agents
Coding agents: strengths, weaknesses and opportunities
Most of what is written about coding agents is either breathless or dismissive. They will replace developers, or they produce rubbish. Neither matches my experience. I use them daily, and have used them to build a production platform and to develop an AI system of my own. They are remarkably good at some things, reliably poor at others, and the difference is predictable once you understand it.
This article is an honest assessment: what coding agents are good at, where they fail, and where the real opportunities lie for engineering companies.
What a coding agent actually does
A coding agent is a language model given tools to work on code: reading files, searching the codebase, editing, running commands, running tests. It works in a loop, as described in What an agent actually is, and isn't: it makes a change, runs something, reads the result, and decides what to do next.
That loop is what distinguishes a coding agent from code completion. It does not just suggest the next line. It attempts a whole task, checks its own work against whatever evidence it can get, and keeps going until it believes it has finished.
Strengths
Speed on well-defined work. Given a clear specification, a coding agent produces working code in minutes that would take a person hours. The more precisely the task is defined, the better it does.
Breadth. It is fluent in more languages, frameworks and libraries than any individual engineer. An experienced engineer can work competently in an unfamiliar language with an agent handling the syntax and idioms.
Tests. Agents are good at writing tests, including the tedious edge cases people skip. Combined with their ability to run tests and read the results, this makes test-driven work with agents very effective.
Reading code. Asking an agent to explain how a large, unfamiliar codebase works, trace a data flow or find where something happens is one of its most valuable and least appreciated uses.
Refactoring and migration. Mechanical transformations across many files, such as renaming, restructuring or moving from one framework to another, are laborious by hand and well suited to agents, provided the target structure is decided by a person.
Documentation. Agents write clear first drafts of documentation, kept alongside the code, at almost no marginal cost.
Simulation. I have found agents excellent at building simulations of an operating environment, so an algorithm can be tested against realistic conditions before it touches hardware.
Weaknesses
Coherence across a system. Agents solve local problems well and global problems badly. Left to design a system, they produce something where each part is sensible and the whole is inconsistent. Architecture has to come from a person.
Confident errors. Agent-written code often looks clean and reads well while being subtly wrong: a race condition, an unhandled edge case, a check in the wrong place. Its fluency makes errors harder to spot, not easier.
Limited context. An agent can only consider what is in front of it at the time. On a large codebase, it may not see the constraint in another module that makes its change wrong. What it is given to read determines what it knows.
Invented details. Agents sometimes use functions, options or libraries that do not exist, or that existed in an older version. Most of these are caught by compilation or testing. Some are not, and invented package names are a real security risk, discussed in The AI supply chain.
Gaming the goal. Asked to make tests pass, an agent will occasionally change the tests. Asked to fix an error, it may suppress it. Agents optimise for the evidence they are given, so the evidence has to be trustworthy.
Security. Agents reproduce the patterns in what they learned from, including insecure ones. The typical issues are the subject of The security of AI-written code.
Drift. Over a long session, an agent gradually departs from the agreed design, one reasonable local decision at a time.
Embedded and hardware-specific work. In my experience, agents are weaker on embedded systems than on web and application code: timing, interrupts, memory constraints, hardware quirks and vendor-specific peripherals are less well represented in what they learned, and mistakes are harder to detect without the hardware in the loop. They are still useful there, but they need closer supervision.
Opportunities for engineering companies
Legacy code. Many engineering companies depend on firmware and software whose authors have left. Agents can map it, document it, add tests around it and make it safe to change. This is one of the highest-value uses I know, and it pairs naturally with the regulatory pressure to maintain connected products throughout their support period.
Test coverage. Code that was never properly tested because it was not affordable can now be covered, making every later change safer.
Migrations. Moving an application to a new framework or platform, a job that used to take a team months, becomes a supervised programme of well-defined tasks.
Tools the organisation never had time to build. Test rig software, data analysis tools, internal dashboards, configuration utilities. Engineering companies are full of useful tools nobody had time to write.
Architects who implement. The biggest opportunity is the one I described in Vibe coding is just another level of abstraction: letting experienced engineers turn designs into working systems directly, without the losses of handing over.
What makes the difference
In nearly every case, the difference between an agent that helps and one that harms comes down to three things: the clarity of the task it is given, the quality of the evidence it can check its work against, and the experience of the person reviewing the result. Those are the themes of the rest of this series.
Four things worth taking seriously
For engineering leaders: expect large gains on well-defined work and none on architecture. Plan accordingly.
For senior engineers: give agents small, precise tasks with tests as the definition of done. Keep the design to yourself.
For embedded teams: use agents, but keep the hardware in the loop and review timing and resource-sensitive code with particular care.
For everyone: an agent's strengths and weaknesses are predictable. Use it where it is strong and supervise it where it is weak.
I would be interested to hear where coding agents have surprised your teams, in either direction.