AI in engineering companies · Agents
Agents at work in engineering
The first articles in this series were about what agents are and when not to use them. This one is about where they are genuinely useful in an engineering company. I describe five agents, each doing a job that most engineering organisations need done and rarely have the time for, and for each I say how much autonomy it should have and what it depends on.
Throughout the earlier series on the five levels of AI value, I followed a single example: a fault that keeps appearing in a connected product in the field. I pick it up again here, because it shows how these agents work together.
1. The fault investigation agent
What it does. Given a unit, or a cluster of units, showing a fault, it works out what they have in common and what changed. It queries the build records, firmware history, telemetry and service tickets, deciding at each step where to look next based on what it has found, and reports its findings with evidence.
Level. This is level 3, connect, from From what the manual says to what the product is doing. It is a genuine agent task, because the path of the investigation cannot be known in advance.
Autonomy. Read-only. It can look at anything the engineer asking could look at, and change nothing. Its output is a report, not an action.
Depends on. Identifiers that match across systems, which is the structured model described in From documents to data. Without it, the agent spends its effort guessing whether two records describe the same unit.
In our example. The agent finds that the affected units share a component batch and a recent firmware update, that the fault codes began shortly after the update, and that eleven other units share the same combination. It presents the evidence and says what it could not determine.
2. The vulnerability triage agent
What it does. When a new vulnerability is published in a third-party software component, it works out whether your products are affected. It checks which products contain the component and in which versions, whether the vulnerable part of that component is actually used, which units in the field are running it, and how exposed they are.
Level. Levels 2 and 3. It depends heavily on knowing exactly what software is in each product.
Autonomy. Read-only analysis, followed by a drafted assessment and, if needed, a drafted notification for an engineer to approve. Under the EU Cyber Resilience Act, a manufacturer that becomes aware of an actively exploited vulnerability in its product must send an early warning within twenty-four hours. An agent that does the analysis in minutes turns a frantic day into a considered decision.
Depends on. An accurate software bill of materials for every product and version. This agent is a very strong argument for building one.
In our example. It becomes central later. In Connected products versus AI-assisted attackers, in the security series, the recurring fault turns out to have a security dimension.
3. The test generation agent
What it does. Given a change, a fix or a new variant, it proposes and writes tests: normal cases, edge cases, failure cases and regression tests for past faults. It runs them, reads the results and refines them.
Level. Level 4, act. Its output is work product for engineers to review.
Autonomy. Can write and run tests in a controlled environment. Cannot change the product code, and cannot alter existing tests without review. As I noted in Building production software with AI coding agents, an agent asked to make tests pass will sometimes change the tests instead.
Depends on. A clear specification of the intended behaviour, and access to the history of past faults.
In our example. Once engineering has a candidate fix, the agent writes the tests that confirm it and a regression test that will catch the fault if it ever returns.
4. The legacy code documentation agent
What it does. Works through a large, poorly documented codebase, often firmware whose original authors have left, and produces an explanation: what each part does, how data flows, what depends on what, and what looks risky or dead.
Level. Levels 1 and 2. It creates findable, structured knowledge where none existed.
Autonomy. Read-only access to the source code. Its output is documentation for engineers to check.
Depends on. Careful choice of where it runs. Source code is usually the organisation's most sensitive asset, which makes this a natural candidate for a local or UK-hosted model, as discussed in Where should your AI live?.
In our example. Before engineering changes the firmware involved in the fault, the agent maps the affected module, its dependencies and its history, so the fix does not break something else.
5. The change impact agent
What it does. When a supplier discontinues a component, changes a specification, or a design change is proposed, it works out everything affected: products, variants, test plans, certifications, documentation, spares and units in the field.
Level. Level 2 for the analysis, level 4 for the drafted impact assessment.
Autonomy. Read-only analysis and a drafted assessment for approval.
Depends on. The structured engineering model. This is exactly the kind of question search cannot answer and structure can.
In our example. If the fix involves changing the component batch or supplier, the agent identifies every product, certificate and document affected.
How they fit together
These five agents are not separate products. They are specialised roles that share the same foundations: a searchable body of knowledge, a structured engineering model, controlled connections to live systems, and an engineer who owns the outcome.
In our example, the investigation agent finds the pattern, the documentation agent maps the code, the test agent proves the fix, the change impact agent identifies what else is affected, and the triage agent assesses whether there is a security exposure. At every point where something changes in the world, a person approves it.
It is also worth noticing what none of them do. None of them ships code, changes a production system, or contacts a customer on its own. That is deliberate, and it is the subject of the next article, Guardrails: autonomy is earned.
Four things worth taking seriously
For engineering leaders: pick one of these five that matches your most painful gap. For many connected product companies right now, it is vulnerability triage.
For anyone planning agents: specialised agents with narrow roles and clear limits are more reliable than a single general agent with access to everything.
For quality and compliance teams: most of these agents depend on the traceability described in the first series. The case for building it keeps getting stronger.
For engineers: these agents do the gathering and the groundwork. The judgement, and the signature, remain yours.
I would be interested to hear which of these five your organisation would find most valuable, and which you would be most nervous about.
Further reading in this series
- What an agent actually is, and isn't
- From documents to data: building an engineering model you own (the five levels series)
- From what the manual says to what the product is doing (the five levels series)
- Closing the loop: letting field data change the product (the five levels series)
- Guardrails: autonomy is earned
- Connected products versus AI-assisted attackers (security series)