Writing · AI and technology leadership

When AI decisions have physical consequences

When hardware in the field acts on an AI classification, human oversight has to be engineered, not assumed.

The ethical problem with AI in business is not that nobody mentions it. It is that the conversation tends to become serious only after an organisation has already decided it wants the capability.

In the discussions I have had about AI-assisted decision-making, the excitement about what a system can do has consistently arrived before the harder question: what it may systematically get wrong, and to whom. When the decision is enacted by hardware in the field, that order matters more than anywhere else.


Opacity and accountability

For certain categories of AI, particularly the more complex machine learning architectures, there is no clean narrative for why the model reached its conclusion. The inputs can be described and the outputs observed, but the reasoning in between is essentially opaque. Organisations tend to treat this as a communications challenge. In my experience it is better understood as a fundamental constraint on where these systems should and should not be deployed.

The accountability question follows directly. If an AI-driven decision turns out to be wrong, not merely statistically suboptimal but wrong in a way that affects somebody’s livelihood or safety, where does the responsibility sit? With the team that built the model, the people who chose to deploy it, or the organisation that collected and curated the training data? The honest answer is usually some combination of all three, which means accountability diffuses until nobody is clearly responsible.

In a software context, that diffusion is uncomfortable. In a connected product context, where hardware in the field may have enacted the decision before anyone noticed the error, it is something worse.


When the decision controls a physical system

Opacity and accountability apply to all AI-assisted decision systems. They become more serious when the system directly controls physical hardware.

In a connected product fleet, whether e-scooters, industrial sensors or commercial vehicles, AI is increasingly present in fault diagnosis, predictive maintenance scheduling, and operational decisions that the hardware enacts immediately. A model that classifies fault codes from telemetry, identifies which units to service before failure, or decides which firmware to push to which cohort of devices is not making a recommendation that a human will review first. It is making a classification that a field technician will generally follow, with limited ability to interrogate why.

Fault classification shows this concretely. A model trained on historical telemetry learns to associate patterns of sensor readings with fault categories. When it classifies a battery thermal event as a firmware issue, the workshop receives that classification as an instruction: reflash the firmware, return the unit to service. The technician has no practical way to see the model’s reasoning, no diagnostic tool that explains why the thermal signature was read as a software fault rather than a hardware one, and, in most field service environments, no cultural expectation that they should question an automated classification. If the classification is wrong, the thermal problem remains, the unit goes back into the field, and the fault returns. Or it progresses until it is a safety problem rather than a maintenance one.

The failure is easy to picture. A fleet of vehicles starts reporting intermittent power cut-outs. The classifier, trained mostly on cases where cut-outs were cured by a firmware update, labels them as firmware faults. Workshops reflash the units and return them to service, and the fault statistics improve for a few weeks because the reset clears the symptom. The real cause, a battery connector degrading under heat and vibration, carries on getting worse. Nobody in that chain was careless; each step did what the system said. That is exactly why the accountability question has to be answered before the system is deployed, not after the first incident.


When the operating environment changes

The training-data problem has a connected-product dimension that the general AI ethics literature rarely addresses. A predictive maintenance model trained on fleet data from one operating environment (temperate climate, well-maintained road surfaces, moderate duty cycles) will still produce confident predictions when deployed in a different one. Higher ambient temperatures, rougher terrain, longer duty cycles or different charging patterns shift the failure profiles in ways the model has never seen.

The model does not know it is operating outside the conditions it was trained on. It keeps producing classifications with the same apparent confidence, and the field team has no signal that the basis for them has quietly become unreliable.

In a connected product business expanding into new markets, the AI layer inherits the same new-market problem as the rest of the product. The assumption that what worked in the first deployment environment will transfer to the next is precisely the assumption most likely to be wrong, and an AI model will not flag the discrepancy. It will simply be wrong more often, in ways that become visible only after enough units are affected to produce a pattern in the warranty data.


Misclassification at scale

I have seen, in connected product work, what happens when misclassification propagates through a decision pipeline. A reading is processed, a classification is made, a decision follows, and by the time the error becomes visible in the field it has already been acted on across a deployed fleet.

When an AI model makes the classification rather than a human analyst, the volume is higher, the reasoning is opaque by design, and the trail back to the original error is harder to reconstruct. A human analyst making the same call on individual units would eventually notice something: the pattern of returns, the technician’s pushback, the warranty data that did not match the diagnosis. An automated system works at a speed and scale that outruns the feedback loop that would catch the error in a human process.

For a CTO deploying AI in any safety-relevant connected product, the human oversight principle needs to be defined more precisely than it usually is in software. Not a human who can theoretically override a recommendation, but a human with enough understanding of the reasoning, and enough diagnostic tooling to surface it, to know when an override is warranted and to justify that call.


What this means for a connected product CTO

I am not sceptical about AI. I use it in my own work and find it genuinely useful in a range of contexts. But the ethical frameworks most organisations have for AI deployment are considerably thinner than the confidence with which they deploy these systems would suggest. The questions that need better answers are not novel. They are the questions any decision system should face: who does this affect, what happens when it is wrong, and how will we know?

In a connected product business, the third question demands the most deliberate engineering. A software company can instrument its AI decisions with logging, A/B tests and rapid rollback. A connected product company whose AI layer makes classifications that hardware in the field then enacts needs something more: diagnostic tooling that makes the AI’s reasoning visible to the people who act on it, not just to the data science team that built it. The field technician needs to see why the system classified a fault the way it did, in the same way they need to see the device’s sensor data and firmware state.

Observability into the AI layer is not a different kind of problem from observability into the hardware layer. It is the same problem, applied to a newer and less familiar part of the system.

My own working position is that human oversight remains non-negotiable in any AI-assisted process where decisions have physical consequences. Not because AI tools are necessarily less reliable than human judgement, but because human oversight is how the feedback loop stays open. A model that is wrong about fault classification in a new operating environment will not correct itself. The technician who can see into the reasoning, and who has the standing and the tools to question it, is what closes the loop between the model’s confidence and the reality of the product in the field.

Building that capability is an engineering investment, and it competes with the same delivery pressures as every other form of observability and diagnostic tooling. The CTO who understands why it matters is the one who can protect it.


Four things worth taking seriously

For boards: ask who is accountable when an AI-driven decision is wrong before the system is deployed, not after the first incident. If the answer is “everyone”, it is nobody.

For connected product CTOs: treat observability into the AI layer as the same problem as observability into the hardware, and fund it the same way. Diagnostic tooling that explains a classification is part of the product.

For field service leaders: give technicians both the tools to see why a fault was classified as it was and the standing to question it. That is where the feedback loop closes.

For anyone taking a connected product into new markets: assume a model trained in one operating environment is less reliable in the next, even though its confidence will not change. It will not tell you it is outside its training conditions.


I would be interested to hear how your organisation lets the people who act on an AI classification see why it was made, and whether they feel able to challenge it.

© 2024 Catherine Ives-Yim. All rights reserved.

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.