AI in engineering companies · Security

How data leaks out of AI systems

When organisations worry about AI and data, they usually worry about one thing: sending confidential information to an AI provider. That is a real concern, and I discussed it in Where should your AI live?. But in the systems I review, it is rarely the only route by which data escapes, and often not the most likely one.

This article sets out the main routes by which data leaks out of AI systems, and what to do about each. The aim is not to discourage anyone from using AI on their engineering data. It is to make sure the data goes only where it is meant to.


Route 1: to the provider

When you use a hosted AI model, what you send it leaves your organisation: the question, the documents retrieved to answer it, and anything an agent passes along the way.

What to do. Decide deliberately which material may go to which kind of model, collection by collection, as described in Where should your AI live?. Use UK or European processing and no-retention arrangements where the provider offers them, and understand what those do and do not protect. Keep the most sensitive material on UK-hosted or local models. Be aware that development tools, including AI coding assistants, may also send code and context to a provider while systems are being built.


Route 2: through the index

A search system over company documents, as described in Your company already knows the answer, holds an organised, searchable copy of everything it has indexed. If the system does not enforce the same permissions as the original documents, anyone who can ask it a question can potentially read anything in it.

This is one of the most common and least noticed problems. A document that was carefully restricted in its original location becomes available to the whole company through the AI.

What to do. Permissions must follow documents into the index. The system should only retrieve material the person asking is allowed to see. Restricted collections, such as HR records, customer personal data or unreleased designs, should be indexed separately or not at all. Treat the index itself as highly sensitive: it is effectively a copy of everything, organised for easy retrieval.


Route 3: through the assistant's reach

An AI assistant connected to live systems, as described in From what the manual says to what the product is doing, can bring together information from several places. If it acts with broader access than the person using it, it becomes a way round the organisation's normal access controls.

Customer telemetry is a particular concern. It often contains location, usage patterns and account details. An engineer diagnosing a fault may need some of it. They rarely need all of it.

What to do. The assistant should act with the permissions of the person using it, not with a master key. Tools should return the minimum data needed for the task. Personal data should be minimised, and the processing should be assessed under data protection law like any other new use. In the connected vehicle platform I designed, each customer's data is held in a separate database instance, which keeps the boundaries structural rather than dependent on every query being written correctly.


Route 4: through the output

An AI can be manipulated, through the prompt injection described in When the data gives the orders, into including data in its output in a way that sends it somewhere. A response might contain a link or an image reference that, when displayed or followed, passes data to an outside address. Or an agent with the ability to send messages might be persuaded to send the wrong thing to the wrong person.

What to do. Control what outputs can do. Do not automatically fetch links or images from AI output. Restrict where agents can send anything, and put a person in front of any external communication. Check outputs for sensitive data before they leave the organisation.


Route 5: through logs and records

Good practice, as I have argued throughout the agents series, is to log everything an AI system reads and does. Those logs therefore contain everything the AI has read: customer data, source code, confidential documents.

Logs are often stored with less care than the systems they describe, kept longer than necessary, and sent to third-party monitoring services.

What to do. Treat AI logs as being as sensitive as the most sensitive data they may contain. Control access to them, set retention periods, and be careful where they are sent.


Route 6: through derived data

To make documents searchable, AI systems convert them into numerical representations, often called embeddings, and store them. These are sometimes treated as harmless because they are not readable text. They should not be. Research has shown that embeddings can sometimes be used to reconstruct a meaningful amount of the original text, and they certainly reveal what the organisation holds.

What to do. Protect derived data with the same care as the source. Where the source is sensitive, so is everything derived from it.


The common thread

Almost every route above has the same cause: the AI system was treated as a tool, rather than as a new place where data lives and moves. Once you think of it as a system that holds, combines and transmits data, the controls are familiar ones: access control, minimisation, careful routing, protection of copies, and records.

Data protection law already expects this. Personal data processed through an AI system needs the same lawful basis, minimisation and security as anywhere else. Designing for it from the start, as I did with the connected vehicle platform's records of processing and consent management, is far easier than retrofitting it.


Four things worth taking seriously

For engineering leaders: ask where each AI system's data goes, including its indexes, its logs and its outputs, not just the model provider.

For data protection leads: treat every AI system that touches personal data as a new processing activity, and assess it as one.

For anyone building AI systems: make permissions follow documents into indexes, and make assistants act as the user, never with a master key.

For everyone: an AI system is not just a tool. It is a new place where your data lives. Secure it like one.


I would be interested to hear whether your organisation's AI search tools respect the same permissions as the documents they index. It is worth checking.


Further reading in this series

Catherine Ives-Yim

Catherine Ives-Yim

Chartered Engineer and independent technical adviser, with a lifetime at the bleeding edge of embedded systems, connected products, data platforms and AI-assisted engineering, who has advised clients across the UK, Europe, the Middle East, the Far East, North America and Africa. Based in Leeds.