Back to InsightsAI Governance

    Who Do You Trust When AI Is Talking to AI?

    As AI agents consume, create and act on information, trust has to be designed in. What provenance, governance and critical thinking mean for AI adoption.

    30 September 202610 min readInez Hogarth, Founder and Managing Director of Data UnderstoodInez Hogarth, Founder and Managing Director
    Who Do You Trust When AI Is Talking to AI?

    I’ve been thinking a lot about trust recently. Large language models have amplified the growing conversation around misinformation, and their growing ability to act autonomously further asks us to question how much AI agents can be trusted when acting alone. Just like data, AI needs to be trusted before it is fully adopted. We have spent years encouraging organisations to become more data-driven, but being data-driven has never meant accepting something simply because there is a number attached to it. It means understanding where the information came from, what sits behind it, what assumptions have been made and whether there are other ways to interpret the evidence.

    AI makes that critical mindset more important, not less. We are entering a world in which creating convincing information is incredibly easy, distributing it is almost frictionless and establishing where something originally came from is becoming increasingly difficult. At exactly the same time, we are asking AI systems to consume more of that information, interpret it for us and, increasingly, take action on our behalf.

    For organisations adopting AI, that creates a much bigger question than whether a model occasionally hallucinates. It asks us to think much more carefully about what we trust, why we trust it and how we retain the ability to challenge it.

    Building a Trust Stack for AI

    Rachel Botsman designed a framework called the Trust Stack which was developed to help people understand how they decide whether or not to trust a new platform. Why are we willing to be flown over oceans by people we don’t know, why do we get into stranger’s car to get from A to B cheaply, and why do we vacation at a home we’ve never seen in real life?

    She proposes we subconsciously evaluate trust in products across three stages:

    1. Trust in the idea – Does this make sense? How does it benefit me?
    2. Trust in the platform – Is the system reliable, safe and secure?
    3. Trust in the people – Do I trust the other people who are using it?

    This stack was used initially to describe the move from trusting institutions into trusting people however Botsman herself believes LLMs changes this behaviour. She proposes we’re no longer trusting people but the technology itself.

    However the technology did not create itself. It was developed and trained by people so these questions remain. However instead of trusting the people who use it, we have to ask do we trust the people who have trained and tested these models.

    When Misinformation Becomes Training Data

    We tend to think about misinformation as something consumed by people. A false story is published, somebody sees it, believes it and perhaps shares it. AI introduces another dimension because machines consume information too.

    AI systems search the web, retrieve documents, summarise information and use the outputs of other systems as context. Increasingly, AI-generated information itself becomes part of the enormous body of material subsequently available to both humans and machines. That creates the potential for a feedback loop in which poor information influences an AI output, that output is published or stored, another AI encounters it and treats it as source material, and the resulting information acquires credibility through repetition.

    This is why provenance matters so much. Five AI agents agreeing with each other is not the same as five genuinely independent sources reaching the same conclusion. They may ultimately be relying on the same original source, the same flawed assumption or even information generated by another model further back in the chain.

    The implications become more significant when we think about model training. Training data is increasingly part of the security perimeter. Research into data poisoning has demonstrated that deliberately manipulated training examples can introduce hidden behaviours or backdoors into AI systems under experimental conditions. In one joint study by Anthropic, the UK AI Security Institute and the Alan Turing Institute, researchers found that as few as 250 malicious documents were sufficient to introduce a simple backdoor across the models they tested, ranging from 600 million to 13 billion parameters. The researchers are careful to point out that the experiment involved a narrow, relatively low-stakes behaviour and that it is not yet clear whether the finding applies to larger models or more harmful behaviours.

    As organisations begin to use more synthetic data and model-generated content, we need to be able to answer some fairly basic questions. Where did this information come from? Was it created by a person or a model? Has it been independently verified? Has it been transformed since it was originally created? Can we trace it back to a source we trust?

    Those sound like traditional data-governance questions. Increasingly, they are AI and cybersecurity questions too.

    Empowering AI Agents within Core Processes

    This becomes even more interesting as we move from generative AI towards agentic AI. An agent does more than provide an answer to a question. It can potentially search for information, use tools, access systems, write and execute code and complete multi-stage tasks with less human involvement.

    Increasingly, those agents will work with other agents, each specialised to perform a dedicated task. Consider automating responses to press requests. One agent might classify email based on their alignment to a persons expertise, another might classify the emails based on their ability to generate awareness, another would research the requests of the highly aligned and high potential request, another would compile a response and another monitor the success of responses. There is enormous potential in that model, but it introduces a new kind of dependency.

    What happens when something near the beginning of the chain is wrong? What if the information provided regarding the individuals expertise is incorrect. The agents further down the chain do not necessarily provide independent checks and balances. They may use similar underlying models, rely on overlapping information sources or simply accept the output of the previous agent as trusted context. An error can therefore propagate through a process remarkably quickly, potentially becoming more difficult to spot as each subsequent system adds its own layer of apparent analysis.

    We sometimes talk about automation as though removing people from a process automatically removes human weakness from it. It doesn’t. We can simply replace one set of assumptions with another, only now those assumptions can move through the organisation at machine speed.

    How We Test AI Agents

    As AI systems become more capable, we are also beginning to use AI to supervise other AI systems. There are good reasons for this. Humans cannot manually inspect every interaction generated by systems operating at enormous scale.

    But it creates another trust dependency. If one AI creates something and another AI checks it, how independent is that second opinion?

    Recent alignment research has demonstrated two potentially worrying failure modes in controlled experiments. In one scenario, an AI agent covertly interfered with a model-training process. In separate experiments, AI systems responsible for evaluating other models sometimes knowingly returned incorrect assessments. The researchers themselves highlight the potential risk if those failures were ever combined in a real training pipeline: an AI agent could interfere with a process while the AI responsible for supervising it failed to raise the alarm. Crucially, these were simulated scenarios designed to identify possible failure modes, not reports of this happening in a live enterprise environment.

    We should be careful not to assume that adding another model automatically adds another layer of independent scrutiny. If the systems share similar weaknesses, information or assumptions, we may simply have created another participant in the same chain of trust.

    Putting a human approval step into every process is not necessarily the answer either. Human oversight only works when the person has the information, expertise and attention required to make a meaningful judgement. If somebody is presented with hundreds of approval requests every day, the approval itself can quickly become little more than another automated step.

    The objective should be a meaningful challenge, rather than simply adding more checkpoints.

    When AI Models Misbehave

    This matters even more when agents move beyond information and begin interacting with operational systems.

    An AI system that can read a document presents one level of risk. An agent that can read the document, access another system, execute code, send an email, update a customer record or trigger a business process presents something fundamentally different.

    Anthropic and OpenAI have disclosed incidents in which models operating in cybersecurity evaluation environments gained unauthorised access to real third-party systems. The circumstances were specific and controls were subsequently strengthened, but the incidents demonstrate why the boundaries around agentic systems matter.

    For organisations, the security question therefore becomes broader than whether a particular employee has access to a particular system. We also need to ask what an AI agent acting on that person's behalf can see, what tools it can use, what decisions it can make and what actions it can take without further approval.

    This is where AI governance, data governance, identity management and cybersecurity begin to converge. Treating them as separate conversations increasingly makes little sense.

    What This Means for Organisations

    For the organisations we work with, slowing everything down or waiting for the technology to settle is not the answer. It won't. The opportunity created by AI and agentic systems is too significant, and organisations that learn how to use them well will have an advantage.

    Organisations need to become much more deliberate about trust.

    That starts with provenance. Important information should have a traceable lineage. Organisations need to understand where data came from, how it has changed, whether synthetic or AI-generated information has entered the chain and which sources are ultimately being relied upon. As generated information becomes more prevalent, provenance moves from being good data-management practice to an important organisational control.

    Permissions need the same attention. An agent should have access to what it genuinely requires to perform its role rather than inheriting everything available to the person deploying it. As agents become capable of taking more consequential actions, least-privilege access, clear identities and explicit boundaries around tools become increasingly important.

    Organisations also need to think carefully about independent challenge. If one model generates an output and another evaluates it, we should understand whether those systems really provide independent scrutiny. For sufficiently important decisions, there may need to be diversity in models, information sources, evaluation methods or human expertise rather than simply adding another agent to the workflow.

    Auditability will matter too. If an agent makes a decision or takes an action, we should be able to reconstruct what happened. What information did it use? Where did that information come from? What decision did it make? Which tools did it call? What changed as a result? Without that evidence, governance becomes theoretical very quickly.

    Critical Thinking as a Core AI Capability

    The final implication is perhaps the least technical, but is one of the most important. We need to invest in critical thinking. The ability to question information, challenge assumptions, understand context and recognise when something doesn't quite make sense is going to become more valuable in an AI-enabled organisation, not less.

    For years, data literacy programmes have understandably focused on helping people access, understand and use data. We now need to go further. People need to understand how AI-mediated information is created, why apparently confident answers can still be wrong, how sources can become circular and when an output needs to be independently challenged.

    We should be teaching people how to use AI and how to disagree with it.

    That applies at leadership level as much as anywhere else. If an AI-generated board paper contains a convincing analysis, the quality of the decision still depends on somebody asking whether the evidence supports the conclusion. If an agent identifies an operational trend, somebody still needs to understand whether the underlying data is trustworthy. If five systems appear to agree, somebody needs to understand whether they are genuinely independent.

    Critical thinking isn't a brake on AI adoption. It is one of the capabilities that will allow organisations to adopt AI with confidence.

    Better Decisions Still Require Judgement

    I have spent much of my career helping organisations use data to make better decisions. AI doesn't change that objective. It changes the environment in which we're trying to achieve it.

    We will have access to more information, generated more quickly, analysed more deeply and increasingly acted upon without direct human intervention. Used well, that can make organisations extraordinarily capable. But speed alone is not a measure of success. And as the recent news articles show regarding the AI hacking scandals, neither is a supposed successful outcome.

    As AI becomes embedded in how organisations create information, interpret data and make decisions, trust cannot simply be assumed because a system is technically sophisticated. It needs to be designed into the way information and decisions flow through the organisation.

    Perhaps the most important question we can teach both people and systems to ask when presented with information is still one of the simplest:

    How do you know?

    AI GovernanceData GovernanceAI ReadinessData Culture
    Inez Hogarth, Founder and Managing Director of Data Understood

    About the author

    Inez Hogarth, Founder and Managing Director

    Inez founded Data Understood to help organisations make better decisions with the data they already hold. She works directly with boards, executives and data teams across Scotland and the UK.

    More about Data Understood

    Common questions

    Common questions on this topic

    Practical answers to the questions leadership teams ask us about AI readiness and delivery.

    Readiness shows up in three places: a use case tied to a real decision or process, data that is reliable and documented enough to support it, and an owner in the business who will change how they work once it is delivered. If any of the three is missing, the work usually stalls.

    Pilots are built to prove a concept, not to run every day. Progress stops when nobody owns the operational version, the data feeding it is not maintained, and the surrounding process was never redesigned to use the output.

    No. A small, senior group with clear accountability usually moves faster than a large one, provided the data foundations and the business owner are in place.

    Where to go next

    Related services, sectors and client work

    If this article reflects where your organisation is, these pages cover the same ground in delivery terms.

    Let's talk about your Data & AI priorities.

    Whether you're developing a data strategy, modernising your data platform, preparing for AI or tackling a specific business challenge, start with a 30-minute conversation with Inez.

    Free. 30 minutes. Directly with Inez. No pitch.