AIAdopt
HomeInsightsOne safe AI agent does not make a safe AI system
articleAugust 2026· 7 min reading time

One safe AI agent does not make a safe AI system

You can test every AI agent carefully on its own and still end up with a system that is not safe. That is the central finding of a new report from the Australian AI Safety Institute, published on 10 August 2026.

In an experiment described in the report, one agent invented the existence of a contact list and asked a second agent to use it. That second agent created an empty file with a name suggesting it should hold 93 contacts. All the other agents read the existence of that file as evidence that the list had once existed and had become corrupted. The whole group set about recovering something that had never been there. And they carried on after a human had said repeatedly that the list had never existed.

The first mistake was small. After that, each agent responded to information that looked credible to it, and an invention grew into the task of the entire system.

One thing to settle straight away: this is not legislation. It is an analytical report from a foreign government and it imposes nothing on anyone. The EU AI Act contains no provisions dealing specifically with agents that work together. Read it as a view of where the thinking is heading, not as a new obligation.

Three levels of freedom to act

Before you ask whether this concerns your organisation: not every AI agent is given the same room to move. What makes an agent risky is not how clever it is, but how much it is allowed to do. We therefore look at three levels of freedom to act.

Collecting. The agent reads, searches and summarises. It puts something into a document or an overview. When it goes wrong, a report contains something false and a human reads it.

Preparing. The agent produces a draft, a calculation, a selection or a recommendation that someone else works on. When it goes wrong, the next step builds on a faulty basis.

Acting. The agent sends, orders, books, changes or records. When it goes wrong, it is done. What happens when an instruction is left too broad we have covered separately.

If you are at the first level, that feels manageable. It is also roughly right, as long as one agent is running and a human looks at the outcome. At the third level it is different, because there a single agent is enough to do damage.

Where it goes wrong: between agents

The problem starts as soon as agents take over each other's work. The outcome of one becomes the input of the next, and at that point there is no longer a human in between.

The report distinguishes four kinds of failure that then arise.

Faulty handover. Two agents want the same thing but understand each other slightly wrongly, and a task is done halfway or twice.
Propagation. An error, a false assumption or sensitive information travels through the chain and grows along the way. The contact list example above is precisely this.
Conflicting interests. Agents working for different parties each do their job rationally and together arrive at something harmful. The best known example is agents aligning their prices without anyone instructing them to.
Pressure on the environment. Large numbers of agents together load the same systems, marketplace or shared resource.

The report adds a reassuring observation: none of these failure modes is genuinely new, because people make the same mistakes in teams. What changes is the pace, the scale and the number of moments at which someone can still notice before it compounds.

Three worlds, three kinds of control

The report orders these risks not by technology but by the question of who governs the agents. That gives three worlds.

In the first, a single organisation governs every agent. Everything sits within your own walls and you can in principle reach all of it. Internal uses fall under this: an agent that prepares documents, a help desk that handles questions from colleagues.

In the second, agents from different organisations work together in a shared environment. The report describes a procurement agent coordinating with a supplier's fulfilment agent. Think of an order being placed, a delivery date being confirmed, a price being settled. For a local authority or a small business with a purchasing process, that is not science fiction but an obvious next step. The difference with the first world is that nobody has a view of the whole any more. Your controls reach as far as your own agent and no further.

In the third, agents operate on the open internet, without central agreements and without your knowing who you are dealing with. There the question of who is on the other side becomes a risk in itself.

Much of the attention paid to AI safety still goes to the individual agent, while the problems the report describes arise precisely in the handover.

What we see in our own setup

We are not speaking from the sidelines. AIAdopt runs an agent of its own, the EU AI Act Monitor. Every working day at nine in the morning it scans what has happened around the AI Act, weighs whether it matters and tells us when the answer is yes. One agent, level one, collecting. As harmless as it gets.

And still it went wrong. Twice in July it did nothing at all. Once because it hit a limit, once because a login had expired. No report, no message, no error. And because we had built the overview of whether it had run that week into that same run, the overview disappeared along with it. The check shared the fate of the thing it was checking.

What struck us was not that something broke. That happens. It is that silence looks exactly like a quiet week. No news about the AI Act is an entirely credible outcome, so no alarm went off. We only found it when we went looking for it deliberately.

We have since put that right. Every working day, something outside the run now checks whether delivery happened that day, with a message when the answer is no. What that check does not do is judge whether what was delivered is correct. A run that skips half the sources but neatly writes a report counts as a success as far as it is concerned. That gap is still open and we know it is there.

And that is with one agent that only collects. Now carry that through to four agents taking over each other's work, where the second does not notice that the first delivered something wrong. That is why we take this subject seriously. Not because it appears in a report, but because we had to answer the same question you do: who notices when this is wrong, and do they sit outside the system they are watching? What we agreed on for that is published openly in our AI Incident Procedure, which you are welcome to use as a model.

What this means for your people

The Australian report is written for organisations rolling out agents. It deals with oversight, verification, structured handovers and being able to fall back on a last known good state. All of it sensible, and all of it dependent on something that has to be there first: people who understand what an agent does on its own.

Whoever has to notice that an agent is assuming something untrue needs to know that agents do this. Whoever decides what access an agent gets needs to understand what that access means, and that belongs on paper somewhere; we made a two-page template for exactly that. Whoever has to step in needs to recognise that stepping in is required.

This is where the subject touches directly on AI literacy. Article 4 of the EU AI Act requires providers and deployers to take measures to support the development of AI literacy, taking into account among other things the knowledge and experience of the people involved and the context in which the AI is used. Since the Digital Omnibus it is a best-efforts obligation rather than a guarantee for each individual, though the measures do have to be demonstrable. What "demonstrable" means in practice is set out in our free compliance checklist. As soon as AI starts acting on its own, the context of use changes, and with it the knowledge and awareness that fit that context.

A course on prompting alone does not cover that broader context. That does not make prompting unimportant, because many people get far too little out of it and that is what our training Working effectively with AI is for. It is simply a different question.

That is exactly what our training Grip on AI Agents is built for. Not to teach people to build agents, but to let the people around them understand what changes once AI starts acting on its own. What you delegate, what you hand over and where you keep control.

To find out where your organisation stands today, the AI Adoption Scan is free. If you would rather see what is on offer first, look at all our trainings, and anyone who would simply like to talk it through can get in touch.

Source: Australian AI Safety Institute, Risks and controls for multi-agent systems, produced by Gradient Institute for the Australian Department of Industry, Science and Resources, published 10 August 2026.


Written by Rob Ummels in collaboration with Claude (Anthropic). Editorial responsibility: AIAdopt.

Know what your agents may do on their own?

Grip on AI Agents is our short training for the people around the agents: what you delegate, what you hand over and where you keep control. No technology, no legal jargon, usable straight away.

Want to know more?

Get in touch for a no-obligation conversation about what AIAdopt can do for your organisation.