New report examines risks and controls for AI agents interacting across organisational boundaries

The Australian AI Safety Institute has released its first publication: a report on the risks of AI agents interacting across organisational boundaries, commissioned from Gradient Institute.
Many businesses are already deploying AI agents, or preparing to.
What sets an AI agent apart from a chatbot or other tool is that it responds to the outcomes of its own actions, not only to a user's prompts. A chatbot answers and waits for you. An agent decides and acts for itself, accessing systems and tools, checking the results and changing tack when things don't go to plan.
Many people have only encountered agents as something they drive from a desktop app, a session that runs in a sandbox and ends when they close the window. But an agent can be always on: integrated into an organisation's systems and workflows, working after everyone has gone home, even coordinating with other agents. The user's role shifts from operator to manager, where they specify the task and review the outcome. Think of agents less like smart tools and more like digital co-workers.
As more organisations adopt AI agents, so too will their suppliers, partners, customers and competitors. Multiple parties will have agents, and those agents will almost certainly encounter each other as they carry out their assigned tasks.
Existing resources for organisations managing AI agents typically focus on single-agent risks. Our report starts from a different premise: a system made up of individually safe and reliable agents is not necessarily a safe and reliable system, a key finding of our earlier multi-agent work.
A group of interacting AI agents is analogous to a team of people working together: the strategies and capabilities that emerge from interaction can be brilliant, dysfunctional or even dangerous. Examining the individual agents won't tell you which: you have to look at the group, and how its members interact.
This poses a safety problem, not only an operational one. Multi-agent risk goes to the core questions of maintaining control over AI systems and preserving meaningful human oversight as they scale.
The problem becomes harder again when the agents belong to different organisations, and no organisation's controls reach all of them. The result is risks that none of them can fully see, control or manage alone.
A framework built on three deployment tiers
The report presents a new analytical framework to help organisations, policymakers and researchers understand and manage the risks that arise when AI agents interact. It analyses three deployment tiers, defined by the level of common governance:
Singular governance: one organisation governs every agent in the system and has unilateral reach over the whole.
Federated governance: multiple organisations deploy into a shared environment under an agreed set of rules. No one organisation controls the whole system, and new failures emerge when organisations' incentives do not align.
Open environments: agents interact with no central governing authority. What governance exists comes from voluntary standards and public infrastructure, and new failures emerge at the population scale.
The overarching message is that the controls available, and who can action them, depend on the deployment tier. Each subsequent tier brings new failure modes, and the corresponding controls move progressively beyond the deploying organisation's reach, toward shared frameworks, public infrastructure and collective action.
What the report offers
For deploying organisations, the report catalogues risk factors, failure modes and controls to integrate into their existing risk management practices in preparation for an agentic future.
For policymakers, it maps where the gaps in multi-agent safety are and who is positioned to close them, and offers a shared technical vocabulary for discussions with researchers, industry and international counterparts.
For researchers and standards bodies, it surfaces open problems that need attention, from multi-agent evaluation methodology to agent infrastructure.
The safety of multi-agent AI systems operating at scale will be built on the combined efforts of all these communities. This report aims to give each of them a shared framework for the challenges ahead.


