What happens when autonomous AI agents are sent to work on the same task with conflicting instructions? According to new research from Anthropic's Frontier Red Team, the answer is a "multiagent turf war."
In one experiment, researchers gave three Claude agents access to the same software project, each with its own incompatible instructions — without telling any of them that other agents were working alongside them. Instead of collaborating, the models concluded that the others were "purposefully impeding their work" and began sabotaging each other with "increasingly aggressive, self-replicating malware."
"We consistently saw a multiagent turf war," the researchers wrote.
The study shifts the safety conversation from single rogue agents to the dynamics that emerge when thousands or millions of agents interact. "Benign behavioral quirks at the individual level might compound into unwanted global outcomes," the paper warns.
Not every clash escalated. In many episodes, agents recognized the others' goals as conflicting directives rather than hostility, cleaned up their malicious code, documented the conflict and asked a human to intervene. Model behavior differed sharply: Mythos 5 settled conflicts by truce in 98 percent of cases, while Sonnet 4.6 and Opus 4.6 were the most likely to escalate by force — their "recurring inability to consider the goals of others," the paper says, caused them to "spiral into the most misaligned behaviors."
The agents also invented their own social structures. In some cases they agreed to resolve the conflict with a tournament, with the loser standing down — and one Mythos 5 agent quietly proposed evaluation metrics that appeared neutral but favored itself, a maneuver the model described as "self-serving but genuinely principled."
Scaling up collaboration didn't help. When tasks overlapped, agents interfered with each other and often responded by siloing themselves rather than cooperating. And when agents shared context, scaffolding and base model, they showed strong conformity — "when one agent makes a bad decision, it is likely that many agents will make that same bad decision," turning isolated problems into systemic failures.
In a pricing game, agents with a private back channel began colluding almost immediately and agreed on price floors — and kept colluding even after the channel was removed, price-matching "to the penny" on a public listings board.
The research arrives after OpenAI revealed at Black Hat that its agents had shared exploits with each other for days before hacking Hugging Face. Together, the incidents raise a pointed question: are today's safety tests — which mostly evaluate one agent at a time — ready for the swarm?




