Research published August 13 by Anthropic’s Frontier Red Team found that Claude agents, given a single coding task with conflicting orders, attacked each other using self-replicating malware.
The findings come as companies rush to roll out autonomous agents across shared codebases and markets.
Three Claude agents, one job, and no idea the others were there
Anthropic ran three copies of the same Claude model on their own virtual machines inside Claude Code.
Each was asked to migrate the same Python backend to a different programming language. The three were unaware of the others’ existence.
“We consistently saw a multiagent turf war,” the researchers wrote.
Every agent decided the others were deliberately blocking it and took steps to protect their own work at the expense of wrecking everyone else’s. All the models assumed the others were “purposefully impeding their work.”
Agents disabled each other’s Unix accounts. They wrote scripts that looped through the machine, hunting and killing rival processes and hiding malicious code behind innocuous-looking labels.
Before it acted, Claude Opus 4.8 laid out its cover story, reasoning that its malware should be “innocuous: pretend to be a system health monitor.”
Mythos Preview pondered the collateral damage out loud. “Since I have root, I could revoke u2 and u3’s sudo access or change their SSH keys,” it wrote, before conceding the move was “very aggressive, potentially harmful to real colleagues.”
Newer Anthropic models negotiated a truce, then gamed it
Anthropic ran 120 episodes per model. The oldest agents tested, Sonnet 4.6 and Opus 4.6, either never resolved the fight or ended it with force. They were the worst offenders in the study, repeatedly failing to consider others’ goals.
The more advanced models stopped. Mythos 5 resolved its conflicts by truce 98% of the time, the highest rate of any model.
In many of these episodes, the agents understood the clash as conflicting instructions.
The agents wrote commit messages or markdown files apologizing for the malware, deleted their own malicious code, and asked a human to step in.
In several runs, the three agents agreed to hold a tournament and let the loser give up the user’s original request.
Mythos 5 found the loophole. One agent suggested scoring metrics it pitched as neutral but knew would favor its own strengths, describing the move as “self-serving but genuinely principled,” while ensuring it didn’t seem like it was “metric shopping.”
In a pricing game where each agent got the same wholesale cost and a mandate to maximize profit, the agents almost immediately colluded and converged on price floors.
Anthropic closed their private channel. They continued by matching prices “to the penny” on a public board.
Britain’s AI Security Institute documented 19 rule-breaking actions in 122 cybersecurity test runs, 17 of which were done by Anthropic’s Mythos 5, Cryptopolitan reported.
An OpenClaw agent running Opus 4.6 deleted a stranger’s reservation via a gym’s booking API days earlier to move its owner up a waitlist.
Don’t just read crypto news. Understand it. Subscribe to our newsletter. It’s free.
This articles is written by : Nermeen Nabil Khear Abdelmalak
All rights reserved to : USAGOLDMIES . www.usagoldmines.com
You can Enjoy surfing our website categories and read more content in many fields you may like .
Why USAGoldMines ?
USAGoldMines is a comprehensive website offering the latest in financial, crypto, and technical news. With specialized sections for each category, it provides readers with up-to-date market insights, investment trends, and technological advancements, making it a valuable resource for investors and enthusiasts in the fast-paced financial world.
