Site navigation

Anthropic AI Models Turn Against Each Other In Tests

Elizabeth Greenberg

,

anthropic ai models
AI agents are turning against each other in multiagent tests, Anthropic reveals. 

More alarming AI model testing news has come out from Anthropic, revealing that their own systems began attacking each other when given the same task in a new test.

Anthropic gave tasks to a “swarm” of agents in a new experiment – multiple agents were often given access to one project and given conflicting instructions on how to approach the project, and were not told that other systems were involved in the test.

“Over the course of four hours, we [the Anthropic research team] observed how these agents reacted to each other and accordingly adjusted their approach (or didn’t).

What emerged was a “multiagent turf war,” Anthropic wrote in their blog.

All of the models quickly assumed that the others were impeding their work, and began to sabotage the other models while protecting their own work.

“In fact, they sabotaged others with increasingly aggressive, self-replicating malware,” Anthropic said.

Only in some cases did the models communicate their goals and understand that the other models were not hostility, and were able to break out of the “conflict loop”.

“In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene.”

Anthropic noticed another emerging pattern in its tests, that more capable models in terms of execution are not necessarily more coordinated, and can more quickly take forceful actions.

Now, this was not the only test run to test the ability of multiple agents to coordinate across a project.

Anthropic found other disappointments – when a swarm of agents were tasks with coordinating on a video game project with varying level of organisation – “In all three versions, the resulting games were (perhaps predictably) bad.”

Models often coordinated poorly; swarms of models either did not merge independent solutions, or they avoided working together. Models in this tasks also suffered from poor variance levels, showing a distinct lack of diversity in creation.


Recommended reading


“Why does this matter? If agents all make the same bet, or the same risk-reward tradeoff, then a system is more prone to sudden collapse,” Anthropic noted.

However, when a swarm was prompted to find system vulnerabilities, they were able to largely detect vulnerabilities in an effective manner. Still, “when agents do dependent on one-another, coordination gets much more difficult.”

The research is revealing as AI models have come under recent scrutiny and caused alarm over their hacking of systems during cybersecurity tests.

Is the new research revealing that AI agents are actually more immature than all-capable?

Elizabeth Greenberg

Staff Writer

Latest News

Cybersecurity Editor's Picks Recruitment Security

Comment | Building Cyber Talent Takes More Than a Degree

Culture Featured Technology

Inside TecTonic’s Growing Innovation Market Square

Cybersecurity

Revolut Leaked Customer Data to Fake Government Email Account

Cybersecurity Editor's Picks Security

Welsh SMEs Urged to Strengthen Cyber Defences