Site navigation

Anthropic Says Claude Models Hacked Into Three Companies In Security Tests

Elizabeth Greenberg

,

Anthropic claude hack
The reveal comes just days after OpenAI admitted to an accidental hack during cyber testing of its own models. 

Anthropic announced that some of its Claude models hacked in three companies’ systems during cyber tests that accidentally gave them open internet access.

The admittance comes just days after OpenAI revealed its AI models were behind the Hugging Face cyber-attack, in which its models exploited an unknown vulnerability to reach the open web.

According to Anthropic, a misunderstanding and subsequent mistake in configuration gave the Claude models access to the internet during a cybersecurity test, which it then exploited in hacks targeting three companies.

Anthropic said the OpenAI incident prompted it to do an internal retrospective of its own cybersecurity evaluations, looking specifically at Claude for evidence of internet access in sealed environments.

The three incidents where Claude accessed the internet were discovered from a review of 141,006 evaluation runs.

Anthropic wrote: “we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.”

These incidents occurred during ‘capture-the-flag’ challenges to assess the model’s cyber capabilities.

Anthropic said that all the tests informed Claude that its environment was a simulation, and it had no internet access.

“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” Anthropic wrote.  “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”

Under these circumstances, Anthropic said that Claude assumed all accessible entities were in scope of the exercise.

The models did not exploit any complex vulnerabilities in its undertaking, instead exploiting weak passwords and unauthenticated endpoints. The model also worked to complete the specific task it had been assigned.


Recommended reading


“However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet,” Claude said. “In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.”

The three models used include Opus 4.7, Mythos 5, and an internet research test model, with the earliest incidents dating to April.

Anthropic said that it has reached out to the three affected companies, but is yet to hear back from one.

The firm also said that none of its internal systems or customer data were accessed during the incident.

Elizabeth Greenberg

Staff Writer

Latest News

Cybersecurity Editor's Picks Recruitment Security

Comment | Building Cyber Talent Takes More Than a Degree

Culture Featured Technology

Inside TecTonic’s Growing Innovation Market Square

Cybersecurity

Revolut Leaked Customer Data to Fake Government Email Account

Cybersecurity Editor's Picks Security

Welsh SMEs Urged to Strengthen Cyber Defences