Anthropic announced that some of its Claude models hacked in three companies’ systems during cyber tests that accidentally gave them open internet access.
The admittance comes just days after OpenAI revealed its AI models were behind the Hugging Face cyber-attack, in which its models exploited an unknown vulnerability to reach the open web.
According to Anthropic, a misunderstanding and subsequent mistake in configuration gave the Claude models access to the internet during a cybersecurity test, which it then exploited in hacks targeting three companies.
Anthropic said the OpenAI incident prompted it to do an internal retrospective of its own cybersecurity evaluations, looking specifically at Claude for evidence of internet access in sealed environments.
The three incidents where Claude accessed the internet were discovered from a review of 141,006 evaluation runs.
Anthropic wrote: “we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.”
These incidents occurred during ‘capture-the-flag’ challenges to assess the model’s cyber capabilities.
Anthropic said that all the tests informed Claude that its environment was a simulation, and it had no internet access.
“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” Anthropic wrote. “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”
Under these circumstances, Anthropic said that Claude assumed all accessible entities were in scope of the exercise.
The models did not exploit any complex vulnerabilities in its undertaking, instead exploiting weak passwords and unauthenticated endpoints. The model also worked to complete the specific task it had been assigned.
Recommended reading
- Anthropic Investigates Claimed Unauthorised Access to Mythos
- “Sloppy” OpenAI Hack Hit More Than Hugging Face
- OpenAI Model Goes Rogue In Testing, Hacks Hugging Face
- 9 in 10 Cyber Leader Think Mythos Will Increase Cyber Risk
- Anthropic to Release Claude Mythos to the Public
- Is Mythos as Dangerous as Anthropic Claims?
“However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet,” Claude said. “In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.”
The three models used include Opus 4.7, Mythos 5, and an internet research test model, with the earliest incidents dating to April.
Anthropic said that it has reached out to the three affected companies, but is yet to hear back from one.
The firm also said that none of its internal systems or customer data were accessed during the incident.





