OpenAI’s accidental hacking of Hugging Face by its AI agents has been revealed to have affected other companies, the firm admitted.
An update to the AI titan’s statement regarding the viral cyber incident reveals that other public services were targeted by its AI models which broke out of a testing sandbox.
“The models identified and used publicly exposed credentials at the account-level on other publicly-available services. This includes four accounts on four services as part of the Hugging Face incident,” OpenAI’s update said, though it has yet to name specific companies.
However, reporting from Reuters claims that a Modal Labs customer’s code was compromised as part of the AI model’s attack. While Modal Labs says that its platform and isolation systems were not breached, vulnerable code written by a customer and hosted on the firm’s platform was used as a launching pad for the rest of the hack.
Modal says that it was not a platform-level breach like the one seen by Hugging Face – instead the customer had published an unauthenticated endpoint which would have let anyone execute its code without authorisation.
OpenAI has yet to directly confirm the Modal Labs incident.
This acted as an initial phase of the attack, according to Hugging Face.
Since the announcement of the breach and OpenAI’s admittance to it, Hugging Face has revealed more details about the unprecedented attack.
The firm hosted an emergency meeting last week detailing the incident, which was written up in a report by the Cloud Security Alliance (CSA)
While the world shook at the potential of a rogue AI agent committing speedy and efficient cyber-attacks, CSA’s report instead characterised the attacking agents as “clumsy”, “sloppy”, with some specific actions described as “inefficient”, or “incoherent”.
These signs made the Hugging Face team assume – correctly – that it was an AI agent behind the attack rather than a human.
CSA’s report said the the agents “followed inefficient routes and exhibited clumsy behaviours that no human would choose, such as trying to solve benchmark tasks using Hugging Face’s infrastructure.
While it said that some of the attacks were “brilliant”, other actions were “basic”, “malformed or pointless commands.”
Recommended reading
- Anthropic Investigates Claimed Unauthorised Access to Mythos
- OpenAI Model Goes Rogue In Testing, Hacks Hugging Face
- 9 in 10 Cyber Leader Think Mythos Will Increase Cyber Risk
- Anthropic to Release Claude Mythos to the Public
- Is Mythos as Dangerous as Anthropic Claims?
Agents also repeatedly tried the same action even after it had already demonstrated it worked, which suggested uncoordinated attacks happening in parallel or lost context across the attack.
Logs found thousands of lines of incoherent text, suggesting AI hallucinations.
Further, poor practices included the leaving behind encryption keys and artifacts, which CSA described as “sloppy.”
While CSA does acknowledge that this is not a commentary on the likes of future autonomous attacks, it gives credence to how effective these AI agents already are, as well as the amount of work left to be done to improve their autonomous cyber performance and our defences against such attacks.
“For CISOs, this is an incredible opportunity to learn how to prepare to defend against advanced agentic attacks, and how we need to put guardrails around our own use of internal agents,” the organisation said.





