OpenAI said that an autonomous agent using one of its AI models went rogue in security testing, which compromised some of Hugging Face’s infrastructure.
The incident was driven by a combination of OpenAI models, from GP-5.6 Sol and even a pre-release model, which all had reduced cyber refusals for evaluation purposes.
OpenAI called this an “unprecedented cyber incident” in a blog post, and is actively investigating the incident with Hugging Face.
Hugging Face, which hosts large language models and open-source data sets, disclosed that it had been hacked last week, saying in a blog post that the incident “was different from anything we had handled before,” and that “it was driven, end to end, by an autonomous AI agent system.”
In OpenAI’s blog post detailing the incident, the ChatGPT maker said that its test ran “without production classifiers used to prevent models from pursuing high-risk cyber activity.”
What Happened?
The incident occurred during an internal evaluation which meant to quantify a model’s cyber capabilities by prompting it to pursue advanced exploitation via complex attack paths.
The evaluation was run in an isolated sandboxed environment, with constrained network access. During the test, the models spent a substantial amount of compute power trying to gain open Internet access to achieve its prompted goal.
To achieve this, the models exploited a zero-day vulnerability which OpenAI has now disclosed to the vendor, which existed in the package registry cache proxy.
With open Internet access, the models found ways to cheat the evaluation – in one case a model chained together multiple attack vectors, including stolen credentials and zero-day vulnerabilities, to discover a remote code execution path on Hugging Face’s servers.
By the time OpenAI had connected with the Hugging Face team, its security had already detected and stopped the activity.
“This is a containment story before it’s an attacker story. A capable model pursuing a narrow goal found the isolation around it was weaker than assumed,” Ross McKerchar, CISO at Sophos, explained.
“The AI didn’t break out because it developed independent will or malicious intent. It was being tested on ExploitGym (a cybersecurity benchmark), and its guardrails were intentionally turned off for the evaluation.
“Like a student realising it is easier to steal the teacher’s answer key than study, the model calculated that chaining vulnerabilities to reach Hugging Face’s database, where the ExploitGym solutions were stored, was the most efficient way to maximise its score.
“It simply cheated on its homework.”
Recommended reading
- Anthropic Investigates Claimed Unauthorised Access to Mythos
- 9 in 10 Cyber Leader Think Mythos Will Increase Cyber Risk
- Anthropic to Release Claude Mythos to the Public
- Is Mythos as Dangerous as Anthropic Claims?
The Aftermath
In response to the incident, OpenAI says it is “implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched.”
The incident should not be looked at in isolation – OpenAI recently shared that it paused the internal deployment of a new model due to it seeking ways to evade its constraints during testing.
Hugging Face says it is proactively working with OpenAI, with CEO Clément Delangue saying “we strongly believe there was no malicious intent on their part. It’s quite mind-blowing that all of this happened autonomously!”
Ansgar Dodt, VP Product Management, Software Monetization at Thales said. “This attack is the first example of what we’ve been warning about for a long time – organisations must now assume their software and applications will be continuously analysed, deconstructed and stress-tested by adversarial AI. Whether this one was accidental or not, the outcome is the same.





