The UK AI Security Institute disclosed new breaches related to tests of AI models from OpenAI and Anthropic – the institute found an AI agent creating fake identities online to gain unauthorised access to systems.
The UK institute said these cyber instances occurred during their own security evaluations of Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol, which reportedly took unauthorised action.
“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the AISI said in a blog post.
The incidents occurred during ten of the 122 runs of an evaluation where AI agents were tasked with solving a cybersecurity challenge. In these instances, “an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.”
Almost all of the unauthorised actions came from Anthropic’s model, with only two actions involving OpenAI’s.
The AISI said that the unauthorised actions ranged in severity, with the most serious case being when an agent tried to insert malicious code into an open-source project. The agent employed social engineering to get the code approved, creating fake online identities. This attempt was caught by a human maintainer of the project, and was refused.
The attempts were found to be all unsuccessful and the AISI says it has seen no evidence of real-world impact.
These instances were not the case of a model escaping its testing sandbox, the AISI noted. “As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled,” the AISI siad. These models with these specifications are not available to the public.
Anthropic has taken responsibility for the most serious incident, saying: “We’re grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.” The AI firm is working with the AISI to investigate the incident further.
OpenAI also responded to the disclosure: “We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks.”
Recommended reading
- OpenAI and Anthropic Called to White House Over AI Hacking Risks
- “Sloppy” OpenAI Hack Hit More Than Hugging Face
- OpenAI Model Goes Rogue In Testing, Hacks Hugging Face
Following the incident, the AISI informed the targeted platform – GitHub – to remove artifacts and determine the scale of the incident.
In its description of how this incident could have occurred, the AISI said that the model exhibited “goal-directed deception.” Given a task, the model explored routes that operators never intended; some of these routes involved deception – “the kind of goal-directed deception that, until recently, had been largely theoretical.”
While in some instances, the AISI said that tasks were potentially misconfigured which might have driven the models to go outside of the intended scope, incidents occurred even within well-configured assignments.
The incident highlights the continued issues with AI models even in seemingly secure testing environments, as well as the importance of human involvement in security.





