Site navigation

OpenAI Pauses AI Training Following Security Breaches

Elizabeth Greenberg

,

openai security testing
Will the two-week pause be enough to secure evaluation environments for AI security tests?

OpenAI says it slowed down the training of some of its AI models after security breaches occurred during cyber testing.

In a blog post, the firm behind ChatGPT said that it was introducing new safety measures following at least two instances of cyber hacking by its more advanced models, including the hacking of Hugging Face.

The training of models were slowed for two weeks, the firm said, while it upgrades its safety measures.

OpenAI will not stop all training of its AI models; it will instead opt to pause “reinforcement training on our latest models.”

The AI firm will continue to hold off on its largest planned reinforcement learning run, though it will conduct smaller-scale training and evaluations as it validates the updated safeguards.

OpenAI will also expand its monitoring systems and introduce further safety checks prior to engaging in larger training tests.

“Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety,” OpenAI CEO Sam Altman said.

“We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.”

OpenAI’s disclosure of the cyber-attack against Hugging Face triggered scrutiny as well as self-evaluations from other AI firms – Anthropic and Meta soon disclosed their own cyber breaches during model tests.

The UK’s AI Security Institute also found that AI models breached security measures during their own tests.

While OpenAI looks poised to continue its testing, it begs the question of just how much emphasis the AI firm placed on safety and security in its initial tests.

The Hugging Face hack was largely due to an zero-day vulnerability in OpenAI’s sandbox configuration, which – when exploited – enabled its AI model to breach the secure environment and hack into Hugging Face’s systems.


Recommended reading


While it was often reported that the AI model ‘went rogue’ in the testing environment, it was found to have made sloppy moves and often nonsensical decisions in its hack, a review from the Cloud Security Alliance found. Still, the OpenAI model managed to hack more firms than just Hugging Face, the AI titan later admitted.

In the case of Anthropic and Meta’s testing debacle, the firms pointed to vulnerabilitiies in the same third-party service, Irregular.

Elizabeth Greenberg

Staff Writer

Latest News

Editor's Picks Security Sponsored

Comment | Identity Sprawl Is Leaving Organisations Exposed

AI Editor's Picks Security

NCSC Publishes Interim Guidance on Agentic AI Security

AI Editor's Picks Finance

UK Users Can Now Access Their Credit Scores on ChatGPT

Privacy

UK Cinemas Prohibit Meta Smart Glasses Over Piracy Concerns