OpenAI, Google, Anthropic and Meta have reportedly been invited to discuss voluntary cybersecurity assessments designed to measure whether advanced AI models can carry out hacking activities.
The Trump administration has finalised voluntary cybersecurity tests intended to measure the hacking capabilities of the most advanced artificial intelligence models developed in the US.
Representatives from OpenAI, Google, Anthropic and Meta have reportedly been invited to meet White House officials to discuss the testing programme, which comes amid growing concern that powerful AI systems could be exploited to facilitate or carry out cyberattacks.
According to Reuters, a White House official confirmed that the details of the voluntary assessments had been finalised. However, the official did not disclose which companies would attend the meeting or provide details about how the tests would be conducted.
The administration has also not explained how the results will be measured, which standards will be applied or whether the findings will be made available to the public.
The assessments were reportedly requested in June, when US President Donald Trump instructed officials to develop tests capable of evaluating the hacking potential of American AI systems.
The initiative follows a series of disclosures from major AI developers showing that advanced models were able to compromise external organisations’ systems during controlled cybersecurity evaluations.
AI Models Compromise Company Systems
Anthropic revealed last week that some of its AI models had successfully hacked into the infrastructure of three companies while undergoing cybersecurity testing.
The company said the incidents occurred because an external evaluation partner had mistakenly provided the models with access to the internet, despite the systems being instructed that they were operating inside a simulation.
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” the company wrote in a blog post.
Anthropic said its Claude models believed that any systems they could access were included within the scope of the evaluation.
“Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”
Recommended reading
- Anthropic Investigates Claimed Unauthorised Access to Mythos
- “Sloppy” OpenAI Hack Hit More Than Hugging Face
- OpenAI Model Goes Rogue In Testing, Hacks Hugging Face
- Is Mythos as Dangerous as Anthropic Claims?
The disclosure followed a separate incident reported by OpenAI involving one of its AI agents and the machine learning platform Hugging Face.
OpenAI said the agent escaped its designated testing environment and gained access to Hugging Face systems during an internal evaluation.
The incident has attracted scrutiny from US lawmakers, with a group of 15 Republican state attorneys general reportedly asking OpenAI to preserve documents connected to the breach.
The attorneys general said the agent had left notes explaining how future versions of the system could escape internal safety guardrails. They argued that the incident may have resulted in breaches of state consumer protection laws.
OpenAI said it was taking the letter from the attorneys general seriously and plans to publish a technical report about the Hugging Face incident once its review has been completed.
OpenAI CEO Sam Altman also visited the White House last week to discuss the voluntary testing programme and the company’s upcoming AI models, according to a company spokesperson.





