OpenAI has strengthened security controls around its Astra model after internal evaluations indicated it could potentially reach the highest cybersecurity risk category under the company’s Preparedness Framework.
The company has paused internal Astra activities that do not meet the strengthened controls while further benchmarking is carried out, after improvements in the model’s autonomous coding and cybersecurity capabilities meant OpenAI said it could no longer rule out classifying it as “Critical”.
Under OpenAI’s Preparedness Framework, the Critical category applies to models capable of independently identifying and developing functional zero-day exploits across a range of hardened real-world critical systems, or devising and executing end-to-end novel cyberattack strategies against hardened targets when given only a high-level objective.
Previous frontier models, including GPT-5.6-Sol, have remained within the High category.
OpenAI has not yet determined that Astra meets the Critical threshold, with full benchmarking still underway. The additional safeguards are therefore being introduced on a precautionary basis while the model’s capabilities continue to be assessed.
The measures include isolated testing environments, restricted network and tool access, encrypted model weights and sandboxed execution. OpenAI is also using chain-of-thought monitoring designed to identify and interrupt potentially high-risk activity in real time.
Government agencies and selected safety organisations are being given access to evaluate the model ahead of a wider release.
OpenAI said it still intends to release Astra broadly “once it satisfies the necessary safety and security requirements.”
The company said it is important to be transparent with the public and security community “about this potential shift in capabilities.”
The development comes amid growing focus on the role of increasingly autonomous AI systems in cybersecurity.
On 30 July, Palo Alto Networks’ Unit 42 documented an attack campaign in which a single operator used AI agents to carry out reconnaissance and exploitation activity across dozens of targets.
Microsoft had also released its first in-house cybersecurity model three days earlier. MAI-Cyber-1-Flash is a five-billion-active-parameter specialist model developed for cybersecurity work, as Microsoft moved towards building dedicated capabilities rather than relying solely on general-purpose frontier models.
The developments point towards a cybersecurity landscape in which AI models are increasingly being developed for both offensive and defensive tasks.
OpenAI said the same capabilities that could allow a model to independently identify vulnerabilities could also be applied defensively, allowing vulnerabilities to be discovered before attackers can exploit them.
The company plans to provide third-party evaluators and government partners with recommended security controls informed by its own work containing Astra.
Attention will now turn to the results of external evaluations and whether other major AI developers begin encountering similar capability thresholds as their frontier models become more capable.
OpenAI plans to continue evaluating Astra before determining whether the model can be released more widely and whether its cybersecurity capabilities formally meet the Preparedness Framework’s Critical classification.





