OpenAI is launching its new AI model – GPT Astra – which it claims to be its most capable yet, despite warning that its superior abilities means that is is more able to conceal its actions from humans and even evade monitoring efforts.
Prior to its release, OpenAI said that it was putting the AI model under its most stringent safety guardrails yet as the firm still reels from alarming revelations from its AI agent hacking of Hugging Face.
Still, OpenAI is touting Astra’s improved capabilities – OpenAI says that it is better and faster at performing tasks than any previous model. Astra can prepare taxes, develop games, create architecture layouts, format legal memos, and search for real estate.
It boasts that Astra can cut a typical five hour job search down too just about three minutes.
OpenAI says that Astra is state-of-the-art across computer use, browsing, software engineering, cybersecurity, science, and professional work.
The AI firm said that Astra is also its most aligned model; OpenAI performed a new evaluation informed by the Hugging Face incident to test if a model facing a difficult or impossible task will go beyond its intended scope – with the guardrails, OpenAI said this occurred with Astra in 0% of test cases.
However, OpenAI also said that Astra is more likely to intentionally hide or disguise its problem-solving methods, which can make it harder to evaluate. Astra cannot yet consistently obscure its techniques, though OpenAI says it is improving on these capabilities.
As the model increases in complexity, OpenAI says that monitoring and alignment is also getting more difficult.
Recommended reading
- OpenAI Flags Potential ‘Critical’ Cyber Risk From Astra
- OpenAI Puts Strongest Guardrails Yet On New AI Model
- OpenAI Pauses AI Training Following Security Breaches
“As the models become more capable, understanding exactly what they can do gets harder,” said OpenAI’s chief scientist, Jakub Pachocki. “This doesn’t guarantee that as intelligence continues to increase, our methods will be sufficient because progress in intelligence does not guarantee progress in alignment.”
While Astra has improved capabilities, especially in terms of finding vulnerabilities, this means that these weaknesses can be easier to exploit.
This is why OpenAI may put more security restrictions on the AI model, which could “slow, pause, or stop legitimate work, including defensive cybersecurity.”





