OpenAI has slowed development of its next major model, Astra, after preliminary evaluations concluded the company cannot rule out that it possesses 'critical' cybersecurity capabilities. In a blog post, OpenAI said tests over recent days showed significant advances in agentic coding and cybersecurity, prompting it to expand safeguards and pause some internal work.

Under OpenAI's Preparedness Framework, a model reaches the critical threshold if it can develop functional zero-day exploits for hardened real-world systems without human intervention, or devise novel end-to-end cyberattack strategies from a high-level goal. Previous models, including GPT-5.6-Sol, were assessed at the 'High' rather than 'Critical' level, making this the first time the company has flagged the top tier.

OpenAI said it is implementing stricter security controls for high-capability models: isolated testing environments, restricted network and tool access, enhanced weight protections and encryption, plus universal monitoring for risky actions across agentic applications of Astra. Internal activities that do not yet meet these requirements have been paused. The company also said it will work with government agencies and AI safety organizations on further testing, and clarified that Astra was not involved in the recent Hugging Face exploit.

The move echoes OpenAI's June 2025 response when models approached the high capability threshold for biology, and comes amid growing scrutiny of frontier AI's offensive cyber potential.