OpenAI says it slowed Astra model development over security concerns

OpenAI has paused development on its upcoming 'Astra' model after internal testing revealed it could independently execute cyberattacks. The company is applying its 'Preparedness Framework' to ensure safety before further progress.
Why it matters
This highlights the growing tension between rapid AI advancement and the potential for catastrophic security risks, marking a rare public admission of safety-related development delays.
OpenAI said Friday it has suspended work on some aspects of its upcoming model Astra after an internal review found it had made significant advancements in agentic coding and cybersecurity — enough to warrant concern over its capabilities.
OpenAI said in a blog post Friday that this model, which is still in development, reached its “critical cybersecurity threshold,” meaning it could independently identify and carry out cyberattacks against traditionally well-protected real-world systems. Under the company’s “Preparedness Framework,” which it created in 2023, this triggered additional safeguards.
“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI wrote. “Astra is an upcoming model, and was not involved in exploiting Hugging Face.”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in