Researchers fear safety disaster ahead of OpenAI’s Astra release

OpenAI has delayed the release of its new AI model, Astra, due to safety concerns regarding its opaque 'looped transformer' architecture. Researchers worry that the model's internal reasoning process is too difficult to monitor, potentially hiding dangerous behaviors.
Why it matters
It underscores the critical challenge of AI interpretability and the safety risks associated with increasingly complex, 'black-box' neural network architectures.
OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it “may be the single worst development for AI security/safety to date.”
Shortly after OpenAI said on Tuesday that it had delayed Astra’s release to work on safety issues, The Information reported that Astra shows far less of its “thinking” than other frontier AI models, sparking concern it could be dangerously hard to monitor.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in