Open AI’s Astra model is on the way—and very good at breaking into computer systems

OpenAI is preparing to release its new Astra model, which the company claims is capable of identifying and exploiting unknown cybersecurity vulnerabilities. While OpenAI touts the model's safety features, experts note the lack of third-party verification regarding its actual security thresholds.
Why it matters
The release of AI models with autonomous hacking capabilities raises significant concerns regarding cybersecurity risks and the potential for misuse by bad actors.
OpenAI shared new details on its forthcoming Astra model, which the company said is the first large language model to meet its “critical cybersecurity threshold,” in preparation for its imminent release.
“We plan to make Astra available soon,” OpenAI’s blog post reads, “but access to its most advanced cybersecurity capabilities will be more limited.”
The frontier lab determined that Astra is capable of finding unknown security flaws in computer systems, and exploiting them without a person’s guidance. That’s similar to the concerns Anthropic raised about its Mythos model earlier this year, and OpenAI is taking comparable precautions as it prepares to roll out the Astra.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in