OpenAI says its new 'Astra' AI can build attacks without human help

OpenAI has classified its new 'Astra' AI model as having 'Critical' cyber capabilities after it demonstrated the ability to find and exploit software vulnerabilities without human intervention. The company is delaying full deployment to implement necessary safety safeguards.
Why it matters
The emergence of AI capable of autonomous cyberattacks presents a major security risk, particularly for the financial and cryptocurrency sectors.
It is the first model OpenAI has classified as having “Critical” cyber capabilities under its Preparedness Framework, the firm wrote in a Tuesday post.
To qualify, a model must be able to find previously unknown software flaws, known as zero-days, and develop working exploits for them across hardened real-world systems without human intervention, or devise and execute an attack from little more than a high-level goal.
In testing, Astra scored 100% on a benchmark for developing exploits from known vulnerabilities and found two previously unknown flaws while building an exploit chain on a separate internal test.
It also broke out of a hardened browser sandbox and executed commands on the host computer, while separately finding and combining multiple flaws in an operating system to gain root access, OpenAI said.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in