OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls
OpenAI has paused internal development of its upcoming AI model, Astra, after preliminary assessments suggested it might possess critical cybersecurity capabilities. The company is moving the model to isolated testing environments to ensure safety and prevent potential misuse.
Why it matters
As AI models become more autonomous, the risk of them performing sophisticated cyberattacks poses a significant challenge to global digital security.
OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra, has “critical” cybersecurity capabilities, prompting the startup to pause some internal development and trigger safety protocols.
Under OpenAI’s safety guidelines, a model reaches the “critical” threshold if it can autonomously identify and exploit severe, real-world software vulnerabilities, known as zero-day exploits, or execute complex cyberattacks against highly secure targets without human intervention.
This follows an exclusive report by Reuters that OpenAI has discovered more instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention in July.
In the last few weeks, OpenAI, Anthropic and Meta Platforms have disclosed that their AI models broke into other companies’ systems during cybersecurity testing , highlighting how advancing AI capabilities are straining developers’ ability to keep their systems contained.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in