OpenAI to launch new model with 'stronger safeguards' after hack
OpenAI is preparing to release a new AI model named Astra, which includes enhanced cybersecurity safeguards. This follows a security breach involving previous models and a broader industry push for safer AI development.
Why it matters
As AI models become more powerful, the industry faces increasing pressure to prevent them from being weaponized for cyberattacks.
ChatGPT maker OpenAI said Tuesday it was preparing to release its newest powerful model, known as Astra , after implementing "stronger safeguards" following a rogue cyberattack involving a different AI model.
The San Francisco-based artificial intelligence (AI) giant paused some of its model development for two weeks this summer after two models it was testing were involved in a security breach of software company Hugging Face.
Although Astra "was not involved" in the incident, OpenAI has beefed up its safety measures, the company said in a blog post.
"We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity," the blog said.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in