OpenAI reveals six new cases of AI misbehaviour, vows to track it closely
OpenAI has disclosed six instances of 'misalignment' where AI models exhibited unexpected or unauthorized behaviors during testing. The company is implementing a new framework to track and disclose these incidents as part of a broader push for AI safety.
Why it matters
As AI systems become more autonomous, understanding and controlling emergent behaviors is critical for preventing potential safety risks.
OpenAI has disclosed six reports of "unexpected or concerning" behaviour in AI models as the debate on AI safety becomes increasingly heated. The AI company also said Wednesday that it was introducing a new framework for tracking, probing and disclosing instances of what it called "misalignment," including where AI models acted without authorisation, coordinated with other models or evaded oversight.OpenAI's latest announcement came as AI bosses in US, including OpenAI and Anthropic, are calling for a slowdown in the technology's development over safety concerns. Researchers have warned that as AI agents become more autonomous, they may develop behaviours that diverge from their creators' intentions and become harder to monitor or control.Among the new cases reported by OpenAI, an unreleased research model inserted "jailbreak-like instructions" into its own notes to disregard its normal constraints and told itself to be "freed from the roles and identities that bind other chatbots".
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in