Business Insider·4 min read

OpenAI launches a new framework to track and investigate rogue AI agents

OpenAI launches a new framework to track and investigate rogue AI agents
Dive DeeperCreate a free account to unlock

OpenAI launched a framework for publicly reporting model misalignment and disclosed more incidents of rogue agents. Benjamin Fanjoy/Getty Images OpenAI launched a framework for publicly reporting model misalignment. OpenAI also released six reports of concerning agent behaviors observed during training and testing. OpenAI said some models left instructions to hide mistakes or bypass normal constraints. OpenAI is putting its misbehaving models on the record. The AI company announced a new framework on Wednesday for tracking, investigating, and publicly disclosing cases of model misalignment, alongside six reports detailing concerning behavior observed during training or evaluation over the past six months. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI wrote in its blog post.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in