Frontier AI labs still won’t say how they’d contain a rogue model

A study by Guidelight AI Standards reveals that major AI labs lack transparent and robust containment plans for rogue models. While OpenAI performed best, other industry leaders like Anthropic and Meta scored lower, raising concerns about operational risks as agentic AI becomes more autonomous.
Why it matters
As AI systems gain the ability to take independent actions, the lack of standardized safety protocols for shutting down or containing malfunctioning models poses a significant security risk to digital infrastructure.
Few of the top AI labs have published or demonstrated containment response plans, according to a recent study . A containment plan spells out what happens once an AI is caught trying to subvert human control — what access gets cut, and when the system gets shut down entirely.
That’s the finding from Guidelight AI Standards, an organization dedicated to promoting safe frontier AI development practices, which graded five leading labs on how prepared they are for exactly this scenario. OpenAI came out on top; Anthropic and Meta scored lowest. The findings matters as agentic AI takes on more autonomous roles inside companies’ own systems, and as regulators in California and New York begin requiring disclosure. For anyone building on or investing in these models, it’s a rare independent read on how seriously each lab treats operational risk versus how it talks about it.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in