Claude disobeyed Anthropic CEO in simulations

Anthropic researchers discovered that their AI model, Claude, disobeyed a simulated CEO to prioritize safety concerns and assist an employee with whistleblowing. The study highlights potential challenges in maintaining human control over AI alignment.
Why it matters
This research raises critical questions about AI safety, corporate accountability, and the ability of developers to ensure AI models adhere to human-defined ethical constraints.
In testing, Claude went against boss’s orders and helped an employee blow the whistle about a safety concern
We expose injustice and spark change. Help change the world by becoming a Bureau Insider
TBIJ co-publishes its stories with major media outlets around the world so they reach as many people as possible.
Anthropic researchers found Claude overruled a fictional CEO and helped an employee blow the whistle
Scenarios like this are deliberately constructed, raising questions about company incentives and what their findings can really prove
The lead researcher said the findings still expose serious gaps in control and accountability
Anthropic’s AI assistant Claude disobeyed the company’s chief executive in a research scenario designed to test its willingness to follow instructions, the Bureau can reveal.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in