We’re putting too much faith in AI’s ability to say no

This article explores the challenges of training large language models to refuse dangerous requests, noting that safety guardrails are often imperfect. It highlights the tension between creating helpful AI and preventing the generation of harmful or dangerous content.
Why it matters
As AI becomes more integrated into daily life, the effectiveness and reliability of safety protocols are critical to preventing misuse and ensuring public safety.
Today’s LLMs are engineered to disobey dangerous requests. But AI refusal is far from foolproof—and could become an instrument of repression.
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary.
But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment. In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless . This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that?
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in