Cybersecurity researchers aren’t happy about the guardrails on Anthropic’s Fable

Cybersecurity researchers are criticizing Anthropic's new AI model, Fable, for overly aggressive guardrails that block legitimate security tasks. Experts argue that the model's keyword-based filtering incorrectly flags benign software engineering requests as cybersecurity threats.
Why it matters
This illustrates the tension between AI safety measures and the practical utility of tools for cybersecurity professionals.
Anthropic released its latest model Fable on Tuesday, billing it as a public and limited version of its powerful and much-hyped cybersecurity model Mythos.
The article presents the complaints of researchers while acknowledging Anthropic's stated safety intentions.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in