Anthropic's 'Watermark' Text Adulteration in Claude Is a Perversion of Writing

The author criticizes Anthropic's decision to implement steganographic watermarking in Claude models. They argue that modifying text output to include hidden fingerprints compromises the quality and integrity of the writing.
Why it matters
This highlights the growing tension between AI safety/regulatory compliance and the desire for high-quality, unaltered generative AI output.
Manage GRC Faster with Drata’s Agentic Trust Management Platform
When I wrote this week about Anthropic’s announcement that all Claude models, worldwide, would soon begin “watermarking” everything they generate, including text, to comply with this EU regulation , we were left to speculate how this was going to work, because Anthropic offered not even a vague description of how it would work — despite the fact that the title of the announcement was, absurdly and insultingly, “ How Claude Marks AI-Generated Content ”.
My initial speculation was that maybe they’d hide invisible non-printing Unicode characters in the text. Just spitballing. Turns out that’s not what they’re going to do. What they’re going to do is apply a form of steganography, where the choice of words (or other token output) at inference time will leave fingerprints that can later, maybe, be detected probabilistically.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in