US Publishers Demand Common Crawl Stop Scraping Their Content

Digital Content Next has issued a cease and desist letter to Common Crawl, demanding the organization stop scraping copyrighted publisher content for its datasets. The dispute highlights the growing tension between AI training data practices and intellectual property rights.
Why it matters
This legal challenge could fundamentally alter how AI companies source training data and how copyright is enforced in the age of generative AI.
Digital Content Next sent Common Crawl a cease and desist letter demanding it stop scraping publisher content and remove protected material from its datasets.
The article focuses heavily on the publishers' perspective and the 'infringement' narrative, though it acknowledges the technical nature of the dispute.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in