Is it legal to train AI models on copyrighted books? It’s complicated

This article explores the legal complexities surrounding the use of copyrighted works to train AI models, citing a recent court ruling involving Anthropic. While the court penalized the company for using pirated content, it affirmed that training AI on copyrighted data is generally analogous to human learning and thus lawful.
Why it matters
This legal precedent significantly impacts the future of generative AI development and the rights of content creators.
You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right?
“I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on,” Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch. “It’s very complex and there are a lot of raw feelings about what is happening, both for and against.”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in