Court Rules Anthropic's AI Training Constitutes Fair Use
A federal judge rules that Anthropic's use of lawfully acquired books to train its AI models qualifies as fair use, marking a major victory for the AI industry. However, the court rules against Anthropic regarding its separate creation of a massive internal library using millions of pirated books.
A federal judge issues a pivotal ruling in the ongoing copyright lawsuit against Anthropic, deciding that training large language models on lawfully acquired copyrighted books counts as fair use. This specific determination represents a significant legal victory for the artificial intelligence company as it defends itself against a class action lawsuit brought by three book authors on behalf of millions of writers. The decision immediately impacts how the tech industry and copyright holders view the boundaries of using published materials for computational research.
Despite this major win on the fair use question, Anthropic loses on the separate issue of how it initially built its massive internal dataset. Judge Alsup rules against the company for downloading millions of copyrighted books from known pirate sites like LibGen and Pirate Library Mirror to create a permanent "central library." The court details how Anthropic acquired at least five million books from LibGen and two million from PiLiMi, alongside purchasing and scanning physical books to amass a repository intended to be retained forever.
The mixed ruling leaves several complex issues unresolved as the litigation continues to move forward. The court still needs to address the specific legal implications of Anthropic using those pirated works from the central library for actual AI training, which remains a distinct question from the fair use of lawfully acquired texts. Legal analysts and advocacy groups continue to monitor the case closely to understand the full impact this decision will have on researchers, authors, libraries, and the future of artificial intelligence development.