Adobe Faces Class Action Lawsuit Over Alleged Use of Pirated Books in AI Training

An author files a proposed class-action lawsuit against Adobe, claiming the company uses pirated books to train its SlimLM artificial intelligence model. The legal action connects Adobe's training data to the notorious Books3 dataset.

An Oregon author named Elizabeth Lyon files a proposed class-action lawsuit against Adobe, accusing the software giant of using pirated books to train its SlimLM artificial intelligence program. The lawsuit claims that Adobe relies on a dataset called SlimPajama-627B, which allegedly derives from the controversial Books3 collection that contains thousands of copyrighted novels and guidebooks without permission.

Adobe states that its small language model uses an open-source dataset provided by Cerebras, but the legal complaint argues this dataset is a manipulated copy of the RedPajama dataset. Because RedPajama incorporates the pirated Books3 library, the lawsuit asserts that Adobe effectively uses stolen literary works to power its mobile document assistance tools.

This legal challenge reflects a growing trend across the technology sector as creators fight back against unauthorized use of their content. Similar lawsuits recently target major companies like Apple and Salesforce over their reliance on RedPajama, showing that the industry faces ongoing and expensive legal battles regarding how artificial intelligence algorithms acquire their training data.

Read More at the original source →