Major AI Developers Secretly Train Models on YouTube Transcripts
An investigation reveals that tech giants like Apple and Nvidia are using transcribed YouTube videos to train AI models without creators' permission. This practice directly violates YouTube's own platform rules.
An investigation by Proof News and Wired reveals that major AI companies are training their models on transcribed YouTube videos without permission from the content creators. Apple, Nvidia, Anthropic, and other tech giants use a dataset called YouTube Subtitles, which contains text from nearly 175,000 videos across 48,000 channels. This massive data collection happens entirely without the knowledge of the YouTubers who produced the content.
The dataset originates from EleutherAI, a group that built the collection to lower barriers to AI development for those outside of large tech companies. The YouTube Subtitles text serves as one component of a larger AI training repository known as the Pile. Apple utilizes the Pile to train its OpenELM model, while Salesforce uses it for an AI model that currently boasts over 86,000 downloads.
The impacted videos span a wide variety of genres, including news, education, and entertainment, featuring content from prominent creators like MrBeast and Marques Brownlee. This practice directly violates YouTube's own rules against harvesting content without explicit consent. To help creators check if their work is affected, Proof News provides a public search tool that scans the dataset for specific channels and videos.