December's AI Crisis Exposes the Ethical Failures of Unconsented Data Scraping
A series of controversies in December, including a major lawsuit from the New York Times and the discovery of harmful content in training datasets, highlights the severe risks of building AI on unconsented data.
December brings a wave of troubling revelations about the generative AI industry as companies face severe backlash over their data collection practices. Major AI models rely on massive datasets built by scraping copyrighted texts, images, and videos from the internet without the knowledge or consent of the original creators. This approach treats intellectual property like stolen goods in a pawn shop, creating a lucrative but fundamentally unethical foundation for leading AI technologies.
The most prominent blow comes from the New York Times, which files a massive copyright infringement lawsuit against OpenAI and Microsoft. The lawsuit alleges that ChatGPT relies on millions of stolen articles to generate content, accusing the companies of building a business model based on mass copyright infringement. Alongside this legal battle, image-generator Midjourney releases an update capable of reproducing near-exact frames from popular movies, further demonstrating how AI regurgitates stolen creative work.
Perhaps the most disturbing discovery involves Stanford researchers finding child abuse images hidden within the LAION dataset, a common resource used to train popular AI models. This horrific finding proves that uncurated, unconsented data scraping introduces dangerous and illegal material directly into the AI development pipeline. Together, these December controversies show that the industry desperately needs strict data sourcing regulations before these unchecked practices cause even more harm.