OpenAI Dismisses New York Times Copyright Lawsuit as Baseless
OpenAI publicly defends its AI training practices against a recent copyright lawsuit from The New York Times, arguing that using publicly available web data constitutes fair use.
OpenAI officially responds to the December copyright lawsuit filed by The New York Times, declaring the legal action entirely without merit. The company maintains that training its generative AI models, such as GPT-4 and DALL-E 3, on publicly available internet data falls under the legal doctrine of fair use. OpenAI insists that this approach allows creators to benefit, supports innovation, and keeps the United States competitive in the global tech landscape without requiring the company to pay for the training examples.
The AI developer also addresses the issue of regurgitation, which occurs when a model outputs its training data verbatim. OpenAI argues that this phenomenon is rare with single-source datasets and suggests that The New York Times intentionally manipulates prompts to force the AI to spit out copied text. Furthermore, the company claims that the examples cited in the lawsuit come from older articles that already exist across multiple third-party websites, implying the newspaper cherry-picked these results from numerous attempts.
This defense arrives as the broader debate over generative AI and copyright reaches a critical peak in the tech industry. Noted AI critic Gary Marcus and visual effects artist Reid Southen recently publish findings that contradict OpenAI's stance, demonstrating that AI systems sometimes regurgitate data even without specific prompting. Their research directly references the ongoing New York Times litigation and challenges the credibility of OpenAI's claims regarding how frequently and easily its models reproduce copyrighted material.