Legal Experts Question OpenAI Over Suspected Game Content in Sora Training Data

OpenAI's new Sora video generator shows an uncanny ability to replicate video game footage and Twitch streams, raising legal concerns about its training data. Experts warn that using copyrighted gameplay without permission could land the company in hot water.

OpenAI launches Sora, a new AI model that generates up to 20-second videos from text prompts or images, but the company keeps the exact training data a mystery. Tests reveal that the tool easily produces footage resembling popular games like Super Mario Bros, Call of Duty, and classic arcade fighters. While OpenAI installs filters to block direct trademarked names, clever prompting bypasses these guardrails and exposes a deep familiarity with gaming aesthetics.

The AI model also demonstrates a striking ability to recreate the distinct layout of Twitch streams, complete with chat boxes and familiar interfaces. In a more concerning development, Sora generates recognizable likenesses of popular streamers like Auronplay and Pokimane, capturing specific details such as tattoos. These highly accurate reproductions strongly imply that the AI ingests vast amounts of copyrighted streaming content during its training process.

Legal experts warn that using this type of copyrighted game and streaming content without explicit permission poses significant legal risks for OpenAI. The company historically remains vague about its data sources, only admitting to using "publicly available" information and licensed stock media. As Sora becomes publicly available, these unresolved copyright questions continue to cast a shadow over the underlying technology powering the impressive video generator.

Read More at the original source →