OpenAI Data Deletion Disrupts New York Times Copyright Lawsuit

Lawyers for The New York Times and Daily News reveal that OpenAI engineers accidentally erased crucial search data used to find copyrighted material in AI training sets. OpenAI disputes the claim and blames the publishers for a system misconfiguration.

Lawyers for The New York Times and Daily News report that OpenAI engineers accidentally delete search data relevant to their ongoing copyright lawsuit. The publishers spend over 150 hours searching OpenAI's training data on provided virtual machines to find their copyrighted articles, but an November 14 erasure wipes out a week's worth of this intensive work.

OpenAI attempts to recover the lost files and largely succeeds in retrieving the raw data. However, the company permanently loses the original folder structures and file names, making the recovered information completely useless for proving where the publishers' articles appear in OpenAI's AI models.

The publishers emphasize they do not suspect intentional destruction of evidence, but they note the incident highlights that OpenAI is in the best position to search its own datasets. OpenAI firmly denies deleting any evidence and shifts the blame to the publishers, claiming their requested configuration change causes the technical issue.

Read More at the original source →