The New York Times and The Daily News allege that OpenAI withheld evidence in their multi-year copyright lawsuit, which accuses the company of training its generative models on the outlets’ journalism and reproducing that content in ChatGPT outputs. The dispute has been ongoing for over two years.
Throughout the litigation, OpenAI has maintained that it could not feasibly search its training corpus, and that retrieving, processing and de-identifying the vast collection of ChatGPT conversation logs would be technically burdensome and raise user-privacy concerns.
What the deposition allegedly revealed
According to the newspapers, in an April court-ordered deposition OpenAI data privacy engineer Vinnie Monaco disclosed that the company had run internal searches and evaluations against its training corpus to look for copyrighted journalism.
Monaco’s testimony also reportedly revealed that, before The New York Times filed its suit, OpenAI had already compiled an internal database of about 78 million de-identified ChatGPT conversations to assess how extensively the model might be infringing on others’ works. The deposition allegedly further disclosed that OpenAI deployed a so-called “Bloom” filter as part of a toolset called “Project Giraffe” that detected and logged instances of regurgitation in model outputs shortly after the lawsuit was filed.
Disputes over discovery materials
The plaintiffs originally sought a sample of 120 million chat logs; through negotiations the request was reduced to 20 million. OpenAI submitted that 20-million-log sample to the court last December, but the plaintiffs say the company redacted so much of the material that, in the court’s words, the sample was rendered “unusable.” The plaintiffs also claim OpenAI deleted billions of ChatGPT outputs after the suit was filed in violation of the court’s preservation order and swapped out millions of logs in the requested sample.
Plaintiffs say these actions made obtaining evidence unduly difficult even though the company had allegedly already collected relevant internal data.
What plaintiffs are asking the court to do
The New York Times and The Daily News have asked the judge to discipline OpenAI for allegedly withholding evidence and obstructing discovery. Their requests include:
- barring OpenAI from using the 20 million chat-log sample as evidence on the grounds that it is unreliable;
- directing the court to accept as fact that ChatGPT logs would have shown substantial regurgitation and grounding in the plaintiffs’ content;
- preventing OpenAI from arguing that the logs it produced do not demonstrate substantial regurgitation; and
- requiring OpenAI to pay plaintiffs’ legal fees incurred while seeking the allegedly withheld evidence.
Ian B. Crosby, lead counsel for the plaintiffs, said in a statement: “If OpenAI genuinely believed that copying our clients’ journalism was fair and legal, it wouldn’t have hid the truth about having done it.”
OpenAI’s response
OpenAI spokesperson Drew Pusateri denied the allegations and accused The New York Times of seeking access to private user conversations as its case weakens. In a statement cited by the outlets, Pusateri said, “As the Times’ case weakens and they’ve been forced to drop claims against us, they’re persisting with their efforts to invade the privacy of people who have nothing to do with this case, including by making these blatantly false allegations. We’ll continue defending our users’ privacy and the long-established principles of fair use.”
Why it matters
The dispute touches on transparency about the sources used to train large language models, enforcement of copyright protections against AI-generated outputs, and the handling of user data during litigation. If the plaintiffs’ allegations are borne out, they could affect future discovery rules, data-retention practices, and legal accountability for AI training data. Further developments in the case will reveal whether the court finds that OpenAI improperly withheld or altered evidence.



