Newly unredacted material in The New York Times’ three‑year‑old copyright lawsuit against OpenAI and Microsoft contains internal statements from the companies acknowledging that mass web scraping for AI training was tantamount to theft and posed a serious threat to news publishers.
According to the filings, a senior Microsoft executive privately referred to the firms’ AI‑training practices as “theft,” and OpenAI leadership described their models as representing an “existential threat” to the publishers and journalists whose work served as training data.
What the newly disclosed material reveals
The filings describe how the companies allegedly acquired and used publishers’ content by circumventing paywalls, scraping large amounts of material (including from the Bing index and Common Crawl), and removing copyright notices from training datasets to prevent models from outputting those notices to users. Much of the newly public information originates from The New York Times’ own legal brief; many underlying exhibits remain sealed, so some quotes appear without full context.
Specific internal admissions and metrics
-
Microsoft internal data cited in the filings showed that its Copilot “answer engine” reduced click‑through rates to The New York Times’ domain by as much as 93% compared with standard Bing search. In a January 2024 internal presentation, Brent Hecht, Microsoft’s Director of Applied Science, described that decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”
-
A Microsoft document quoted in the filing states: “It is highly unusual that an end‑product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’ ”
-
Microsoft CEO Satya Nadella testified earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and said that had he known OpenAI scraped paywalled material, he would have required OpenAI to retrain its models.
-
OpenAI’s Head of ChatGPT, Nick Turley, wrote internally that publishers face an “existential threat” from chatbot‑style products that are “largely substitutive” and will become more so as they improve.
-
OpenAI President Greg Brockman described the models as “excellent at news.” Nadella testified that conversing with chatbots “has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”
-
A Microsoft document warned of a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”
Scale of copying and composition of training datasets
The filings present figures on the scope of content copying: OpenAI’s mid‑training datasets reportedly include more than 91,692 copies of works published by The New York Times, Daily News, and the Center for Investigative Reporting. A Common Crawl‑derived dataset reportedly contained over 2 million documents from nytimes.com alone.
In a January 2023 internal memo, Brent Hecht called the practice “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.”
The filing also alleges that OpenAI transferred its entire GPT‑3 training dataset to Microsoft, which Microsoft used to evaluate integrating OpenAI’s models into its commercial products. It says Microsoft likewise provided data to OpenAI via initiatives labeled Project Taxi and Project Mango. The companies allegedly assembled Project Mango into a training dataset containing copies of at least 160,903 unique works from the news publishers.
Paywall circumvention and removal of copyright notices
The filings say OpenAI employees developed methods to bypass paywalls without detection. When OpenAI researcher Nick Ryder told Greg Brockman about a “hack to get around nytimes paywall,” Brockman replied: “ah nice,” according to the filing.
The documents also describe the construction of WebText and WebText2 datasets that relied heavily on scraped news content and the ingestion of millions of articles from Common Crawl. The filings further allege deliberate efforts to strip copyright notices from training material so that models would not reproduce those notices in outputs.
Legal context and responses
These unsealed materials are the latest escalation in the three‑year legal dispute that began when The New York Times sued OpenAI and Microsoft, alleging unlawful use of its copyrighted content to train generative AI models. Whether training on copyrighted material constitutes lawful “fair use” remains unsettled in absolute terms; courts have so far often been receptive to AI firms’ fair‑use arguments. Earlier this month, the Trump administration filed an amicus brief defending OpenAI’s unlicensed use of copyrighted material for LLM training.
However, several of the companies’ own admissions in the filings—particularly about substitution and market harm—could undermine a fair‑use defense that requires a use not to supplant or harm the market for the original works.
OpenAI and Microsoft did not respond to requests for comment.



