New unredacted information in the copyright lawsuit The New York Times brought against OpenAI and Microsoft three years ago reveals that AI companies privately acknowledged their training practices amounted to theft, and that their products pose an existential threat to publishers.
The Theft Admission
Per the lawsuit filings, a top Microsoft executive described the companies' AI training practices as theft, and OpenAI's own leadership said its AI models posed an existential threat to the publishers and journalists whose work trained them.
The admissions come from newly unsealed material in the three-year-old lawsuit, in which The New York Times alleged the firms violated copyright law by training generative AI models on its content. While judges have largely been favorable to AI companies' fair use arguments, these new internal documents cut directly against that defense.
The Doom Loop
Microsoft's own data showed its Copilot answer engine caused click-through rates for The New York Times domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation described the decline as a doom loop that would hurt the performance of our models and the entire web at the same time.
The presentation stated: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain.
Microsoft CEO Satya Nadella testified in a deposition that anything that is paywalled should be licensed by anyone who wants to use it for grounding or training, and that had he been made aware OpenAI had scraped paywalled content, he would have invoked Microsoft's right to require OpenAI to retrain its models.
Existential Threat to Publishers
OpenAI's head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an existential threat from products like the chatbot, which are largely substitutive and will get more and more substitutive as they get better.
OpenAI President Greg Brockman described the models as excellent at news. Nadella agreed under oath that conversing with chatbots has substituted giving you the information right there on the website on the AI platform versus needing to go to the underlying source.
The Scale of Copying
The documents reveal for the first time that OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.
The filing details how the companies assembled Project Mango, a training dataset containing copies of at least 160,903 unique works from the news publishers. OpenAI delivered its entire GPT-3 training dataset to Microsoft, which used it to evaluate how to implement OpenAI's models. Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango.
Paywall Bypass and Copyright Stripping
The filings show that when OpenAI researcher Nick Ryder told Brockman about a hack to get around the NYT paywall, Brockman replied: ah nice.
OpenAI employees allegedly built training datasets like WebText and WebText2 that disproportionately relied on scraped news content. They also allegedly pulled millions of articles from Common Crawl, a free, open repository of web crawl data.
Perhaps most damning, the findings describe deliberate efforts to strip copyright notices from training data before it reached the model, since researchers would not want model outputting copyright notices to users.
A Real Risk to Employment
A Microsoft document states that there is a real risk that generative AI could significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.
The question of whether AI firms can legally use copyrighted material to train AI has no clear answer. Earlier this month, the Trump administration contributed a brief in defense of OpenAI's use of copyrighted material. But the new admissions, particularly around market substitution and internal awareness of harm, could prove pivotal.
OpenAI and Microsoft did not return requests for comment.




