Key facts
- Microsoft executive Brent Hecht called AI scraping "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."
- OpenAI's Head of ChatGPT, Nick Turley, stated publishers face an "existential threat" from AI models.
- Microsoft's Copilot "answer engine" caused a 93% drop in click-through rates for The New York Times' domain.
- Microsoft CEO Satya Nadella stated paywalled content should be licensed for AI training.
- OpenAI's mid-training datasets contained over 91,692 copies of works from The New York Times, Daily News, and Center for Investigative Reporting.
- OpenAI employees allegedly devised a plan to circumvent paywalls without detection.
New unredacted information has emerged in the three-year-old copyright lawsuit filed by The New York Times against OpenAI and Microsoft, revealing internal admissions that AI training practices constitute "theft" and pose a significant threat to publishers.
A top Microsoft executive privately described the companies’ AI training practices as "theft," and OpenAI’s own leadership acknowledged that its AI models presented an “existential threat” to the publishers whose work was used for training.
The unsealed material details how the companies allegedly bypassed paywalls undetected, conducted mass scraping to build training datasets, and deliberately removed copyright notices from the data. Much of this new information comes from The Times’ own brief, with underlying exhibits remaining sealed.
While the legality of using copyrighted material for AI training is still debated, judges have often favored AI companies' "fair use" arguments. However, several new admissions in the filing appear to contradict OpenAI's fair use defense, particularly regarding the requirement that the use does not substitute for or harm the market for the original work.
Microsoft's internal data indicated that its Copilot "answer engine" led to a drop of up to 93% in click-through rates for The New York Times' domain compared to traditional Bing searches. An internal Microsoft presentation by Director of Applied Science Brent Hecht referred to this decline as a "doom loop" that would negatively impact both AI models and the web.
Microsoft CEO Satya Nadella testified that any paywalled content should be licensed for AI training and stated he would have required OpenAI to retrain its models if he had known they scraped paywalled information.
OpenAI's Head of ChatGPT, Nick Turley, wrote in internal communications that publishers face an "existential threat" from AI chatbots, which are "largely substitutive" and will improve over time. OpenAI President Greg Brockman described the models as "excellent at news," and Nadella agreed that chatbots substitute for visiting original sources.
A Microsoft document highlighted a "real risk" that generative AI could disrupt the employment of individuals who created the data used to train foundation models. The scale of copying is substantial, with OpenAI's mid-training datasets alone containing over 91,692 copies of works from The New York Times, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com.
Hecht, in a January 2023 memo, characterized the situation as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." The filing also detailed how OpenAI and Microsoft acquired content, including scraping from the Bing Index. OpenAI provided its GPT-3 training dataset to Microsoft, which Microsoft used for its commercial products, while Microsoft also supplied training data to OpenAI through initiatives like Project Taxi and Project Mango. OpenAI employees allegedly developed a plan to circumvent paywalls without detection, with Brockman responding positively to a "hack to get around nytimes paywall."
