Explosive Filings: Microsoft Exec Called AI Scraping 'Largest Theft of Labor'
Explosive Filings: Microsoft Exec Called AI Scraping 'Largest Theft of Labor'
Unsealed court documents reveal Microsoft and OpenAI privately viewed AI data scraping as 'theft' and an 'existential threat' to publishers. Discover the shocking internal admissions in the NYT lawsuit.
Newly unsealed court documents from The New York Times' copyright lawsuit against OpenAI and Microsoft have revealed astonishing internal admissions, painting a picture of deep concern within both tech giants regarding their AI training practices. A top Microsoft executive privately described AI data scraping as nothing less than “the largest theft of labor in human history.” Adding to the gravity, OpenAI's own leadership acknowledged that their AI models posed an “existential threat” to the publishers and journalists whose work fueled these very systems.
The filings detail how these companies allegedly obtained copyrighted content by bypassing paywalls, building massive training datasets through mass scraping, and even stripping copyright notices from the data. While much of this new information comes from The Times' brief, not direct exhibits, the implications are profound.
This latest escalation in the three-year-old lawsuit challenges the common “fair use” defense AI companies often employ.
Fair use typically allows copyrighted work to be used without permission for purposes like parody or news reporting, provided it doesn't harm the market for the original work. However, these new revelations directly contradict that defense.
For instance, internal Microsoft data showed its AI-powered Copilot “answer engine” caused click-through rates for The New York Times’ domain to plummet by as much as 93% compared to traditional Bing searches. A Microsoft internal presentation from January 2024 by director Brent Hecht ominously described this decline as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.” The document further noted, “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’
Even Microsoft CEO Satya Nadella testified under oath earlier this year that any paywalled content should be licensed for AI training. He stated he would have “invoked [Microsoft's right to] require OpenAI to retrain its models” if he had known OpenAI scraped paywalled information.
OpenAI's head of ChatGPT, Nick Turley, also warned internally that publishers face an “existential threat” from chatbots, calling them “largely substitutive” and noting they “will get more and more substitutive as they get better.” OpenAI President Greg Brockman described their models as “excellent at news,” further highlighting the direct competition.
Nadella himself agreed that conversing with chatbots substituted the need to visit original sources. Another Microsoft document pointed out a “real risk” that generative AI could
significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.
The sheer scale of the operation is staggering, with documents revealing OpenAI's mid-training datasets alone containing over 91,692 copies of published works.
These revelations could significantly impact the future of AI development and copyright law.