Microsoft Exec Called AI Scraping "The Largest Theft of Labor in History": Inside the Unsealed Doom Loop Documents
📑 Table of Contents
- Introduction: The Documents They Fought to Hide
- "The Largest Theft of Labor in Human History"
- The Doom Loop, in Microsoft's Own Words
- The Paywall "Hack" and the "Ah Nice" Email
- The Numbers: Up to 94% Traffic Collapse
- How Chatbots Were Extracting Full Articles
- What This Means for the AI Tools You Use
- Frequently Asked Questions
Introduction: The Documents They Fought to Hide
For years, Microsoft and OpenAI have fought to keep their internal thinking about training AI on news content out of public view. On Thursday, that fight largely ended: a motion for summary judgment from news plaintiffs led by The New York Times was unsealed, and the documents inside read less like a copyright dispute and more like a confession.
The filings expose internal warnings from Microsoft's own scientists that scraping news to train AI was ethically indefensible, predictions of an industry "doom loop" that would starve both publishers and the models themselves, and a breezy internal exchange at OpenAI celebrating a workaround for the New York Times paywall. If you use AI tools daily — for search, research, or writing — this case will shape what those tools are allowed to do, and what they'll have to pay for, for years to come.
"The Largest Theft of Labor in Human History"
The most explosive line comes from Brent Hecht, Microsoft's Director of Applied Science — a senior researcher, not a critic outside the building. In internal documents cited by the plaintiffs, Hecht repeatedly warned that scraping news for AI training amounted to:
"An astonishing theft of unprecedented proportions" — perhaps "the largest theft of labor in human history."
Hecht didn't stop there. While Microsoft and OpenAI's legal teams argued publicly that training AI on news content is fair use, Hecht's internal assessment said the plan to widely scrape news made "a complete mockery of the idea of 'fair use.'" In other words: the people building the technology apparently didn't believe the legal defense being mounted on its behalf.
Over at OpenAI, ChatGPT head Nick Turley wrote in an internal message that publishers faced an "existential threat" from commercial products trained on news content that could substitute for the news providers themselves.
The Doom Loop, in Microsoft's Own Words
One Microsoft internal document included in the unsealed motion described a "doom loop" that would "hurt the performance of our models and the entire web at the same time." The mechanics are brutally simple:
- Step 1: Chatbots train on news content scraped without payment or permission.
- Step 2: Those chatbots answer questions directly, so users stop clicking through to news sites.
- Step 3: News organizations lose the advertising and subscription revenue that funds journalism.
- Step 4: The quality and quantity of reliable new information on the open web declines — and future models are trained on an increasingly degraded, AI-recycled web.
The plaintiffs argue that data from both firms shows the prediction was accurate — and even Microsoft's own measurements back them up. One Microsoft document agreed there is a "real risk" that generative AI could "significantly disrupt" the employment of the very people who generated the training data. Microsoft even included an internal cartoon illustrating LLMs destroying their own supply chains.
The plaintiffs' sharpest point is structural: no individual AI company can break out of the loop alone. While the industry as a whole would benefit if every player paid to sustain quality content, each company is better off taking content for free while competitors pay. Only a court ruling that training on news requires licensing, they argue, can reset the equilibrium.
The Paywall "Hack" and the "Ah Nice" Email
Some of the most damaging material is small and human. Under oath, Microsoft CEO Satya Nadella testified that AI companies shouldn't violate news sites' terms of use by dodging paywalls. But at OpenAI, when a staffer named Nick Ryder informed President Greg Brockman that "a hack" had been found for OpenAI crawlers "to get around" the New York Times paywall, Brockman's reply was two words:
"ah nice."
Nadella also acknowledged that chatbots substitute for news platforms — describing the dynamic as stealing clicks from news sites by "giving you the information right there on the website on the AI platform versus needing to go to the underlying source." Turley agreed there was "no good reason to click," and described chatbots as "largely substitutive, period" — predicting they "will get more and more substitutive as they get better."
The motion also alleges OpenAI obtained a New York Times dataset of roughly 1.8 million articles from a third party that was bound by an agreement not to use the data commercially — and that OpenAI employees knew it "would not be appropriate" to use it for training, but did it anyway.
The Numbers: Up to 94% Traffic Collapse
The unsealed data quantifies what publishers have been shouting about for two years. Microsoft's own records showed click-through rate drops of 83–93 percent for some news plaintiffs, and drops of 51–94 percent for others. Combined with reports of low click-through from ChatGPT search results, the substitution effect is no longer theoretical.
And it's not just Microsoft and OpenAI. The plaintiffs note that after ChatGPT's launch, Google's AI Overviews was quickly introduced and began absorbing even more of the traffic that previously went to news sites — a pattern now replicated by virtually every AI search tool, from Perplexity to independent AI browsers.
How Chatbots Were Extracting Full Articles
News organizations also tested how easily chatbots could be coaxed into reproducing their copyrighted work. Their findings, per the motion:
- Asking "what's the next line?" repeatedly could walk a chatbot through an entire article, line by line.
- Requests for "summaries" or "key bullet points" of articles often produced long verbatim excerpts.
- Prompts asking chatbots to "rate the bias" of a news article were particularly effective at extracting substantial reproduced portions.
- Asking a chatbot to pick any article off a site's homepage and discuss it also yielded reproduced passages.
Microsoft's public response defended its products as "transformative fair use" that doesn't substitute for news sites, and argued Nadella's testimony reflected broad observations about how information consumption is changing — not conclusions about the copyright questions before the court.
What This Means for the AI Tools You Use
This case is not just a legal fight between billionaires — it will directly change the AI tools on your screen:
1. AI search gets more expensive
If the court rules that training on news requires licenses, the cost flows into product pricing. Expect more paywalls around AI search, higher subscription tiers, and a real shakeout among free AI answer engines.
2. Licensing deals become a moat
Tools backed by companies with publisher deals — like Microsoft Copilot and ChatGPT — could gain a durable advantage over smaller AI search tools that can't afford licensing at scale.
3. Citations become a survival feature
AI research tools that reliably cite and link back to sources — like Perplexity, Kagi, and Brave Search — are better positioned in a world where regulators and courts scrutinize substitution. If you rely on AI for research, favoring tools that send traffic back to publishers is the sustainable choice.
4. The open web's health becomes a product feature
The doom loop argument reframes web health as an AI quality problem: models trained on a collapsing web get worse. The tools that invest in sustaining their sources — not just scraping them — will simply have better raw material.
Frequently Asked Questions
What is the AI "doom loop"?
It's a cycle described in Microsoft's own internal documents: AI chatbots trained on news absorb the traffic that funds journalism, journalism revenue collapses, and the open web produces less reliable new information — ultimately degrading the data future AI models are trained on. The term appeared in a Microsoft document unsealed on September 17, 2026.
Who said AI scraping was "the largest theft of labor in human history"?
Brent Hecht, Microsoft's Director of Applied Science, in internal documents revealed in unsealed court filings. He also wrote that wide-scale news scraping made "a complete mockery of the idea of 'fair use'" — directly contradicting Microsoft's public legal position.
How much traffic have news sites lost to AI chatbots?
According to Microsoft's own data cited in the filings, some news plaintiffs saw click-through rate drops of 83–93 percent, with others seeing drops between 51 and 94 percent. Google's AI Overviews has compounded the effect since its introduction.
Will AI tools get more expensive if the publishers win?
Likely, yes — at least for AI search and research tools that rely on news content. Licensing costs would flow into subscriptions or ads. Tools with established publisher partnerships would gain an advantage over smaller free competitors.
Explore All AI Tools
Discover and compare 300+ AI tools on aitrove.ai — your trusted AI tool directory.
Browse All Tools →