Newly unsealed filings in the New York Times copyright lawsuit reveal internal Microsoft and OpenAI documents where a senior Microsoft scientist called news scraping for AI training an "astonishing theft" and possibly the "largest theft of labor in human history." Ars Technica and TechCrunch report that the filings also show OpenAI staff discussing paywall workarounds and acknowledging that chatbot outputs could substitute for news sites.
This article aggregates reporting from 2 news sources. The TL;DR is AI-generated from original reporting. Race to AGI's analysis provides editorial context on implications for AGI development.
The unsealed language from Microsoft and OpenAI is unusually blunt. When a senior Microsoft director calls his own employer’s data practices an "astonishing theft" and warns of a "doom loop" where chatbots starve their news sources, it undermines the industry’s public narrative that training on news is obviously fair use. Legally, this gives the New York Times and other publishers powerful ammunition that even insiders thought scraping went too far. Strategically, it exposes deep anxiety that frontier AI business models depend on appropriating content from an ecosystem they might then hollow out. ([arstechnica.com](https://arstechnica.com/tech-policy/2026/09/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history/))
For the race to AGI, this dispute is not just about back payments to publishers. If courts or settlements force labs to license large swaths of premium text and multimedia, the effective cost of training frontier models could rise sharply. That would advantage labs backed by cash rich platforms, and might slow the pace of open model proliferation. On the flip side, a clear licensing regime could de risk training and make large scale data access more sustainable.
The most important takeaway is about trust. If the two most powerful Western labs privately viewed their own scraping practices as ethically dubious, regulators are more likely to treat future assurances on safety or copyright with skepticism. That, in turn, raises the odds of heavier handed AI specific regulation, particularly around data provenance and training transparency.

