TechCrunch reports that newly unredacted filings in The New York Times’ copyright lawsuit against OpenAI and Microsoft reveal a senior Microsoft executive describing AI training data scraping as an "astonishing theft" and "the largest theft of labor in human history." The filings also quote OpenAI leaders calling their own models an "existential threat" to publishers and detail alleged paywall circumvention and mass copying of news content.
This article aggregates reporting from 2 news sources. The TL;DR is AI-generated from original reporting. Race to AGI's analysis provides editorial context on implications for AGI development.
The unsealed language in the Times v. OpenAI/Microsoft case is one of the starkest windows yet into how frontier AI companies talked about data internally. When a senior Microsoft executive calls web scraping for AI training an "astonishing theft" and the "largest theft of labor in human history," it undercuts the industry’s public reliance on fair use arguments and signals real legal exposure around how foundational models were built.([techcrunch.com](https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/?utm_source=openai))
For the race to AGI, this matters because every cutting‑edge model is downstream of massive, messy datasets assembled under legal gray zones. If courts deem large swaths of that data unlawful to use, companies may need to rebuild pipelines around licensed or synthetic data, raising costs and potentially slowing the cadence of new frontier models. On the other hand, big incumbents like Microsoft and OpenAI are best positioned to strike multi‑billion‑dollar licensing deals and retrofit their stacks, which could entrench their lead over smaller labs.
Strategically, this case crystallizes the emerging bargain: AGI‑class systems may be technically feasible this decade, but only if the industry can secure a sustainable, legally durable data supply. Expect more publishers to demand equity‑like economics or revenue sharing rather than flat fees, and expect regulators to probe whether data partnerships become another vector for platform dominance.

