Skip to content
KoishiAI
ไทย
← Back to all articles
ai-news ai-safety openai microsoft copyright-lawsuit ai-benchmarks

OpenAI and Microsoft Knew Their AI Training Was a 'Doom Loop'—Here's What It Cost

Unsealed court documents from The New York Times' copyright lawsuit reveal that OpenAI and Microsoft executives acknowledged their data scraping practices created a 'doom loop' threatening the web and amounted to 'the largest theft of labor in human history,' despite public claims of innovation and safety.

AI-drafted from cited sources, fact-checked and reviewed by a human editor. How we work · Standards · Report an error
Close-up of a blue screen error shown on a data center control terminal.
Photo by panumas nikhomkhai on Pexels

TL;DR: Unsealed court documents reveal that OpenAI and Microsoft executives acknowledged their data scraping created a ‘doom loop’ threatening the web, with one executive calling it the largest theft of labor in human history. This admission exposes a contradiction between public safety claims and internal awareness of predatory practices driven by competitive panic against Google.

Key facts

  • Unsealed documents from The New York Times’ copyright lawsuit reveal OpenAI and Microsoft executives acknowledged their data scraping created a ‘doom loop’ threatening web ecosystems.
  • Microsoft Director of Applied Science Brent Hecht described unlicensed content scraping as ‘the largest theft of labor in human history’ in internal filings signed by leadership.
  • Internal Microsoft documents warned that generative AI products could kill ‘the entire web,’ contradicting public narratives about safety and innovation.
  • In a June 12, 2019 email, CTO Kevin Scott told CEO Satya Nadella and Bill Gates Microsoft was ‘multiple years behind’ Google in machine learning scale.
  • Microsoft invested $1 billion in OpenAI to close the competitive gap with Google, which Nadella acknowledged as its ‘biggest AI competitor’ at the time.
  • In January 2024, the FTC sent investigative letters to Alphabet, Amazon, Anthropic, Microsoft, and OpenAI regarding competition in the generative AI space.

The Unsealed Truth Behind AI’s Web-Wrecking Cycle

Internal documents from The New York Times’ copyright lawsuit against OpenAI and Microsoft have revealed that executives at both companies were aware their data scraping practices posed a systemic threat to online content ecosystems [1][5]. These filings, submitted as part of a request for summary judgment, describe the training of large language models (LLMs) as creating a ‘doom loop’—a self-reinforcing cycle where AI models degrade the quality and value of web content, which in turn harms the data used to train future models [1][5].

The term ‘doom loop’ appears in multiple internal Microsoft documents, with one describing generative AI products as capable of killing ‘the entire web’ [1][5]. This acknowledgment comes despite public narratives from both companies promoting their AI technologies as transformative and beneficial.

The Theft That Was Known

Among the most striking revelations is a characterization by Brent Hecht, Microsoft’s Director of Applied Science, who described the unlicensed scraping of online content as ‘the largest theft of labor in human history’ [1][5]. Another internal document referred to the use of unlicensed material as an ‘astonishing theft of unprecedented proportions,’ noting that creators were not compensated for their work.

While Microsoft has since distanced itself from Hecht’s remarks—claiming they reflect an individual academic perspective rather than corporate policy—these statements were made within official filings and were signed off by multiple executives, including OpenAI leadership [1]. The fact that such language appears in legal submissions underscores the seriousness of the claims.

A Competitive Panic That Accelerated Risk-Taking

The origins of this risky strategy trace back to 2019. Internal Microsoft documents show that CTO Kevin Scott warned CEO Satya Nadella and co-founder Bill Gates that Microsoft was ‘multiple years behind’ Google in machine learning scale [2][3][4]. In a June 12, 2019 email, Scott noted it took six months to replicate Google’s BERT language model due to infrastructure limitations.

This competitive pressure prompted Microsoft to invest $1 billion in OpenAI—a move aimed at closing the gap with Google, which Nadella later acknowledged as its ‘biggest AI competitor’ before the partnership [7]. The investment was framed as a long-term bet on stable collaboration, aligning with Microsoft’s ethos as a platform partner.

Regulatory Scrutiny and Corporate Contradictions

The revelations come amid growing regulatory scrutiny. In January 2024, the Federal Trade Commission (FTC) sent investigative letters to Alphabet, Amazon, Anthropic, Microsoft, and OpenAI, probing how their investments in AI startups affect competition in the generative AI space [6]. While some internal emails about Microsoft’s fears of Google were released during US Justice Department antitrust proceedings against Google, the specific documents detailing the ‘doom loop’ and ‘theft’ originated from The New York Times’ lawsuit, not those cases [1][5].

Microsoft has attempted to downplay Hecht’s statements, with spokesperson Alex Haurek stating they reflect an individual viewpoint and do not represent company policy [1]. Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, described Hecht as having ‘divergent, academic’ views and emphasized he does not speak for Microsoft on the broader implications of AI [1]. However, these disclaimers do not negate the fact that such language was included in court filings submitted by Microsoft’s legal team.

The Cost of Innovation Built on Unlicensed Data

The contradiction between public-facing messaging and internal admissions has intensified criticism. While companies tout AI as a force for innovation, the unsealed documents suggest the technology was built on practices many now label as predatory—using vast amounts of unlicensed human-generated content without compensation [1][5].

Critics argue these internal acknowledgments prove that the industry’s narrative of responsible development is at odds with its actual practices. As AI models increasingly replace human-authored content, the long-term sustainability of online ecosystems remains in question.

The legal battle continues, but one thing is clear: OpenAI and Microsoft executives were not blind to the risks they were creating. The question now is whether accountability will follow—or if the ‘doom loop’ will continue unchecked.

Sources

  1. OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web (www.theverge.com) — 2026-09-18
  2. ‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft (www.404media.co) — 2026-09-17
  3. Microsoft’s OpenAI investment was triggered by Google fears, emails reveal (theverge.com) — 2024-05-01
  4. Microsoft’s OpenAI investment was triggered by Google fears, emails reveal (www.theverge.com) — 2024-05-01
  5. Microsoft’s OpenAI investment was triggered by Google fears, emails reveal (www.theverge.com) — 2024-05-01
  6. Nadella tells us that before the OpenAI partnership, Google was its biggest AI competitor. (www.theverge.com) — 2026-05-11
  7. FTC investigating Microsoft, Amazon, and Google investments into OpenAI and Anthropic (www.theverge.com) — 2024-01-25

Frequently asked questions

What is the 'doom loop' that OpenAI and Microsoft acknowledged in their internal documents?
The term 'doom loop' refers to a self-reinforcing cycle where AI models degrade the quality and value of web content, which subsequently harms the data used to train future models. Internal Microsoft documents described this process as capable of killing 'the entire web.'
Did Microsoft executives admit that AI data scraping was a form of theft?
Brent Hecht, Microsoft’s Director of Applied Science, characterized the unlicensed scraping of online content as 'the largest theft of labor in human history.' Another document referred to this practice as an 'astonishing theft of unprecedented proportions' because creators were not compensated for their work.
Why did Microsoft invest so heavily in OpenAI according to internal emails?
Microsoft invested $1 billion in OpenAI in 2019 to close the gap with Google, which CTO Kevin Scott warned was 'multiple years behind' in machine learning scale. This competitive pressure drove Microsoft's risky strategy regarding data acquisition and AI development.
Does Microsoft deny the 'theft' comments made by its director?
Microsoft spokespersons claim that Brent Hecht’s remarks reflect an individual academic perspective rather than corporate policy. However, these statements were included in official legal filings signed off by multiple executives from both Microsoft and OpenAI.
Where did these unsealed court documents revealing AI risks come from?
The documents originated from The New York Times’ copyright lawsuit against OpenAI and Microsoft, specifically during a request for summary judgment. While some related emails were released in US Justice Department antitrust proceedings against Google, the specific 'doom loop' details came from the NYT case.