OpenAI and Microsoft Concede LLMs Create 'Doom Loop' by Consuming the Web They Need

OpenAI and Microsoft have stated that large language models both depend on and erode the open web through large-scale extraction of content without consent.

The news

OpenAI and Microsoft have described the current state of large language models as a "doom loop." The companies acknowledge that these systems are built on the mass collection of existing web content and that this same process is degrading the quality and availability of that content for future use. A central quote from their position notes that millions of people will soon view the hoovering of their work by large models as an astonishing theft of unprecedented proportions.

Context

Prior to this admission, public discussion of training data sources often centered on scale and capability rather than ownership or long-term sustainability. The companies now frame the relationship between models and the web as self-undermining: models require fresh, high-quality human-generated material to improve, yet their outputs and data practices reduce incentives for people to publish that material openly. This shifts the prior assumption that the web would remain an abundant, freely accessible resource for model training.

The shift appears in how the two companies present the dependency. Earlier framing treated the open web as raw material that could be drawn upon indefinitely. The new description instead treats continued large-scale extraction as a process that undercuts its own foundation. No other parties are named in the statement as sharing this view, and the companies do not outline any alternative data sources they intend to pursue.

Details

The admission centers on two linked claims. First, the models were trained by ingesting vast quantities of existing web text without explicit permission from creators. Second, the resulting systems accelerate the decline of original web content because users and publishers see their work extracted at scale. The quoted assessment ties these issues directly to perceptions of theft rather than technical necessity or fair use arguments. No specific volumes of data, lists of affected sites, or timelines for mitigation appear in the statement.

The language used is explicit about scale and consent. The phrase "hoovering up all their work" is presented as the lens through which millions of people will judge the activity. The companies do not qualify the claim with references to robots.txt compliance, licensing deals, or opt-out mechanisms. They also do not dispute the characterization of the activity as theft; they simply record that the public will reach that conclusion.

Why it matters

The concession removes the usual corporate framing that treats web scraping as a neutral or inevitable step in model development. For developers and publishers who rely on search traffic or open distribution, the "doom loop" language signals that the companies themselves expect the supply of freely shared human writing to shrink. This changes the practical environment for anyone building tools or businesses on top of public web data: continued extraction now carries an explicit internal warning that the resource base is being consumed faster than it is replenished.

The result is a narrower path for future model improvement that does not rely on ever-larger unlicensed ingestion. Teams that had planned to train successive generations of models on the same open web corpus now face an acknowledged feedback loop in which the act of training reduces the quality and volume of the next corpus. Publishers who once accepted indexing for visibility must weigh whether that visibility still justifies the downstream extraction of their text. The companies' own wording supplies the evidence that this trade-off is no longer theoretical.

The admission also affects how external observers should interpret future claims about model performance. Any assertion that newer models will continue to improve at historical rates now sits alongside the companies' description of a self-limiting data supply. Readers can therefore treat statements about capability growth as conditional on either new data sources or a reversal of the incentives the companies themselves have described.

---

Sources:

{
  "sources": [
    {
      "publisher": "404 Media",
      "title": "‘Doom Loop’: OpenAI and Microsoft Admits LLMs Are Destroying the Web and Built on Theft",
      "url": "https://www.404media.co/doom-loop-openai-and-microsoft-admits-llms-are-destroying-the-web-and-built-on-theft/",
      "published_at": "2026-09-17T22:01:43.000Z",
      "summary": "\"Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions.\""
    }
  ],
  "word_count": 612
}

No comments yet