The filings
Newly unsealed documents from the New York Times lawsuit against OpenAI and Microsoft contain an internal Microsoft assessment that described the companies’ data collection as “the largest theft of labor in human history.” The passages had been redacted until recently. They show both companies knowingly accessed Times articles behind paywalls, incorporated the text into training datasets, and circulated internal messages that the activity would remove revenue from news organizations. The same messages flagged the risk that models trained on this material would later generate summaries and articles that compete directly with the original sources.
What the documents describe
The unsealed sections include emails and internal notes in which Microsoft staff discussed OpenAI’s scraping of paywalled content. The material states that both organizations built datasets from the scraped articles. Microsoft executives used the phrase about labor theft in correspondence that also raised the possibility of a self-reinforcing cycle: AI systems trained on news output would produce synthetic content that reduces traffic and subscriptions to the original publishers, leaving those publishers with fewer resources to create new reporting. The filings do not list exact article counts or the precise technical steps used beyond the general description of scraping paywalled material.
Prior state of the case
Before the unsealing, only redacted versions of the filings were public. The newly visible portions add concrete internal language from Microsoft about the scale and consequences of the data practices. They also document that staff at both companies recognized the direct link between training data sources and future publisher viability. The documents reference concerns that continued scraping would accelerate a feedback loop in which synthetic output displaces original reporting and further reduces publisher income.
Internal risk discussions
The released messages show Microsoft and OpenAI personnel raising the same set of issues. One line of discussion centered on the immediate effect of removing paywalled text for training. Another line examined the longer-term outcome once models began producing news-like text at scale. The term “doom loop” appears in the filings as shorthand for the scenario in which reduced publisher revenue leads to less original reporting, which in turn leaves future models with even lower-quality or synthetic data to draw from. These exchanges occurred while both companies continued to use the scraped Times content.
Why it matters
The disclosures place Microsoft in the position of having documented, in its own words, the scale of value removal from publishers while participating in the same activity alongside OpenAI. For teams that build or fine-tune large language models, the filings supply on-the-record evidence that the companies understood the revenue threat to news sources yet proceeded. News organizations now hold clearer documentation that major AI developers anticipated the economic consequences and still chose the data-collection path. The practical effect is added weight to ongoing legal and regulatory questions about how future models obtain training text. Any organization that depends on web-scale data for model development now faces a narrower set of options that can withstand similar scrutiny, and the cost of those options will be determined in courtrooms and licensing negotiations rather than in engineering decisions alone.
---
Sources:
{"word_count": 612, "sources_used": ["TechCrunch", "Ars Technica"]}
No comments yet