Seattle Times and Newsday Sue OpenAI and Microsoft Over Training Data

Seattle Times and Newsday sue OpenAI and Microsoft, claiming unauthorized use of their articles for AI training and reproduction of passages in model outputs.

The News

The Seattle Times and Newsday filed suit against OpenAI and Microsoft. The complaints claim the companies used the outlets' copyrighted journalism as training data without permission and that the resulting models sometimes reproduce verbatim passages from the original articles when users ask related questions.

Context

This action follows a pattern already set by earlier plaintiffs. The New York Times, Ziff Davis, Merriam-Webster, and Encyclopedia Britannica have brought comparable cases against the same defendants. The new filings name both OpenAI and Microsoft, consistent with prior complaints that treat the two companies as joint actors in the development and deployment of the models.

The suits focus on two distinct harms. First, the ingestion of full articles during training. Second, the models' ability to surface extended excerpts in response to ordinary prompts. Both organizations argue that no license was granted for either step.

Details

The Seattle Times and Newsday describe their content as original reporting that required substantial investment. They state that OpenAI and Microsoft incorporated that material into training sets for successive generations of large language models. The complaints further allege that user queries can elicit near-verbatim reproductions of sentences and paragraphs that first appeared in the two publications.

These assertions mirror claims made in the earlier suits. The New York Times case, for example, documented specific instances in which GPT models returned long passages from its archives. The Seattle Times and Newsday filings add two more regional and local voices to the same line of argument.

Microsoft is included because it has provided substantial computing resources and capital to OpenAI and because its own products now surface the models in question. The complaints treat the relationship as integrated rather than arm's-length.

No public response from either defendant appears in the available filings. The pattern of litigation suggests that OpenAI and Microsoft will likely raise fair-use defenses, though those arguments have yet to be tested at trial in this specific context.

Why it Matters

The addition of two more plaintiffs increases the legal and financial exposure facing OpenAI and Microsoft. Each new case raises the possibility of separate discovery processes, different damage calculations, and additional precedent-setting rulings on whether training on copyrighted news text qualifies as fair use. For newsrooms that have not yet sued, the filings demonstrate that smaller and mid-sized organizations can also pursue claims.

For developers and product teams that rely on the current generation of models, the suits underscore ongoing uncertainty about training data provenance. If courts ultimately limit the use of unlicensed news text, future model releases may require narrower datasets or explicit licensing agreements. The cost of such licenses would fall on the companies that build and host the models rather than on the end users who query them.

The cases also highlight a practical distinction between search and generation. Traditional search engines surface links to original articles. The complaints argue that generative models can deliver the substance of those articles without directing traffic back to the source. Whether that distinction changes the fair-use analysis remains for the courts to decide.

The Seattle Times and Newsday suits therefore do not stand alone. They extend an existing line of litigation and test whether the same legal theories will hold when applied to additional publishers and different scales of operation.

---

Sources:

No comments yet