Seattle Times and Newsday Add to Growing Legal Push Against OpenAI, Microsoft Over AI Training
The Seattle Times and Newsday each launched a federal lawsuit alleging that OpenAI and Microsoft harvested their articles without consent to train generative AI systems. The filings assert that the companies copied protected news content to improve large language models that now power services such as ChatGPT, infringing the papers' copyrights and cutting off potential licensing revenue.
Both complaints were filed in the U.S. District Court for the Western District of Washington and the Southern District of New York, respectively. The plaintiffs contend that the defendants employed automated scrapers to amass thousands of pieces—including pay‑walled material—and fed that text into the datasets used to train OpenAI's models, which Microsoft subsequently offers via its Azure cloud platform. The suits seek injunctive relief to stop further use of the newspapers' material, as well as monetary damages and a share of any commercial profits generated by the AI products.
The dispute hinges on a longstanding clash between the open‑internet methodology that underlies AI model building and the intellectual‑property rights of content creators. Under U.S. copyright law, reproducing a protected work without a license may constitute infringement, even when the material is later transformed by machine‑learning algorithms. Defendants have countered that the practice falls under “fair use” because it is transformative and does not replace the original articles, a defense that courts have applied inconsistently in recent years.
Seattle Times and Newsday join an expanding roster of media outlets taking the issue to court. Earlier in the year, the Associated Press, Reuters and several other publishers filed comparable actions against OpenAI, sparking vigorous debate within the publishing sector about safeguarding digital content in an AI‑driven world. Those prior cases have underscored how difficult it is to prove that particular excerpts were incorporated into training sets, given the opacity of proprietary datasets.
The lawsuits highlight a wider anxiety among journalists that AI tools can reproduce news stories verbatim or generate summaries that siphon traffic away from the original sources. Media companies argue that such loss of readership directly erodes advertising and subscription revenue, threatening the financial health of newsrooms already coping with declining income.
OpenAI and Microsoft have declined to comment on the pending cases, though both have previously maintained that their models are trained on publicly available data and that they respect copyright law. In public remarks, OpenAI has stressed its ongoing work to develop “responsible AI” practices, including exploring licensing agreements with content providers.
Legal commentators note that the rulings will likely turn on how courts interpret the fair‑use defense in the context of machine‑learning training. Some observers predict that a decision favoring the publishers could compel AI developers to negotiate licensing deals, while others warn that overly restrictive judgments might stifle innovation in the fast‑moving field.
If the courts grant the requested injunctions, OpenAI could be forced to remove the contested material from its training pipelines and possibly compensate the newspapers for past use. Conversely, a dismissal might embolden other tech firms to keep aggregating online content with little oversight. Both outcomes carry major implications for the balance between open data ecosystems and creators' rights.
The cases come at a time when policymakers are intensifying scrutiny of AI’s impact on intellectual property, privacy and misinformation. As the litigation unfolds, both the publishing industry and the technology sector are watching closely, aware that the verdict could set a precedent shaping the future relationship between news media and artificial‑intelligence developers.
Comments (0)
Be the first to comment.
Join the discussion