35 US Newspaper Publishers Sue OpenAI, Microsoft Over Alleged Copyright Infringement
Thirty-five US newspaper publishers have filed a copyright lawsuit against OpenAI and Microsoft, alleging unauthorized use of their content to train AI models like ChatGPT and Copilot.
Intelligence analysis by Gemini 2.5 Flash

The lawsuit, filed in New York, claims that OpenAI and Microsoft scraped articles, including paywalled content, from nearly 400 news outlets, removed copyright information, and incorporated them into training datasets without permission or compensation, leading to alleged memorization and reproduction of copyrighted material.
Imagine you write amazing stories for your school newspaper. Now, a super-smart robot comes along, reads all your stories (even the ones behind a special club password), copies them, and uses them to learn how to write its own stories. But it forgets to say where it got the ideas, like taking your name off your articles. The newspaper writers are now saying, 'Hey, you used our hard work to get super rich, but you didn't ask or pay us!'
Analysis
The Publishers' Core Allegations
Thirty-five local and regional newspaper publishers, representing nearly 400 outlets across 33 US states, have launched a significant copyright infringement lawsuit against OpenAI and Microsoft. The core of their complaint centers on the alleged systematic scraping of their journalistic content, including material behind paywalls, to train large language models such as ChatGPT and Microsoft Copilot. The publishers claim that automated crawlers, utilizing tools like Dragnet and Newspaper, not only copied articles to the defendants' servers but also deliberately stripped away crucial copyright management information (CMI), such as author bylines, publication names, and terms of use, before incorporating the text into training datasets like WebText and Common Crawl. This alleged removal of CMI is a critical component of their Digital Millennium Copyright Act (DMCA) claim, suggesting an intentional effort to obscure the origin of the content and make infringement harder to detect.
The scale of the alleged infringement is highlighted by token counts presented in the complaint. Analyses of open-source dataset approximations, like OpenWebText and C4 (a filtered snapshot of Common Crawl used for GPT-3), reveal millions of tokens sourced from the plaintiffs' websites. For instance, Ogden Newspapers alone accounted for over 71 million tokens in C4, demonstrating the extensive use of their material. The publishers argue that these AI models have not only been trained on their content but have also 'memorized' portions, reproducing them in response to user prompts, thereby directly infringing on their registered copyrights and undermining their business models.
Legal Battlefronts: Copyright and DMCA
The lawsuit asserts three distinct legal counts. The first, direct copyright infringement, is brought by five publishers with registered copyrights, alleging unauthorized reproduction, storage, and distribution of their works during model training and through model outputs. This claim targets the fundamental act of using copyrighted material without permission. The second count, vicarious copyright infringement, extends liability to Microsoft and OpenAI's parent entities, arguing that they controlled and profited from the infringement carried out by their subsidiaries and partners, possessing the ability to prevent such actions.
Perhaps the most expansive claim, and one that includes all 35 plaintiffs regardless of copyright registration, is the DMCA violation under 17 U.S.C. § 1202. This count specifically alleges the knowing removal of CMI with the intent to conceal infringement. The publishers contend that the extraction tools were chosen precisely because they were known to strip away identifying information, which, if retained, would have linked the content to its original source and alerted users to its copyrighted status. This DMCA claim is significant because it provides a legal avenue for all plaintiffs to seek damages, even those without formally registered copyrights, broadening the scope and potential impact of the lawsuit.
AI's Growth vs. Content Compensation
The complaint places these allegations within the context of the defendants' immense financial success. OpenAI, valued at $852 billion after a $122 billion funding round in March 2026 and reportedly generating $2 billion in monthly revenue, confidentially filed for an IPO in June 2026, with projections exceeding a $1 trillion valuation. Microsoft, a key partner, also reported substantial quarterly revenue. The publishers' central grievance is that none of this vast revenue, generated from AI models that reportedly serve over 900 million weekly active ChatGPT users and are utilized by more than 92% of Fortune 500 companies, has been shared with the content creators whose work allegedly formed the foundational training data. This lawsuit is not merely about past infringement but about establishing a framework for fair compensation and ethical data sourcing in an AI-driven economy. It highlights the growing tension between the rapid advancement and commercialization of AI technologies and the rights of content creators whose intellectual property fuels these innovations, potentially setting a global precedent for how AI companies engage with copyrighted material.
Key points
- 35 US newspaper publishers are suing OpenAI and Microsoft for alleged copyright infringement.
- The lawsuit claims AI models like ChatGPT and Copilot were trained using scraped articles, including paywalled content, without permission or compensation.
- Publishers allege that copyright management information (CMI) was deliberately removed from their content before it was used for training.
- The complaint includes claims for direct copyright infringement, vicarious copyright infringement, and a DMCA violation for CMI removal.
- The lawsuit highlights the immense financial growth of OpenAI and Microsoft, contrasting it with the lack of compensation for content creators.
This lawsuit could lead to clearer guidelines and fair compensation models for content creators whose work is used to train AI, fostering a more equitable digital ecosystem. It might also encourage AI developers to prioritize ethical data sourcing and transparency, potentially leading to more robust and trustworthy AI systems.
The legal battle could be prolonged and costly, potentially stifling AI innovation due to increased regulatory uncertainty or high licensing fees. Publishers might also face an uphill battle against well-resourced tech giants, with an unfavorable outcome potentially diminishing the value of journalistic content in the AI era.



