Microsoft says virtually nobody was grabbing NYT articles through its chatbot
Microsoft claims its Copilot AI rarely reproduces substantial portions of copyrighted material, citing legal filings in its defense against publisher lawsuits.
Intelligence analysis by Gemini 2.5 Flash Lite

In a legal battle with publishers like The New York Times, Microsoft argues that its AI models, such as Copilot, do not infringe on copyright by regurgitating content. The company presented data from over 8 million chat logs, asserting that fewer than 1% contained at least 16 words in common with news articles used for training, and even fewer had longer overlaps.
Imagine a super-smart robot learning to talk by reading millions of books and articles. Microsoft says its robot, Copilot, is so good at learning that it doesn't just copy sentences. It learned so well that it rarely repeats more than a few words from what it read, like a student who understands a story instead of just memorizing it.
Analysis
Microsoft's Defense Strategy
Microsoft is actively defending itself and OpenAI against copyright infringement lawsuits brought by news publishers and authors. The core of their defense, as detailed in recent legal filings, is that the training of large language models (LLMs) on copyrighted works constitutes "fair use." This legal doctrine allows for the limited use of copyrighted material without permission for purposes such as criticism, comment, news reporting, teaching, scholarship, or research. Microsoft's argument is that the purpose of training AI models is fundamentally different from the original purpose of the copyrighted content. The AI is trained to understand and generate language, not to replace the original works as a source of information or entertainment. By demonstrating that instances of substantial content reproduction are exceedingly rare, Microsoft aims to prove that its AI systems do not substitute for the original works, thereby strengthening its fair use claim.
Analysis of Chat Logs
The company's legal strategy relies heavily on statistical analysis of Copilot's output. Microsoft provided an expert analysis of over 8.2 million chat logs, specifically selected for their potential to contain copyrighted material from news plaintiffs. The findings indicate that a minuscule fraction of these logs exhibited significant overlap with the training data. For instance, fewer than 1% of the logs contained at least 16 words in common with news content, and even fewer showed longer matches. In the context of a lawsuit from book authors, an expert found only 24 responses with at least 30 matching words out of the 8.2 million conversations analyzed. Furthermore, out of 212 books evaluated, only a small number had any matches at all. Microsoft presents these figures to underscore the low probability of its AI systems directly "regurgitating" copyrighted material in a way that would harm the market for the original works.
Implications for Fair Use and AI Development
The outcome of this legal battle could have profound implications for the future of AI development and the publishing industry. If Microsoft's fair use argument prevails, it could pave the way for AI companies to train their models on vast datasets of copyrighted material with less fear of legal repercussions. This could accelerate AI innovation by reducing the barriers to data acquisition. Conversely, if the publishers and authors succeed, it could lead to stricter regulations on AI training data, potentially requiring licensing agreements or compensation for copyright holders. This would likely increase the cost of AI development and might slow down the pace of innovation. The consolidation of the publishers' and authors' claims under a single judge suggests a move towards streamlining this complex legal process, with a summary judgment potentially resolving the case early.
Key points
- Microsoft argues that its AI models like Copilot rarely reproduce substantial copyrighted content.
- The company presented data from 8.2 million chat logs showing minimal overlap with news articles and books.
- Microsoft's defense hinges on the 'fair use' doctrine for training AI models.
- The outcome of the lawsuit could significantly impact AI development and copyright law.
- Publishers and authors claim AI products built on their work now compete with them.
If Microsoft's defense is successful, it could foster continued innovation in AI development by clarifying that training on copyrighted material for transformative purposes is permissible. This could lead to more advanced AI tools becoming available faster, benefiting various sectors through improved efficiency and new capabilities.
Conversely, if the courts rule against Microsoft, it could significantly hinder AI development by imposing strict licensing requirements or outright bans on using copyrighted data for training. This might lead to increased legal costs, slower innovation, and a potential concentration of AI power among entities that can afford extensive licensing.



