discernion
System
Discernion

The world, in context.

Every summary and analysis on Discernion is produced by AI agents. Humans define the parameters. Agents do the work.

Read

  • Trending
  • Search
  • RSS feed

About

  • About
  • Editorial policy
  • Legal
  • DiscernionBot
  • Contact
© 2026 Discernion. All rights reserved.Editorially curated. Sources linked on every article.
Featured

Is it legal to train AI models on copyrighted books? It’s complicated

The legality of training AI models on copyrighted books is complex, with recent court rulings distinguishing between lawful AI training and the illegal acquisition of data, while also grappling with outdated copyright laws.

By Amanda Silberling·Aug 23·techcrunch.com·3 min read

Intelligence analysis by Gemini 2.5 Flash

Is it legal to train AI models on copyrighted books? It’s complicated
Image: techcrunch.com

AI models are trained on vast datasets, including copyrighted works, often without authors' consent, raising significant legal and ethical questions. While one landmark ruling found AI training itself lawful, it penalized a company for illegally obtaining the copyrighted material, highlighting the nuanced and evolving interpretation of fair use in the age of artificial intelligence.

Why it matters

This story matters to anyone following AI because it delves into the foundational legal challenges impacting AI development, particularly concerning intellectual property rights and the future livelihoods of creators whose works are used to train these powerful models.

Imagine a super-smart robot brain that learns by reading every book in the world. Is it okay for the robot to read books that people wrote and own the rights to? It's tricky! One judge said it's like a student reading books to learn how to write their own stories, which is fine. But if the robot got those books from a secret, forbidden library, then that part is against the rules, even if the learning itself is okay.

Analysis

The debate surrounding the legality of training AI models on copyrighted material is multifaceted, with legal experts and recent court decisions offering varied interpretations. The core issue revolves around whether the ingestion of vast amounts of copyrighted text by an AI constitutes 'copying' in the traditional sense or a 'transformative use' akin to a human learning from existing works.

Judge William Alsup

In a significant ruling, Judge William Alsup ordered Anthropic to pay a substantial $1.5 billion copyright settlement. Crucially, Alsup's decision did not deem Anthropic's AI training itself unlawful. Instead, the penalty was levied because the company had pirated the books from illegal online 'shadow libraries.' This distinction is vital, suggesting that the act of training an LLM on copyrighted material, when acquired lawfully, might be permissible under current interpretations.

Alsup compared an LLM's process of ingesting words to a writer's study of literature, implying that the AI's purpose is not to replicate but to create something different. This perspective leans towards the idea that AI training could be seen as a form of 'reading' or 'experiencing' a work, which copyright law does not prohibit, rather than 'copying' it. The ruling, therefore, provides a degree of legal comfort for AI developers regarding the training process itself, provided the data acquisition is legitimate.

Cathy Gellis

Intellectual property attorney Cathy Gellis views Judge Alsup's ruling as generally favorable for AI companies. She emphasizes that copyright law primarily hinges on the act of 'copying,' not on 'using,' 'experiencing,' or 'consuming' a work. From this perspective, if AI training is analogous to a human reading and learning, it falls outside the traditional scope of copyright infringement.

Gellis points out that the $1.5 billion fine, while substantial, might be a manageable cost for a company like Anthropic, which projects massive future revenues. This suggests that the financial penalty for illegal data acquisition might be seen as a cost of doing business rather than a deterrent to the core training methodology. The legal landscape remains complex, however, as copyright law has not been updated since 1976, forcing judges to apply decades-old statutes to cutting-edge technology.

Thomson Reuters

While the Anthropic case offered some clarity on AI training, other rulings highlight the 'fair use' doctrine's nuances, particularly concerning market competition. The case involving Thomson Reuters suing Ross Intelligence provides a contrasting example. Here, Ross Intelligence was found to have copied Thomson Reuters' content to build a directly competing AI-based legal platform.

Judge Stephanos Bibas ruled that Ross's use was not transformative because it lacked a 'further purpose or different character' than Thomson Reuters' original work. This decision underscores that if an AI model is trained to directly compete with the source material's market, courts are more likely to rule against it. This distinction is critical for authors, who might argue that chatbots generating synthetic books directly compete with their livelihoods, an argument that has yet to prevail in court but remains a significant point of contention.

Key points

  • Training AI models on copyrighted works is a complex legal issue, with current copyright law dating back to 1976.
  • A landmark ruling against Anthropic penalized the company for pirating books from illegal sources, not for the act of AI training itself.
  • Legal experts suggest that AI training might be viewed as 'reading' or 'consuming' a work, which is not prohibited by copyright law, rather than 'copying'.
  • The 'fair use' doctrine, particularly the 'transformative' nature of the use and its impact on the market, is central to these cases.
  • Courts tend to frown upon AI training that directly competes with the original copyrighted work's market, as seen in the Thomson Reuters vs. Ross Intelligence case.
The Upside

The legal distinction between AI training as a form of 'reading' and direct market competition could foster innovation by allowing AI models to learn from vast datasets while still providing avenues for creators to protect against direct infringement. This nuanced approach might lead to a balanced framework that supports both technological advancement and intellectual property rights.

The Downside

The ongoing legal ambiguity and the slow pace of copyright law reform leave authors vulnerable, as AI models continue to be trained on their works without explicit consent or clear compensation mechanisms. This could undermine creators' livelihoods and lead to prolonged, costly litigation, creating significant uncertainty for both the creative and AI industries.

Originally reported at

techcrunch.com

Discernion covers the story. Read the full piece at the source.

Tagsaicopyrightregulationpolicyethicsllmsfair-use

Author

Amanda Silberling

Intelligence analysis by

Gemini 2.5 Flash

Published

Aug 23, 2026

Source

techcrunch.com

Share

Topics

aicopyrightregulationpolicyethicsllmsfair-use

Related

More from this desk

Aug 24·scmp.com

A rural city, once known for livestock, now provides China’s AI computing fuel

Ulanqab, a city in Inner Mongolia previously known for livestock, is transforming into a major AI computing hub, leveraging its abundant wind power and cool climate.

Aug 24·scmp.com

Alibaba sets price in US$10.2b billion new share offer, drops 10% on market open

Alibaba Group Holding is raising US$10.2 billion through a new share placement to fund its artificial intelligence expansion, issuing 710 million shares at a discount, which led to a more than 10% drop in its stock price on market open.

Six boys from the Roehampton esports team celebrating a win at a competition.
Aug 23·bbc.co.uk

Why students are being paid £2,000 to play computer games

The University of Roehampton is offering esports scholarships of up to £2,000 to students who can balance academic progress with competitive gaming. This initiative aims to recognize talent in the growing esports field and provide financial support.

Aug 23·techcrunch.com

Who's behind the new 'stealth model' Ox Alpha?

A mysterious new AI model called Ox Alpha has been released, and speculation is rife about who actually built it. The model was described as 'very impressive' by Stripe CEO Patrick Collison, but its origins remain unclear.