Anthropic’s $1.5 Billion Settlement: Not the AI Copyright Reckoning You Think It Is
You’ve probably seen the headlines: AI company Anthropic has agreed to a landmark $1.5 billion settlement with book authors and publishers. It’s been hailed as the largest payout in US copyright history and a “Napster moment” for the AI industry. From a European perspective, where copyright is often seen through a lens of strict authorial rights and enumerated exceptions, this might look like a straightforward victory against an AI giant that indiscriminately scraped copyrighted works.
However, the reality is far more nuanced and, for the AI industry, surprisingly encouraging. This settlement is not about whether training AI on copyrighted material is legal. In fact, on that core question, the AI company won. This case pivots on a much simpler, more classic issue: the difference between using legally acquired material and using pirated material.
Let’s unpack what really happened.
The Real Ruling: Training AI is “Exceedingly Transformative” Fair Use
The most crucial part of this story happened in June 2025, before any settlement was reached. The judge in the case, William Alsup, made a critical ruling that AI developers should pay close attention to. He found that when Anthropic legally acquired books (by purchasing and scanning them), its subsequent use of them for training its AI models was a fair use.
For those in the EU unfamiliar with the concept, “fair use” is a flexible, four-factor balancing test that courts in the U.S. apply on a case-by-case basis to determine if a use of copyrighted material is permissible without a license. The most important factor is often whether the use is “transformative”: that is, whether it re-purposes the original work for a new function or meaning.
Judge Alsup’s ruling was decisive on this point. He declared that AI training is “exceedingly transformative” and that “the technology at issue was among the most transformative many of us will see in our lifetimes”. He also found that a market for licensing works for AI training is “not one the Copyright Act entitles Authors to exploit”. In essence, the court affirmed the central argument of the AI industry: the act of learning from data is a transformative process protected by fair use, provided the material was obtained legally.
The Real Sin: Sourcing from “Shadow Libraries”
The case against Anthropic had two parts. The first was about the act of training, which the judge largely blessed. The second, however, was about Anthropic’s data acquisition practices. The court found that alongside its legally purchased books, Anthropic had also downloaded and stored millions of books from pirate “shadow libraries” like Library Genesis and Pirate Library Mirror.
This is where Anthropic’s fair use argument collapsed. The judge ruled that knowingly downloading and storing pirated material was illegal infringement. It didn’t matter that the intended future use (training) might have been fair; the initial act of acquiring pirated content was not.
This distinction is key. As Brandon Butler, Executive Director of the US copyright advocacy group Re:Create, put it, the settlement is “only really about *how Anthropic obtained* some of its training data…, not *what Anthropic did with the data”.
This led to a simple, but terrifying, financial calculation for Anthropic, thanks to another unique feature of US copyright law: statutory damages. In the US, a copyright holder can sue for pre-determined damages—up to $150,000 per infringed work if the infringement is willful—without having to prove actual economic harm. When you multiply that figure by millions of pirated books, the potential liability is astronomical. Faced with this “extreme duress,” where a defendant could risk “trillions,” the $1.5 billion settlement becomes a pragmatic business decision to mitigate an existential risk.
Who Are the “500,000 Authors”? The Class Action Complexity
The settlement is set to pay $3,000 each to 500,000 authors. But the idea of a single, unified group of wronged creators is a legal simplification that masks a messy reality. As a brilliant analysis by the Authors Alliance points out, the collection of books from Library Genesis is far from a simple list of recent bestsellers.
Giving full credit to the Authors Alliance for their research, their deep dive into the LibGen metadata reveals several complicating factors:
- A Heterogeneous Mix: The dataset is incredibly diverse, containing everything from recent commercial fiction to decades-old academic textbooks, small press publications, and works that may already be in the public domain.
- Identification Issues: A huge portion of the books in the dataset lack the ISBN or ASIN identifiers required to even be a part of the class action—less than half of the fiction and just over half of the non-fiction titles have them. This makes identifying and notifying potential rightsholders a massive challenge.
- Public Domain and Open Access: The dataset includes works by long-dead authors like Shakespeare and Charles Dickens, as well as thousands of books published under open-access Creative Commons licenses that may permit the very use Anthropic was engaged in.
This complexity shows that the narrative of “AI vs. Authors” is too simplistic. The interests of a major publisher of a current bestseller are vastly different from those of an academic author whose 30-year-old monograph was openly licensed. Treating them all as a single class with identical grievances is problematic.
Key Takeaways: A Settlement Driven by Flaws in the System
While this case offers lessons on data sourcing, the most profound takeaways are not about AI, but about the peculiarities of the US legal system. The outcome was shaped less by the merits of copyright in the AI era and more by systemic pressures that force pragmatic, albeit unsatisfying, resolutions.
- The Perverse Incentive of Statutory Damages. The single biggest factor driving this settlement is the US copyright law’s provision for statutory damages. This system allows for damages of up to $150,000 per willfully infringed work, regardless of the actual economic harm caused. When dealing with a dataset of millions of books, Anthropic faced a theoretical liability in the trillions of dollars. This threat creates “extreme duress,” where even a “righteous person would still likely choose to capitulate”. The $1.5 billion settlement, therefore, should be seen less as a fair valuation of harm and more as a risk-management strategy against a disproportionately punitive legal weapon.
- The Legal Fiction of a Unified “Class”. The case proceeded as a class action, which assumes that a small group of plaintiffs can represent the interests of a much larger group. However, as the Authors Alliance has argued, the court certified this class with “almost no meaningful inquiry into who the actual members are likely to be”. The dataset is not a monolithic collection of recent bestsellers; it’s a complex mix of works spanning decades, from major publishers to small academic presses, including works that are likely in the public domain or openly licensed. Lumping them together into a single class oversimplifies their diverse rights and interests.
- A One-Size-Fits-All Settlement Ignores Reality. The combination of extreme financial pressure and a poorly defined class leads to a settlement that papers over crucial complexities. A solution that treats every book identically—whether it’s a recent novel, a public domain work by Dickens, or an openly licensed scientific paper—cannot adequately reflect the “diverse rights, licenses, and interests actually at stake”. As the Authors Alliance concludes, future AI litigation would benefit from a far more careful analysis of what’s actually in these datasets before assuming all uses and all works raise the same legal questions.
As Europe defines its own course on AI regulation, this case serves as a powerful cautionary tale. We can only hope that the EU’s framework will prove more resilient and foster innovation by providing clarity and fairness without the distorting influence of a litigation system prone to such perverse incentives.
Written by Caroline De Cock, LL.M. , Head of Research.
