Anthropic's $1.5 billion settlement over pirated training books establishes a per-work price benchmark for AI copyright liability even as the DOJ argues lawfully acquired training data is fair use.
Anthropic's $1.5 billion settlement over pirated training books establishes a per-work price benchmark for AI copyright liability even as the DOJ argues lawfully acquired training data is fair use.

Anthropic's $1.5 billion copyright settlement — the largest in U.S. history — has set a $3,000-per-book benchmark for pirated AI training data, even as the Justice Department argues lawfully acquired material requires no license at all.
"My big fear when this was all set up is that not all publishers keep great records of what books they've reverted rights to," said Mary Rasenberger, chief executive of the Authors Guild, which represents more than 18,000 writers. "So they should be taking it off their catalog."
The settlement, approved July 20 by a judge in the Northern District of California, covers roughly 482,000 books Anthropic downloaded from pirate libraries Library Genesis and Pirate Library Mirror to train its Claude chatbot. Each affected author or publisher receives about $3,100 per covered work. But as the settlement administrator begins notifying rights holders of competing claims, authors and publishers have staked overlapping ownership on repeated occasions, Rasenberger said. She said she did not believe publishers were acting in bad faith but acknowledged the process was "ripe for such misunderstandings."
The payout creates a clear price signal for non-consensual training data use, even as the DOJ's September 1 statement of interest in the consolidated OpenAI copyright litigation argues that training on lawfully acquired material is "extraordinarily transformative" fair use. If courts adopt the government's position, creators would need to secure compensation through licensing deals or new legislation rather than litigation.
The settlement administrator has begun informing authors and publishers of discrepancies in ownership claims across the more than 482,000 titles covered by the agreement. For many titles, ownership is undisputed, but competing claims have emerged repeatedly, Rasenberger said. The Authors Guild chief said publishers may not maintain complete records of books whose rights have reverted to authors, creating the potential for funds to be directed to the wrong party.
The dispute layer adds complexity to what was already the largest copyright class-action settlement in U.S. history. The $1.5 billion figure — roughly $3,100 per covered work — reflects the specific harm of downloading books from pirate repositories rather than purchasing them, a distinction that matters for the broader legal environment. The settlement does not address whether AI training on lawfully acquired material constitutes fair use, a question that remains unresolved at the appellate level.
The DOJ's filing in In re OpenAI, Inc. Copyright Infringement Litigation — MDL No. 25-md-3143 before U.S. District Judge Sidney H. Stein in the Southern District of New York — argues that training large language models on copyrighted works constitutes transformative fair use under Section 107 of the Copyright Act. The government cited Authors Guild v. Google (2d Cir. 2015), where the Second Circuit upheld Google's right to scan millions of books for search indexing without reproducing the books themselves.
The position aligns with Judge William Alsup's June 2025 ruling in Bartz v. Anthropic, which described training on lawfully acquired books as "quintessentially transformative" and "among the most transformative many of us will see in our lifetimes." But the Anthropic settlement covered pirated downloads — not lawfully acquired material — leaving the core fair use question unresolved. Two federal judges in the Northern District of California reached broadly similar conclusions on comparable facts in June 2025, but neither opinion controls the SDNY proceeding.
The DOJ's filing also invokes national security stakes, arguing that restrictive copyright interpretations could hand competitive advantages to Chinese AI developers operating under different intellectual property frameworks. That framing aligns with the White House's March 2026 National AI Legislative Framework, which stated that AI training on copyrighted material does not inherently violate copyright law.
The Copyright Office's May 2025 report took a more cautious position, concluding that certain AI training uses cannot be defended as fair use and endorsing voluntary licensing markets. The report came one day after the Trump administration dismissed the Register of Copyrights, drawing scrutiny about political pressure on the Office's findings.
Meanwhile, several AI companies have moved toward licensing deals as litigation risk mounts. Disney struck a $1 billion Sora character licensing agreement with OpenAI in December 2025, covering more than 200 characters from Mickey Mouse to Marvel heroes. Warner Music Group settled copyright disputes with Suno and Udio in November 2025, agreeing to launch licensed AI music platforms.
Judge Stein has ordered The New York Times to show cause by September 11 why its case should not be stayed pending summary judgment rulings in other MDL cases. Defendants may respond by September 18. The first appellate ruling on AI training fair use will shape the industry for a decade, determining whether AI companies must license training data or can rely on fair use for lawfully acquired material. For creators, the court-based avenue for compensation remains open but now runs against the executive branch's declared position — pushing the compensation question toward Congress and the licensing table.
This article is for informational purposes only and does not constitute investment advice.