Essay
25

Essay · Published Jun 28, 2025 · Updated Sep 7, 2026

Anthropic's Fair-Use Win Did Not Excuse Its Pirated Library

The court treated AI training as transformative while leaving Anthropic exposed for keeping a library of pirated books.

A field-guide drawing of a Pacific banana slug
In this article5 sections

In Bartz v. Anthropic, Judge William Alsup treated training on copyrighted books as fair use. That did not excuse how Anthropic obtained and kept millions of pirated books. The distinction matters more than a simple “AI companies won” headline.

I use Claude and other coding agents every day. That makes the provenance question more practical, not less. A useful model does not turn an unlawful copy into a lawful one, and a lawful training purpose does not excuse every way a company assembled its corpus.

Looking back This post described the June 2025 summary-judgment ruling. On July 20, 2026, the court granted final approval to the class settlement and entered judgment. The case did not proceed to the jury trial anticipated below.

Anthropic moved from piracy to purchase

Anthropic’s data collection history explains why the ruling split in two. The company started with pirated collections, then changed course.

Co-founder Ben Mann downloaded Books3 in early 2021 — 196,640 pirated books. By June 2021, he had downloaded at least five million books from Library Genesis. In July 2022, Anthropic added two million more from the Pirate Library Mirror. All of these sources were known to contain unauthorized copies.

Then Anthropic changed course entirely. In February 2024, they hired Tom Turvey, former head of partnerships for Google’s book-scanning project. His mission: obtain “all the books in the world” while avoiding “legal/practice/business slog.”

Turvey’s team spent millions buying print books, often used. They stripped bindings, cut pages to size, scanned them into PDFs, and discarded the physical copies.

The ruling

Judge Alsup’s 32-page decision draws lines that will define AI copyright law.

Fair Use

AI Training: The court called training LLMs on copyrighted books “spectacularly transformative.” The judge compared it to how humans learn from reading — forcing people to pay “each time they read, each time they recall from memory, each time they later draw upon it when writing new things” would be unthinkable.

Purchased-and-Scanned Books: Converting bought print books to digital for internal use is fair use, though on narrower grounds. The court treated it as format shifting.

Not Fair Use

Pirated Central Library: Maintaining a permanent digital library of millions of pirated books is not fair use. The court stressed that Anthropic kept pirated copies even after deciding not to use them for training.

What this decision means for AI companies

  • Training on copyrighted material was fair use on this record because the court found Anthropic’s use transformative and not a substitute offered to the public
  • Purchase and one-to-one digitization for Anthropic’s internal library was fair use in this decision
  • No special carveout for AI: The court stated plainly, “There is no carveout from the Copyright Act for AI companies”
  • Piracy isn’t excused by downstream fair use

This was one federal district-court decision on a specific record, not a national rule that settles every training case.

What happened next

Anthropic and the author class settled the piracy claims before trial. The July 2026 final-approval order also required Anthropic to destroy original files downloaded from Library Genesis and Pirate Library Mirror, subject to legal-preservation obligations.

Other AI copyright cases continue on different facts. That is another reason not to stretch this ruling into a universal answer about model training.

Unanswered Questions

  • What counts as “transformative” use across different AI contexts?
  • How does fair use apply to images, videos, or code?
  • What licensing models emerge to serve both creators and AI companies?

The ruling protects a purpose, not every path used to reach it. Training may be fair use while acquiring and retaining the training material is still infringement. For anyone building AI systems, provenance did not become optional just because the model use was transformative.

One quick signal

Did this earn your time?

What was missing?

Thanks. That gives me something concrete to check.