Anthropic had millions of print books sliced, scanned, and then destroyed to train its AI models.
Leaked internal documents show the project was deliberately kept secret. The revelations have ignited fierce global debates over copyright, cultural heritage, and the future of knowledge.
For years, the fight focused on AI training with illegally downloaded ebooks. Now the spotlight shifts to a new playbook. Instead of scraping the web, AI firms are buying physical books at industrial scale. The spines are hydraulically cut off, every page is rapidly scanned, and the books are then pulped.
Court filings say this happened at Anthropic under the internal codename Project Panama—an initiative that, according to internal memos, aimed to “destructively scan every book in the world.” Staff were explicitly told not to discuss the project publicly.
What exactly is Project Panama?
Project Panama launched in early 2024, as Anthropic faced mounting legal pressure over using copyrighted works to train
Claude, the company’s AI chatbot.
Thousands of pages of court records indicate Anthropic spent millions of dollars buying massive lots of used books. Former Google Books executive Tom Turvey was brought in to build an industrial-scale scanning operation.
According to the documents, the process ran as follows:
- millions of books were purchased;
- the spine was hydraulically removed;
- every page was scanned individually;
- the physical books were discarded;
- the digital scans were used internally for AI training.
The court noted this method produced higher-quality data than web sources, since books contain fewer errors and are less polluted by AI-generated text.
Why is this happening now?
The timing is no accident.
Anthropic has been under fire for using millions of books from pirate libraries like LibGen to train Claude.
In June 2025, a U.S. judge ruled that training AI on lawfully acquired books can, under certain conditions, qualify as fair use. At the same time, acquiring pirated copies remains tied up in further litigation.
By simply buying physical books, Anthropic aims to sidestep much of that legal risk.
The company owns a lawful copy before digitizing it.
That makes the legal question fundamentally different from downloading illegal digital copies.
Used book sellers report sudden surges
The impact is now showing up across the global secondhand book market.
Multiple dealers say they once sold a few dozen specialty titles per week, but since last year have been flooded with hundreds of orders.
In particular:
- foreign editions;
- academic works;
- older technical literature;
- rare titles;
- out-of-print books
have become especially sought after.
Several sellers suspect many of these purchases are ultimately headed to AI companies.
Alarm over rare and fragile titles
That’s exactly what worries libraries, collectors, and authors.
While many destroyed books exist in large print runs, dealers report that hard-to-find copies are also being snapped up.
Critics fear some editions with only a handful of physical copies left could disappear for good.
There’s more.
The scans aren’t being made public.
They vanish into closed AI training datasets, inaccessible to researchers, libraries, or the public.
Critics argue this doesn’t build a digital archive for society—it creates private training corpora for commercial AI systems.
Legal doesn’t mean uncontroversial
The U.S. court drew an important line.
Digitizing a lawfully purchased book and destroying the physical copy was treated as a transformative use, akin to earlier digitization efforts.
But that ruling doesn’t erase the backlash.
Authors note they receive no compensation for the downstream AI training, even as their work is used to make commercial systems smarter.
Moreover, the ruling doesn’t change earlier accusations about using illegal digital copies.
ISBNdb now sells AI-focused bulk book data
And the practice appears to be expanding.
ISBNdb, a major international book database, now offers bulk services that let AI companies purchase hundreds of thousands to millions of books.
According to the company’s documentation, books are selected by language, topic, and publication year, then prepared for large-scale scanning.
ISBNdb itself acknowledges the PR problem. A headline like “AI company destroys two million books” would understandably spark backlash, the company says.
The AI race enters a new phase
The revelations underscore how crucial high-quality training data has become.
As the internet fills up with AI-generated text, AI firms are hunting for reliable, original sources.
Books are a prime target because they are:
- carefully edited;
- fact-checked;
- often written before the rise of generative AI.
That’s why companies are willing to spend millions of dollars on physical books that are destroyed immediately after digitization.
Why this fight is only getting louder
Project Panama touches a far bigger question than copyright alone.
Who owns human knowledge once it’s converted into AI training data?
Supporters argue the information survives in digital form, helping AI systems perform better.
Critics see a shift where cultural heritage disappears into closed, commercial models that no one outside the companies can access.
With more lawsuits targeting AI developers like Anthropic, OpenAI, Meta, and Google, the battle over training data is far from over. The coming years will likely set the legal and ethical boundaries for digitizing humanity’s knowledge for artificial intelligence.