Amazon buys rare used books, then destroys them to train AI

News
Tuesday, 18 August 2026 at 19:04
Amazon koopt zeldzame tweedehands boeken en vernietigt ze voor AI-training
Amazon is buying physical books at scale, slicing off the spines, and scanning the loose pages to harvest training data for artificial intelligence. That’s the conclusion of an investigation by 404 Media. Reporters tracked a shipment with a hidden tracker to an Amazon facility in Las Vegas, where staff, according to the report, process books for AI training.
The investigation exposes a striking legal pathway previously used by Anthropic. AI companies can legally buy physical books, digitize them, and destroy the paper copies. In 2025, a U.S. judge ruled in a case involving Anthropic that this specific method could fall under “fair use.”
That creates an uneasy reality: a company can buy a used book for a few dollars, digitize its entire contents, and use that text to build commercially valuable AI systems—without automatically securing a separate license from the author or publisher.
So the big question isn’t just what Amazon is doing with these books, but whether current copyright rules still fit the scale at which AI firms process human knowledge.

How 404 Media uncovered Amazon’s book pipeline

404 Media cracked the operation by literally tracking a physical book. Journalist Emanuel Maiberg teamed up with a seller who supplied a batch of hard-to-find titles, suspected to be destined for an AI company.
A tracker was hidden in one of the books. The shipment moved across the United States and, according to 404 Media, ended up at an Amazon warehouse in Las Vegas, Nevada.
Workers at that site told the outlet that large quantities of printed books are received there. The spines are removed so the loose pages can be scanned much faster. The original book is destroyed in the process.
404 Media identifies the site as VGT3. The team there reportedly even has its own logo: a dinosaur holding a book.
According to the investigation, Amazon has built infrastructure to convert physical books into digital data at industrial scale.

Why secondhand books are AI gold

Used books are a rich data source because much of their content isn’t freely available or neatly structured online. Older, specialist, foreign-language, and barely digitized books often contain text missing from existing training datasets.
Large language models need vast amounts of text to learn patterns in human language. Books are especially valuable because they offer long-form, edited, and coherent writing.
There’s a new twist: the public internet is increasingly flooded with AI-generated text. Older books, by contrast, guarantee material from before ChatGPT and its peers existed.
A forgotten technical manual from 1986 may be worthless to most collectors but still valuable to an AI company for the knowledge and human-written prose it contains.
AI is giving secondhand books a whole new economic role: they’re turning into training fuel.

Amazon mirrors Anthropic’s earlier playbook

Amazon isn’t the first major AI company to use physical books this way. Anthropic, maker of Claude, previously bought millions of print books to build a vast digital library.
AI Wereld reported in 2025 that Anthropic bought, scanned, and destroyed millions of physical books. The company enlisted industry veteran Tom Turvey, formerly of Google Books.
Court documents showed Anthropic aimed to build a central library containing virtually “all the books in the world.” It bought physical copies at scale, removed covers and bindings, and digitized the pages.
The parallels with the new Amazon findings are hard to miss. Amazon also has close business ties with Anthropic and invested billions in the AI firm. Still, 404 Media’s investigation describes an Amazon-run operation and does not, on its own, prove the two companies are jointly executing their book-scanning projects.

Why destroy the books at all?

Destroying books isn’t a technical requirement. Libraries and archives have digitized books intact for years.
But destructive scanning is faster and simpler at industrial volumes. Once a book’s spine is sliced off, loose pages can run through automatic document scanners. Optical Character Recognition (OCR) software then converts the images into digital text.
But there’s also a key legal angle at play.
In June 2025, U.S. federal judge William Alsup ruled in Bartz v. Anthropic that, under Anthropic’s specific circumstances, digitizing lawfully purchased physical books could qualify as fair use. The company replaced the purchased physical copy with an internal digital version and did not distribute that copy further.
That distinction is crucial.
The ruling does not say every AI company can freely copy any copyrighted book. It’s also a U.S. district court decision, not a binding nationwide precedent.
Still, the case points AI companies to a legally intriguing path: buy the physical book legally, make an internal digital version, destroy the original, and use the resulting data internally.

Is this a loophole in copyright law?

Practically, the setup looks like a loophole, but legally that’s too strong. U.S. fair use is intentionally flexible, and courts assess uses against several factors.
Yet the Anthropic ruling exposes a striking gap between the economics of a book and those of AI training.
A consumer who buys a used book typically doesn’t pay the author again. That fits the principle that a lawfully purchased physical copy can be resold.
An AI company can buy that same copy—but for a very different purpose. It’s not just reading the book. It can turn the contents into machine-readable data and use them to train models that serve millions of users and could be worth billions.
The price of a used book and the potential economic value of the training data are wildly out of sync.
That’s exactly why the AI-and-copyright debate is shifting from “was the copy legally obtained?” to a bigger question: should commercial use as AI training data require a separate right or licensing model?

But Anthropic also got a $1.5 billion bill

The Anthropic case also shows why lawful sourcing matters. Anthropic didn’t rely solely on purchased books.
According to court filings, the company also downloaded millions of books from piracy libraries like LibGen and PiLiMi. Judge Alsup drew a sharp line between those files and the physically purchased books.
Scanning the bought copies was deemed fair use in that case. Building a permanent library of illegally obtained books was not.
Anthropic ultimately settled with authors and rightsholders for $1.5 billion. The court granted final approval to that agreement in July 2026.
That’s relevant for Amazon. The legal takeaway from Anthropic isn’t simply that training AI on books is “legal.” How a company acquires the source can be just as important.

Is scanning used books good or bad?

Honestly, both sides have defensible arguments.
Digitization can preserve knowledge that might otherwise vanish. Used bookstores, libraries, and warehouses hold vast numbers of books with little commercial demand. Some are eventually discarded or recycled.
From that angle, digitization creates value. The text survives and can fuel AI systems that make knowledge more accessible.
But there’s a hard limit.
An internal Amazon dataset is not a public archive.
If a scarce book is destroyed and its scan exists only inside a tech company, a library, researcher, or citizen doesn’t automatically get access. The knowledge is preserved digitally—but effectively privatized.
That’s fundamentally different from a public library digitization project.

Rare titles create the toughest dilemma

The word “rare” also needs care. A rare book isn’t automatically a valuable or unique historical artifact.
A little-sold manual from 1974 might now exist in only a few copies without commanding collector prices. The same goes for local publications, foreign-language works, old textbooks, and small print runs.
404 Media says it’s deliberately not naming the titles in the tracked shipment. According to the outlet, few copies were in circulation, partly because some works were originally printed in small runs or published in less widely spoken languages.
That makes it hard to pinpoint the exact cultural damage.
There’s a world of difference between destroying one copy of a book that exists elsewhere in dozens of editions, and destroying a copy with unique notes, provenance, or other historical value.
That oversight becomes crucial with industrial-scale processing.

Authors don’t automatically get paid for AI training

For writers and publishers, the real pain point lies elsewhere. Buying a physical secondhand copy doesn’t mean the original creator gets paid again.
The AI company pays the seller of that copy. Then it can use the content to build a commercial AI model.
That tension now runs through nearly the entire AI industry. Meta, OpenAI, Anthropic, and others have all faced legal fights over when and how copyrighted material can be used for model training.
Political pressure is rising too. The European Parliament has proposed rules to give rightsholders more control and compensation when protected material is used as training data.
The European debate matters because U.S. fair use doesn’t map cleanly onto European copyright law.

The battle for quality AI data moves into the physical world

Amazon’s research ultimately shows just how valuable human-written text has become. The biggest AI players have massive compute, but powerful chips are worthless without high-quality data to learn from.
So begins a new hunt for sources that aren’t yet tapped out.
First came large-scale scraping of the open web. Then licensing deals with media companies and databases. At the same time, lawsuits exploded over books, news, music, and images.
Now even the secondhand book market is part of the race.
Here’s the twist: a book can have two wildly different values. For a bookstore, an obscure title might be unsellable. For an AI company, that same publication could hold a fragment of unique training data.
Whether that’s good or bad depends on more than the act of scanning. The key questions are which books are destroyed, whether other copies exist, how the digital text is used, who gets access to that digital copy, and whether the original creators share in the economic value it generates.
Until those issues are clearly settled, large AI companies have every incentive to push the edges of existing copyright law.
Amazon now shows the hunt doesn’t stop at the internet. Even dusty stacks in secondhand shops have become valuable feedstock for the AI economy.
loading

Loading