The used-book trade has spent decades running on predictable rhythms: collectors chasing first editions, students hunting cheap textbooks, estate sales moving boxes of paperbacks nobody wants. That pattern broke recently. Orders started arriving that made no sense to anyone who knows the business — hundreds of unrelated titles at once, no interest in condition, no haggling over price, no apparent theme beyond one shared trait.
A used bookseller receives an order for 400 volumes spanning history, botany, regional law, and German-language economics. The only thing the titles have in common is that each one carries an ISBN. That's the tell.
Why books printed before 2022 became premium training data
The value here isn't nostalgia. It's contamination.
The demand centers on books published before 2022, prior to the point when synthetic text became widespread online. Anything written after generative models went mainstream carries the risk of being partly machine-written — and feeding model output back into model training degrades quality. The hazard has been compared to inbreeding, and it's more commonly called model collapse. Printed books from before the boom are sought precisely because they contain no AI-generated text.
There's a second reason print wins over web scraping. ISBNdb, which says it operates the world's largest book database, makes the pitch directly on its site: the best AI training data available is sitting on a shelf. Books, the company argues, represent curated, peer-reviewed, domain-specific human knowledge structured in a way no web crawl can match — dense, edited, authoritative.
A scraped forum thread is noise. A university press monograph is signal that's already been edited, fact-checked, and organized by someone who knew the field.
How the buy-scan-destroy pipeline works
The mechanics are industrial, not delicate.
Books go through a hydraulic cutting machine that removes the pages from the binding. The loose pages then run through high-speed industrial scanners. What's left of the physical book is discarded or recycled.
Non-destructive scanning exists — flatbed rigs, robotic page-turners — but the destructive route is faster and cheaper, so that's the one companies choose. The result is the destruction of millions of books.
The middlemen and the anonymity layer
Most of this activity doesn't happen under a lab's own name.
ISBNdb says it can coordinate purchases ranging from 1,000 to 1 million books, sourced from secondary markets where libraries, retailers, and individuals offload used inventory. The company also offers non-disclosure agreements to buyers.
It's candid about why. Its own marketing acknowledges that the optics problem is real — that a headline about an AI company destroying two million books isn't one that wins public sympathy.
Worth noting for accuracy: there is no direct confirmation linking these specific bulk purchases to named AI companies. The anonymity is the point of the arrangement.
The court ruling that made destruction the rational choice
The legal groundwork was laid in Bartz v. Anthropic, where U.S. District Judge William Alsup of the Northern District of California ruled in June 2025 that digitizing legally purchased print books and using the resulting files to train large language models qualified as fair use.
The reasoning turned on substitution rather than multiplication. Alsup wrote that each purchased print copy was copied to save storage space and enable searchability, and that the print original was destroyed — one replaced the other. Because the digital copy was never displayed, shared, or sold outside the company, he found the process clearly transformative.
Read that closely and the incentive structure becomes obvious. Keeping the physical book means paying to store it and weakens the "replacement" argument. Destroying it strengthens the fair use claim and eliminates the warehouse. The ruling handed the industry a template, and companies now cite it explicitly. Booksellers have pointed to the same precedent as the reason acquisitions have continued.
Anthropic's own operation ran at serious scale. Court documents reported by the Washington Post describe "Project Panama," which spent tens of millions of dollars using contractor Datamation to scan books for Claude's training data. The company brought on Tom Turvey, formerly head of partnerships for Google Books, to handle acquisition, and bought volumes from used-book sellers including Better World Books and World of Books. One vendor proposal referenced converting between 500,000 and two million books over a six-month window.
What sellers are seeing on the ground
The buying patterns are distinctive enough that dealers now recognize them. The markers are abnormal volume, subject-agnostic ordering, and complete indifference to price. Dealers have said the purchases don't resemble collector or resale behavior at all.
The activity is international. Booksellers in the Netherlands, Germany, Switzerland, and Spain have reported bulk orders for niche and specialized titles. Pieter de Vries of De Vries & De Vries in Haarlem fielded a request for nearly 3,000 titles from a Singapore-based firm, and a German seller described overnight orders from the Canadian company Zoom Books covering unrelated academic works.
For sellers sitting on dead inventory, the money is real and the feelings are mixed. One bookseller told 404 Media that weekly sales jumped from roughly 20 books to several hundred once AI buyers entered the market. He described it as financially good and useful for clearing stock unlikely to move otherwise — while worrying that uncommon and out-of-print titles are being permanently lost after scanning.
The titles that don't come back
That's the part with no easy answer. Rare, foreign-language, and out-of-print volumes face genuine extinction risk when the last surviving physical copies are cut apart and pulped. The text survives inside a private training corpus. The object doesn't, and neither does public access to it — the digital copy is never shown, shared, or sold outside the company, which is exactly the condition that made the process legally defensible in the first place.
Regulation in other countries hasn't solved the issue either. The EU's Digital Single Market Directive makes it harder for rights holders to opt out of text and data mining, so authors and their estates have little real power to control how their work is used.
The reporting that surfaced much of this came from Emanuel Maiberg at 404 Media.

