News
AI Companies Are Buying Up and Destroying Old Books in Search of Clean Data

Used-book dealers in the Netherlands, Germany, Switzerland and Spain are receiving orders for thousands of niche titles that end up not with collectors but under a scanner's blade. AI labs are hunting through printed books from before 2022 for the last remaining stores of knowledge uncontaminated by machine-generated content.
Contents
Used-book dealers in several European countries have noticed the same unsettling pattern in recent months: someone orders hundreds of random, niche titles in bulk, pays without haggling, and never asks about the condition of the cover. The books don't end up with collectors or readers, though, they go straight into industrial scanners and are then destroyed. Having exhausted the internet's easily accessible resources, the AI industry has turned to printed books from before 2022 as a last source of knowledge created entirely by humans.
Dutch used-book dealers were the first to sound the alarm. Pieter de Vries of De Vries & De Vries in Haarlem received an email from someone identifying herself as Nataly from Singapore-based 2077AI, with a list of nearly three thousand titles to buy. Similar orders began arriving at shops in Germany, Switzerland and Spain, and local media, from the Dutch outlet BNR to Swiss broadcaster SRF, described it as a new, systematic buying model.
Used-book dealers under pressure
For dealers, the situation is a mixed blessing. On one hand, bulk orders mean quick cash for books that had sat unsold on shelves for years. On the other, sellers have no certainty about what actually happens to the copies after the sale, and the suspicion that unique or rare academic editions end up as pulp after a single scanning pass has drawn pushback from those concerned with preserving literary heritage.
Canadian company Zoom Books operates differently from 2077AI, it doesn't send lists of titles, but instead buys random batches of niche publications every day, in the middle of the night. The company itself describes this activity as a standard book recycling and trading model, without specifying who ultimately buys the scanned data.
Anthropic and Project Panama
The fullest picture of what such an operation looks like at scale came from court documents related to Anthropic that were unsealed this year. Since 2024, the company had been running an internal effort called Project Panama, under which it hired Tom Harvey, previously of Google Books, bought stock from the nonprofit Better World Books, and contracted scanning, including destructive scanning, to a company called Datamation. Order volumes ranged from half a million to two million volumes over roughly six months.
The legal dispute over these practices ended with Anthropic settling with book authors for $1.5 billion, as nowosci.ai reported earlier. The newly unsealed documents reveal something beyond just the size of the settlement, though: the scale and method behind acquiring physical books, a practice that turns out to be far broader than a single case involving one company.
Why books from before 2022
The reason behind this buying frenzy is a phenomenon known as model collapse. Language models trained on content that earlier models themselves generated gradually lose quality, repeat errors and become more prone to hallucination. The internet, after the wave of generative tools, is increasingly contaminated with this kind of material, so AI labs are looking for data from before large language models became widespread, meaning roughly before 2022, when text online and in print was presumed to be exclusively human-made.
A printed book has an additional edge over web text: it has gone through editing, proofreading and the test of time, which in the eyes of AI companies makes it a more reliable material for teaching reasoning and clear writing than fragmentary, often unverified content from the web.
A secret under NDA
ISBNdb, which holds one of the largest databases in the publishing market, has turned itself into a broker handling orders for AI labs ranging from a thousand to a million copies at a time, always under strict non-disclosure agreements. The service itself admits on its own site that the discretion is no accident.
The image problem is real. "AI company destroys two million books" is not a headline that wins sympathy. - ISBNdb
The legal question has already been partly settled by federal judge William Alsup, who ruled that digitizing a purchased book falls within fair use, since the physical copy is destroyed in the process rather than duplicated. That precedent gives AI companies a legal basis to keep buying up and destroying books at scale, despite growing public pushback.
What this means for the book market
For Poland's used-book market, the phenomenon still feels distant, but the mechanism is the same everywhere: demand for training data is growing faster than the supply of uncontaminated text, and used-book dealers around the world are sitting on physical resources the internet doesn't have. If the practice described in the Netherlands, Germany, Switzerland and Spain spreads further, other markets, including those in Central Europe, could see the same sudden interest from buyers whose only criterion is volume, not collector's value.
For now, the biggest open question is how many rare editions, ones that exist nowhere else, have already vanished for good under the blade of an industrial scanner before anyone could digitize them non-destructively. Dealers and librarians are starting to demand transparency about who is buying and for what purpose, but with orders shielded by confidentiality clauses, it's hard to get a full picture of the scale involved.

