News
AI Companies Are Buying and Destroying Old Books to Get AI-Free Training Data

Global AI labs are buying up secondhand book collections and subjecting them to destructive scanning in search of text written before the ChatGPT era. Canadian firm Zoom Books is already seeking more than 800,000 Polish titles.
Contents
Hundreds of thousands of old books are disappearing from secondhand bookshops around the world. They aren't ending up on new readers' shelves, but in scanners, after which they're destroyed. Companies training large language models are paying for them in bulk, chasing something the internet is starting to run short of: text written by humans before the web was flooded with content generated by AI itself.
The phenomenon was documented by journalists at Gazeta.pl, citing reports from secondhand bookshops in Poland and abroad. The scale of the buying is large enough that some booksellers have started refusing to deal with companies that make no secret of their purpose: the books are meant solely as training material for AI models, not to return to circulation among readers.
Why old books specifically
The reason isn't sentimental, it's technical. Successive generations of language models are increasingly trained on data that itself largely comes from an internet flooded with content generated by earlier models. Researchers call this process model collapse, a degradation in quality where the algorithm replicates its own errors, oversimplifies knowledge, and gradually drifts away from reliable sources.
Books published before ChatGPT's public debut in late 2022 are therefore a rare resource of text uncontaminated by AI: coherent narrative, vetted specialist terminology, and language untouched by an algorithm. For AI labs, it's raw material comparable to clean spring water in a world of polluted internet data rivers.
Why destruction, not just scanning
What matters is how companies handle the physical copy after digitization. They don't resell the scanned books or donate them to libraries, they cut the spines and pulp the paper. That's not incidental, it's a consequence of the ruling in the 2025 Bartz v. Anthropic case, in which a US federal court found that training AI on scanned books falls within fair use, provided the number of copies in circulation doesn't increase.
Physically destroying the original is therefore proof that only the medium changed, not that the work was duplicated. Anthropic runs an operation called Project Panama for this purpose, aiming to digitize as many of the world's existing books as possible. The company had earlier paid a $1.5 billion settlement over training models on pirated digital libraries such as Books3.
That would be completely at odds with my understanding of what a bookseller does - Marcin Gałązka, Zakładka secondhand bookshop, Warsaw
The Polish angle
The phenomenon has already reached Poland. Zoom Books, a Canadian company operating as Acirassi Books Ltd. and describing itself as an eco-friendly leader in mass resale of used books, is seeking more than 800,000 Polish titles marked with ISBN numbers, bought in bulk rather than individually. Offers went out to, among others, the Zakładka secondhand bookshop in Warsaw and Szarlatan in Wrocław, run by Stanisław Karolewski. Both refused.
Marcin Gałązka warned that thousands of books could soon be headed for digitized pulping. Aldona Mackiewicz of the Toruń Antykwariat Księgarski pointed to a different problem: booksellers have no way to verify whether a buyer actually intends to resell the books to readers or will send them straight to a scanner.
What it means for the market
For publishers and authors, the situation is ambiguous. On one hand, bulk sales of old, often low-print-run and hard-to-find titles bring booksellers real income. On the other, the books physically vanish from the secondhand market, and their content ends up solely in closed AI model databases, with no chance for a reader to borrow or buy a copy again.
Similar legal disputes are already playing out beyond the secondhand book market. Encyclopedia Britannica sued OpenAI, accusing the company of copying and mimicking encyclopedia content while training its models. These cases show that the line between lawfully sourcing training data and copyright infringement is still being contested in US courts, despite the ruling in the Anthropic case that favored the AI industry.
For now, Poland has no regulations directly addressing the buying up of books for AI training, and the decision to sell collections to foreign companies rests with individual booksellers. Some, like Gałązka and Karolewski, treat refusal as a matter of professional ethics, while others view the transaction in purely business terms.
