AI Companies Are Bulk-Buying Secondhand Books — The Strange Gold Rush for “Unpolluted” Training Data

Introduction: The Strange Orders Landing at Bookshops

On August 15, 2026, The Guardian reported that secondhand booksellers across the UK and Ireland are being hit by a wave of bulk orders from mystery buyers — and the working theory is that AI companies are buying the world’s used books to feed their models.

The orders are bizarre. One recent request included an Estonian translation of John le Carré’s The Mission Song, a copy of Anne Brontë’s Agnes Grey from a specific imprint, and the October 1983 edition of Warship, a monthly magazine. “Normally a book order is on a theme but this is all over the place,” Stuart Manley, co-owner of the famous Barter Books in Alnwick, Northumberland, told the paper. His theory about who’s paying: “masses and masses of AI money being squandered.”

If that theory is right, the story is bigger than quirky bookshop gossip. It’s the clearest physical-world sign yet that the AI industry has hit its data wall — the point where all the easily-scraped internet text has been consumed, and companies are turning to printed books published decades before ChatGPT existed: text that is genuinely fresh to a model and untouched by AI-generated slop.

Anatomy of the Buying Spree

The Guardian spoke to sellers across the UK and Ireland, and reports from the US, Australia, and continental Europe describe the same pattern: large, thematically random orders placed through marketplaces like Biblio and the Amazon-owned AbeBooks, arriving from buyers in the US, Canada, and Europe.

Seller What They’re Seeing
Barter Books (Alnwick, UK)Orders began ~3 months ago; hundreds of random books sold to three buyers for about £4,000
MW Books (Claregalway, Ireland)Large, varied orders since May — “sporadic, presumably machine-driven”; from 18th-century African agriculture to 1950s racing-driver biographies
BookLovers of Bath (UK)Rush of ~200 orders between May and July; “it was overwhelming for a while”
Anonymous UK seller6,000 books ordered since January, buyers paying “top whack” with no bulk discount, deliveries converging on freight warehouses near Heathrow

Several sellers flagged the same fingerprints: opaque buyer aliases, orders from different accounts shipping to the same address, and zero interest in the discounts any normal trade buyer would demand. One buyer cited by multiple sellers, the Canada-based Zoom Books, describes itself as “North America’s leading book recycling company.” As one UK seller put it: “It’s disrupting the secondhand bookselling business because this isn’t typical for the trade.”

The Anthropic Precedent: Spines Sliced, Scanned, Recycled

The booksellers’ suspicions have a concrete anchor. The Washington Post revealed this year that Anthropic — the company behind Claude — spent tens of millions of dollars acquiring books, slicing off their spines so the contents could be scanned, and then sending the books for recycling, all under a program literally labeled “data acquisition.”

Anthropic’s response to the Guardian is a rare on-record glimpse into how frontier labs source text: “Claude is trained on a mix of publicly available web data, commercially acquired datasets, and data we generate ourselves. Sourcing books is a widely used approach for training large language models across the AI industry. None of our data acquisition programs buy and destroy rare or antiquarian books.”

In other words: buying books at scale is standard practice — and Anthropic says it is not alone. Last month, 404 Media reported that the books database ISBNdb was telling booksellers “the world’s best AI training data is sitting on a shelf,” arguing pre-2022 books are ideal because their text is unpolluted by chatbot-generated material. (ISBNdb later called that page a “test of market interest” and took it down.)

The detail that matters: the demand isn’t for rare first editions — it’s for ordinary, out-of-print, pre-internet text in bulk: the profile of a training-data pipeline, not a collector’s wishlist. Booksellers noticed the same puzzle — some orders include public-domain classics like Gulliver’s Travels that are free online, suggesting automated buying scripts that don’t always discriminate.

Why Books? The “Unpolluted” Data Premium

Why would a lab pay “top whack” for a 1983 naval magazine? Two words: data quality. Models are trained on vast amounts of text, and the pool of high-quality human writing is finite. Worse, everything published online since late 2022 is increasingly likely to contain AI-generated content — and training new models on model-written text risks model collapse, the degradation loop where synthetic output crowds out genuine human signal.

Pre-2022 printed books solve both problems at once. They are human-written by definition, professionally edited, frequently absent from the open web, and legally cleaner to acquire than scraping — buying a physical copy through normal channels looks more like licensing than the mass scraping that has dragged AI companies into copyright court. The result is a strange new commodity market where your grandparents’ bookshelves are a premium data mine.

Hitting the Data Wall

The used-book gold rush is one front in a broader scramble: labs repricing APIs as demand strains supply, publishers cutting licensing deals, and startups building distillation and synthetic-data pipelines to squeeze more from less. Frontier labs now treat proprietary data — books, archives, licenses, human feedback — as a strategic moat, much like chips and power. Data provenance is becoming part of a product’s spec sheet, not a footnote.

What This Means for AI Tool Buyers in 2026

Want to compare AI tools that publish their data practices? Browse the vetted catalog on aitrove.ai.

Frequently Asked Questions

Are AI companies really buying secondhand books?

It’s the leading theory, not confirmed fact. The Guardian reported on August 15, 2026 that booksellers in the UK, Ireland, the US, Australia, and Europe are receiving large, thematically random, machine-like bulk orders from opaque buyers. What is confirmed: Anthropic has spent tens of millions of dollars buying books to scan for “data acquisition,” per the Washington Post.

Why do AI companies want books published before 2022?

Text written before ChatGPT’s late-2022 explosion is guaranteed to be 100% human-written and often isn’t available online at all. It’s “unpolluted” training data that avoids model-collapse risk and adds genuinely new signal — which is why some buyers reportedly pay full price without haggling.

Does Anthropic destroy the books it buys?

The Washington Post reported Anthropic’s process involves slicing off spines to scan contents before recycling the books. Anthropic says none of its programs “buy and destroy rare or antiquarian books,” and that sourcing books is a widely used approach across the AI industry.

Is buying books for AI training legal?

It’s less legally exposed than mass web scraping, but training on copyrighted books is still being litigated worldwide — and buying physical copies doesn’t automatically grant training rights, which is why many labs now pursue formal licensing deals with publishers.

What should I ask an AI tool vendor about training data?

Three questions: Where does the training data come from (licensed, public domain, scraped)? Does the vendor indemnify customers against IP claims? And is the vendor transparent about data provenance? Vendors with clean answers face less risk of court-ordered disruptions to your workflow.

Pick AI Tools You Can Trust

Compare 300+ vetted AI tools — coding assistants, writing tools, agents, and more — on aitrove.ai, your trusted AI tools directory.

Explore the Directory →