Secondhand booksellers in UK and Ireland suspect AI firms behind ‘strange’ bulk orders
Secondhand booksellers in the UK and Ireland, and globally, are receiving unusual bulk orders for diverse books. They suspect AI companies are acquiring these for data acquisition to train large language models, disrupting the traditional market.
Intelligence analysis by Gemini 2.5 Flash

Secondhand booksellers across the UK, Ireland, and other countries are experiencing a surge in 'strange' bulk orders for highly varied titles from opaque buyers. They suspect AI companies are behind these purchases, aiming to scan the books for data to train large language models, a practice previously reported with companies like Anthropic.
Imagine big computer brains that learn by reading lots and lots of books. Now, people who sell old books are getting strange orders for all sorts of books, not just ones on a certain topic. They think these computer brain companies are buying the books to quickly read them all and learn from them, like a super-fast student, before the books are recycled.
Analysis
Stuart Manley's Observations
Stuart Manley, co-owner of Barter Books in Alnwick, Northumberland, highlighted the unusual nature of the bulk orders, which began three months prior. Unlike typical thematic purchases, these orders comprised a "strange combination of books," including an Estonian translation of John le Carré and a specific imprint of Anne Brontë's Agnes Grey. Manley reported selling "hundreds" of such random books to three buyers, totaling approximately £4,000, indicating a significant, non-traditional purchasing pattern.
This trend is not isolated to Barter Books; Jim Shaughnessy of MW Books in Claregalway, Ireland, described similar "sporadic, presumably machine-driven" orders lacking thematic currency, ranging from 18th-century agricultural implements to 1950s racing car biographies. David Gower-Spence of BookLovers of Bath also noted a rush of orders for around 200 books from buyers cited by other sellers, including Canada-based Zoom Books. These accounts collectively paint a picture of a widespread, coordinated, and unconventional buying spree.
Anthropic's Precedent
The speculation among booksellers is significantly bolstered by previous reports concerning AI companies' data acquisition strategies. The Washington Post revealed that Anthropic, a prominent AI company and developer of the Claude chatbot, had spent tens of millions of dollars acquiring books. Their method involved slicing off book spines to facilitate scanning, after which the books were sent for recycling. This practice provides a concrete example of the suspected activity, lending credibility to the booksellers' theories.
Anthropic confirmed that "sourcing books is a widely used approach for training large language models across the AI industry," stating that Claude is trained on a mix of publicly available web data, commercially acquired datasets, and self-generated data. While they clarified that their programs do not target rare or antiquarian books, their admission underscores the industry's reliance on physical books as a crucial data source for developing advanced AI models. This context helps explain the current surge in demand.
The Value of Pre-2022 Books
The specific appeal of secondhand books for AI training was further illuminated by a report from the tech news site 404 Media, which cited a books database, ISBNdb. This database reportedly suggested that books published before 2022 were ideal material for AI companies. The rationale was that their text was "unpolluted" by content potentially generated by chatbots, making them a pristine source of human-authored data. Although ISBNdb later stated the webpage was a "test of market interest" and was taken down, the underlying logic remains pertinent.
AI tools like chatbots require vast and diverse datasets to effectively generate responses and build new models. Secondhand books, many published decades before the widespread emergence of generative AI like ChatGPT and often not available in digital formats, represent a rich, untapped reservoir of such "fresh data." This makes them highly valuable for tech companies seeking to enhance their language models, even if it means disrupting traditional book markets and potentially leading to the pulping of physical copies after scanning.
Key points
- Secondhand booksellers in the UK, Ireland, and globally are receiving unusual bulk orders for diverse titles.
- Buyers are often opaque, paying full price, and orders do not follow typical thematic patterns.
- Booksellers suspect AI companies are acquiring these books for data acquisition to train large language models.
- Anthropic, an AI company, was previously reported to have spent millions on books for scanning and recycling for its Claude chatbot.
- The practice is disrupting the traditional secondhand book trade and raises questions about the future of physical books as data sources.
The increased demand for secondhand books could provide an unexpected revenue boost for booksellers, potentially revitalizing a niche market and creating new opportunities for businesses involved in data acquisition and processing.
The bulk acquisition of diverse books by AI companies could disrupt the traditional secondhand book market, leading to shortages of certain titles for individual readers and collectors, and potentially driving up prices or altering the cultural value of physical books.



