Amazon, which started off selling books, is destroying rare texts to train AI
Amazon is destroying rare books to train its AI models. The company buys rare texts, cuts off their spines, and scans them for AI training.
Intelligence analysis by Llama

Amazon is using rare books to train its AI models, cutting off their spines and scanning them for training data. This is a new source of coveted training data for the company.
Imagine you're trying to teach a kid to read. You'd want to use real books, not fake ones. Amazon is doing something similar with its AI, but instead of using real books, it's using fake ones to teach its AI to read. This is a problem because the AI might get confused and start making mistakes.
Analysis
Amazon's AI Training Data Conundrum
Amazon's decision to destroy rare books to train its AI models has sparked controversy. The company's need for unfathomably large amounts of text to train its LLMs has led it to seek out new sources of training data. Rare books, especially those that are out of print or impossible to find on the internet, offer a valuable source of training data. However, this raises questions about the value and preservation of these texts.
The Risks of Model Collapse
When LLMs train on AI-generated text, they risk
Key points
- Amazon is destroying rare books to train its AI models.
- The company is using rare texts as a new source of training data.
- This raises questions about the value and preservation of rare books.
If Amazon's use of rare books for training data leads to breakthroughs in AI, it could have a positive impact on various industries, such as healthcare and finance.
The destruction of rare books for AI training data could lead to a loss of cultural heritage and historical knowledge, which could have negative consequences for future generations.



