Breaking Lab srl
Feeding the Machine at Any Cost
The race to build more capable artificial intelligence has always been, at its core, a race for data. The internet has been scraped, digitized, and indexed to exhaustion, leaving AI developers scrambling to find new sources of high-quality text that their models have not already seen. Amazon, it now emerges, has found one such source — and the method it is using to access it is drawing sharp criticism: the company is reportedly acquiring rare, out-of-print books and destroying them in order to feed their contents into large language model (LLM) training pipelines.
Why Rare Books Are So Valuable for AI
To understand why a tech giant would target physical rare texts, you need to understand the data problem facing modern AI. Today’s leading LLMs have effectively consumed everything that exists in a clean, accessible digital form — news archives, Wikipedia, academic papers, forums, social media. The models have, in a sense, already read the internet. What they have not read is the enormous body of human knowledge that was never digitized: obscure historical volumes, early-edition scientific treatises, regional literature, unpublished manuscripts, and specialized reference works that exist only in a handful of physical copies worldwide.
These texts are extraordinarily valuable precisely because of their scarcity. They contain linguistic patterns, vocabulary, and knowledge structures that are simply absent from standard training corpora. For an AI company trying to push its models beyond the current plateau, gaining exclusive access to this material — even temporarily — represents a genuine competitive advantage.
Destruction as a Business Decision
What makes Amazon’s reported approach particularly troubling is not merely the acquisition of these materials, but the decision to physically destroy them after digitization rather than preserve, donate, or return them to libraries and archives. From a narrow cost-benefit standpoint, the logic is coldly rational: acquiring, shipping, storing, and ultimately disposing of physical inventory is treated as a straightforward operational expense in the pursuit of training data. The books become raw material, no different from any other input in a supply chain.
But rare books are not fungible commodities. Many of these volumes exist in single-digit quantities globally. Some are the last surviving copies of texts that took centuries to produce. Once destroyed, they are gone permanently — not just as physical objects, but as cultural and historical artifacts whose full scholarly value may not even be understood yet. The idea that a corporation can unilaterally decide that a book’s highest purpose is to serve as a one-time data source, after which it is discarded, is a profound departure from how civilizations have historically treated their written heritage.
The Broader Ethical Landscape
Amazon is hardly alone in pushing aggressive boundaries to secure AI training data. The entire industry has faced mounting legal and ethical challenges over the use of copyrighted material, scraped web content, and data obtained without clear consent. Several high-profile lawsuits from authors and publishers are already working their way through courts in the United States and Europe. The destruction of rare books adds a new and arguably more irreversible dimension to these concerns.
Unlike a copyright dispute — which can theoretically be resolved through licensing fees or the removal of infringing content — the physical destruction of a rare text cannot be undone. There is no settlement that restores a lost first edition. There is no licensing deal that reconstitutes a manuscript burned or pulped to reduce operational overhead.
A Question of Priorities
Amazon has not publicly confirmed the details of this practice, and it is unclear how widespread or systematic the destruction has been. But the story points to a structural tension that the AI industry has yet to seriously confront: the assumption that any data source is fair game if the resulting model is sufficiently useful. That logic has so far been applied primarily to digital content, where the costs of aggressive data acquisition are diffuse and often invisible. When it is applied to physical cultural heritage, the consequences become concrete, immediate, and permanent. Regulators, libraries, and the broader public may soon need to decide whether the appetite of AI training pipelines should be allowed to consume objects that belong, in a meaningful sense, to all of humanity.







