Why AI Giant Anthropic Is Destroying Old Books to Save Its Future
eyesonbrasil
In a quiet industrial facility, a heavy steel blade slices through the sewn spine of a century-old hardcover book. Pages that survived wars, economic collapses, and decades of silent storage in forgotten attics fall away, suddenly unbound. Within seconds, a high-speed sheet feeder swallows them whole, converting thousands of words of forgotten human thought into high-resolution digital tokens.
The book itself—the physical artifact—is discarded as trash.
This is destructive scanning, a process that has placed Anthropic, one of the world’s leading artificial intelligence safety startups and the creator of the Claude AI model, square in the crossfire of ethics, preservation, and technology.
To bibliophiles, archivists, and critics, this practice feels like a high-tech edition of Fahrenheit 451—a corporate sacrilege where history is guillotined in the name of progress. But to AI developers, this mechanical destruction is not an act of malice; it is a desperate attempt to solve an existential crisis threatening the entire AI industry: the synthetic paradox.
The Poisoned Digital Well
To understand why a multi-billion-dollar AI company is systematically destroying physical books, one must first understand the catastrophic state of the internet.
For the past decade, Large Language Models (LLMs) were built by scraping the public web. They consumed Wikipedia, Reddit, news outlets, digitised forums, and blogs. This vast ocean of digital human banter allowed models like ChatGPT and Claude to learn syntax, nuance, and logic.
However, that ocean has been poisoned.
Since the public explosion of generative AI in late 2022, the internet has been flooded with a tidal wave of AI-generated content. Search engines are choked with low-effort, AI-written SEO articles. Social media is dominated by bots replying to bots. News sites are churning out automated summaries of automated reports.
AI models are now trapped in a dangerous feedback loop. As they scrape the modern web for fresh training data, they are increasingly eating their own output.
In computer science, this phenomenon is known as “Model Collapse” or AI autophagy (self-digestion). When an AI is trained on data produced by another AI, it slowly degrades. The subtle nuances, rare vocabulary, and deep structural logic of human language begin to erode. Over generations, the model turns into a digital imbecile—hallucinating wildly, repeating bland cliches, and losing the ability to reason complex thoughts. It is the algorithmic equivalent of inbreeding.
The Rush for Uncontaminated Human Thought
To prevent Model Collapse, tech companies desperately need “pure” data—text written strictly by humans, untouched by the synthetic echo chamber of the post-2022 internet.
And where does that pure, uncontaminated human thought exist? In physical books published before the advent of the consumer web.
Books represent the crown jewel of training data. Unlike a casual tweet or a chaotic Reddit thread, a book contains sustained, highly structured, and heavily edited human thought. It offers rare vocabulary, complex narrative arcs, and deep domain expertise.
While millions of modern books are available digitally, copyright laws and paywalls make them a legal minefield. Out-of-print, public domain, and obscure physical books, however, offer a treasure trove of “virgin” human language.
The problem? Digitisations using non-destructive scanners—where a human or robotic arm gently turns every page—is painfully slow, labor-intensive, and wildly expensive.
Destructive scanning, on the other hand, is fast. By slicing off the spine of a book, high-speed document feeders can digitize hundreds of pages a minute. For AI companies racing to train their next-generation models before their competitor does, speed is the only metric that matters.
The Faustian Bargain: Burning the Past to Build the Future
This urgency has led to a profound cultural and ethical clash.
Critics argue that Anthropic’s approach treats human heritage as disposable fuel for corporate algorithms. When a rare or out-of-print book is un-bound and destroyed, a physical link to our history is severed forever. If the digitized file is corrupted, lost, or trapped behind proprietary company walls, the original knowledge is effectively erased from the physical world.
“We are burning the furniture to keep the digital furnace warm,” critics argue, pointing out the tragic irony of the situation: AI companies are destroying physical artifacts of human intelligence to teach machines how to simulate human intelligence.
Furthermore, it raises uncomfortable questions about cultural ownership. Does a private tech company have the moral right to consume physical human history, liquidate its form, and monetise the resulting “intelligence” without returning anything to the public domain?
The Irony of the Artificial Future
Anthropic has built its brand on being the “ethical” and “safety-conscious” alternative in the AI race. Yet, the destruction of physical books highlights the raw, predatory pragmatism underlying the AI revolution.
We have reached a bizarre inflection point in human history. By filling our digital world with artificial, cheap, automated text, we have made our online environment unlivable for the very machines we built to inhabit it. Now, to keep those machines alive, we are forced to reach back into the physical past, raiding libraries and slicing open old volumes to extract the last remaining drops of authentic human thought.
The ultimate irony remains to be seen. If AI models successfully devour all the world’s physical books, what happens next? Once the last spine is cut, and the physical past is fully ingested, the machines will be left alone again—staring at an internet of their own making, starving for a human voice.










