Crypto VC News – Crypto Press Release Distribution & Guest Posting Site

collapse
Home / Daily News Analysis / AI Book Burning? Companies Are Destroying Millions of Books to Feed Chatbots

AI Book Burning? Companies Are Destroying Millions of Books to Feed Chatbots

Jul 30, 2026  Twila Rosenbaum 6 views
AI Book Burning? Companies Are Destroying Millions of Books to Feed Chatbots

In a development that echoes the dystopian novel "Fahrenheit 451," some artificial intelligence companies are not just reading books—they are systematically destroying them. As developers race to build more powerful AI models, a shadowy industry has emerged to supply them with millions of physical books that are ripped apart, scanned into training datasets, and then discarded. The practice, first reported by 404 Media, has ignited a firestorm of criticism from authors, publishers, and book lovers who fear that rare and out-of-print works are vanishing forever.

The Mechanics of Book Destruction

AI companies, eager to train their chatbots on high-quality human-authored text, have turned to intermediaries who specialize in bulk book sourcing. These intermediaries advertise their ability to locate hundreds of thousands of titles while promising confidentiality for AI clients. The anonymity reflects the sensitivity surrounding the practice: companies do not want to be publicly associated with destroying books. The books are typically purchased in bulk from used-book sellers, libraries, and estate sales. Once acquired, they are physically cut apart—often by guillotine-like machines—and each page is scanned at high resolution. The digital files are then fed into machine learning models, while the original books are pulped or discarded. Critics argue that this process is wasteful and destructive, especially for rare editions that cannot be replaced.

Why Now? The Race for Human-Written Data

The surge in demand for physical books is driven by a fear among AI developers that the internet is becoming polluted with AI-generated text—often called "AI slop." Books published before the rise of generative AI in 2023 are especially valuable because they represent a pure source of human knowledge, untainted by machine output. AI models trained on such data are believed to produce more accurate, coherent, and original responses. This has led to a scramble to acquire as many pre-2023 books as possible, with some companies building vast libraries of scanned texts. The practice mirrors Anthropic's "Project Panama," which digitized millions of books through destructive scanning. Alphabet's Google has also been involved in large-scale book scanning for years, though its approach has been less secretive.

Impact on the Used-Book Market

The AI-driven buying spree is reshaping the used-book market in ways both profitable and troubling. One unnamed bookseller told 404 Media that weekly sales climbed from roughly 20 books to several hundred after AI buyers entered the market. While this has been a financial boon for sellers with large inventories of obscure or foreign-language titles, there is a growing unease. The same bookseller admitted, "I don't like the end-use, and I don't like that uncommon books are being pulped." Rare-book dealers report that demand for out-of-print titles has surged, driving up prices and making it harder for collectors and libraries to acquire them. Some fear that the destruction of these books represents a cultural loss—once a book is scanned and discarded, its physical presence is gone, and if the digital copy is lost or corrupted, the knowledge it contained may be gone forever.

Legal Battles Over Fair Use and Copyright

The legal landscape surrounding destructive scanning is complex and evolving. In the copyright lawsuit Bartz v. Anthropic PBC, a federal judge in San Francisco ruled that scanning legally purchased physical books into digital copies, even when the originals were destroyed, constituted transformative fair use. This ruling, along with similar decisions in cases involving OpenAI and Meta, suggests that courts are sympathetic to the argument that AI training constitutes fair use, as it transforms the text into a functional tool rather than a mere reproduction. However, a separate case in the same district recently approved a $1.5 billion copyright settlement requiring Anthropic to pay thousands of authors about $3,000 per book for using pirated copies of their works to train the Claude model. This highlights the tension between fair use and the protection of copyright, especially when companies rely on illegally obtained datasets.

Ethical Concerns and Public Backlash

The practice of destroying books has drawn sharp criticism from authors, publishers, and even some tech executives. Elon Musk, who leads Tesla and SpaceX, has spoken out against the practice, urging companies to preserve rare books in libraries and scan them without damaging the originals. "I've asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning," Musk wrote on social media. Other critics argue that the destruction of books is a symbol of the tech industry's disregard for cultural heritage. They point out that many of the books being destroyed are from marginalized or niche genres that may not have digital backups. The irony is not lost on observers: AI companies, which tout their ability to preserve and democratize knowledge, are actively erasing physical records of that knowledge.

Broader Implications for Knowledge Preservation

The trend of book destruction for AI training raises profound questions about how knowledge will be preserved in the digital age. While digital copies are convenient, they are also fragile. Hard drives fail, file formats become obsolete, and cloud services shut down. Physical books, on the other hand, have survived for centuries with minimal maintenance. The wholesale destruction of physical copies in favor of digital scans could create a future where entire libraries exist only as data in the servers of private companies. If those companies go bankrupt or change their priorities, the knowledge could be lost. This has led some scholars to call for a moratorium on destructive scanning until clear rules are established for preservation and public access. Meanwhile, libraries and archives are racing to digitize their collections using nondestructive methods, but they face funding and staffing shortages.

Industry Responses and Future Outlook

Some AI companies have defended the practice, arguing that it is necessary to build high-quality models that can serve the public good. They point to the transformative potential of AI in education, healthcare, and scientific research. Others have promised to donate digital copies to libraries or to preserve the scanned books in secure archives. However, critics remain skeptical, noting that digital copies are often locked behind paywalls or subject to restrictive licenses. The debate is likely to intensify as more companies adopt the practice and as courts refine the boundaries of fair use. In the meantime, used-book sellers are caught in a moral dilemma: profiting from a business they suspect is harmful to cultural heritage. The future of book destruction for AI training may depend on the outcome of ongoing litigation, public pressure, and the development of alternative data sources that do not require the physical destruction of books.

The Race Against AI Slop

One of the driving forces behind the book-buying frenzy is the fear that the internet is being flooded with AI-generated content, which degrades the quality of training data for future AI models. This phenomenon, sometimes called "model collapse," occurs when AI models are trained on AI-generated text, leading to a loss of diversity and accuracy. To avoid this, developers seek out datasets that are guaranteed to be human-authored. Physical books, with their editorial oversight and historical pedigree, are seen as the gold standard. However, the rush to acquire them has created perverse incentives: the more books that are destroyed, the fewer remain for future generations. This paradox has led to calls for a more sustainable approach, such as licensing digital collections from publishers or partnering with libraries that can provide access without destruction.

Historical Parallels and Cultural Memory

The destruction of books for AI training draws unsettling parallels to historic book burnings. While the motives are different—profit and progress rather than censorship—the outcome is similar: knowledge is being physically erased. In her novel "Fahrenheit 451," Ray Bradbury imagined a world where firemen burn books to suppress dissent. Today, AI companies are burning books in the name of progress, but the result may be the same: the loss of physical artifacts that embody our collective memory. Critics argue that by prioritizing speed and cost over preservation, the tech industry is repeating the mistakes of past civilizations that lost their libraries to conquest or neglect. The difference is that today's destruction is deliberate and systematic, driven by a hunger for data that seems insatiable.

What Can Be Done?

Several solutions have been proposed to mitigate the harm caused by destructive scanning. Some advocates suggest that AI companies should be required to deposit digital copies of scanned books in public archives, similar to copyright deposit requirements for libraries. Others call for a voluntary code of conduct that prohibits the destruction of rare or unique items. There is also hope that technological advances, such as high-speed non-destructive scanners, could eliminate the need to cut books apart. However, these devices are expensive and slower, making them less attractive to companies operating on tight deadlines. In the meantime, book lovers and scholars are urging the public to be vigilant and to report suspicious bulk purchases to preservation organizations. The debate over book destruction is far from over, and its outcome will shape the future of both AI training and cultural heritage preservation for decades to come.


Source:Decrypt News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy