Book Database Company Scrubs AI Training Service From Website After Backlash Over ‘Destruction Narrative’

In a striking example of corporate damage control, ISBNdb, a third-party book database company, has quietly removed portions of its website that advertised services to source and scan physical books for artificial intelligence training purposes. The company had previously marketed itself as a “streamlined partner for sourcing printed books in bulk, tailored to your LLM training needs, delivered at the scale AI demands.” Now, facing intense scrutiny over the practice of destroying books after scanning them, ISBNdb claims the advertised service never actually existed, dismissing it as merely “a test of market interest.”

The controversy highlights growing concerns about how AI companies acquire the massive amounts of text data needed to train large language models. While tech giants have faced lawsuits over scraping copyrighted material from the internet, the physical book pipeline represents a different and arguably more troubling dimension of the data-harvesting ecosystem. The notion of physically destroying books—even after digitizing their contents—evokes uncomfortable historical parallels and has sparked outrage among authors, publishers, and book lovers worldwide.

The Hidden Supply Chain Behind AI Training Data

The original investigation by 404 Media revealed a shadowy supply chain connecting AI companies to bulk book buyers who purchase large quantities of printed materials, scan them for text extraction, and then dispose of the physical copies. This process allows AI developers to acquire training data while maintaining plausible deniability about the destruction of books. Third-party intermediaries like ISBNdb positioned themselves as crucial links in this chain, offering to handle the logistics of sourcing millions of books at scale.

Before removing the content, ISBNdb’s website included a blog post titled “Reframing the Destruction Narrative” that attempted to justify the practice with philosophical arguments. The company wrote: “The book is not destroyed. Its value has migrated. The paper returns to the material cycle; the knowledge enters the intellectual one.” This framing drew sharp criticism from literary advocates who argue that reducing a book to extractable data fundamentally misunderstands the nature of reading and knowledge acquisition. The idea that an AI chatbot’s summary carries equivalent value to the experience of reading a complete work strikes many as a profound category error.

Suspicious Buying Patterns Alert Booksellers

While ISBNdb now claims its service was never operational, evidence suggests that someone is actively purchasing books at unusual volumes. According to reporting by Guardian Australia, secondhand booksellers across the globe have noticed dramatic shifts in their online sales patterns. Multiple sellers have described receiving “waves of orders that did not fit with usual customer patterns,” suggesting systematic acquisition rather than typical consumer behavior. These purchases often target older, obscure titles—exactly the kind of out-of-print works that might fill gaps in AI training datasets.

The book industry has historically operated on thin margins, and many small booksellers initially welcomed the uptick in sales without questioning the buyers’ motivations. However, as awareness has grown about AI data-harvesting practices, some sellers have begun refusing orders from suspected bulk buyers. The secondhand book market, worth billions annually, has become an unexpected battleground in the broader conflict over AI training data and intellectual property rights.

Legal and Ethical Questions Remain Unresolved

The practice of scanning books for AI training raises complex legal questions that courts are only beginning to address. Several major lawsuits are currently working through the American legal system, with authors and publishers arguing that using copyrighted works to train AI models constitutes infringement regardless of whether the original text is reproduced. AI companies counter that their use falls under fair use provisions, transforming the source material into something fundamentally new. The outcome of these cases could reshape the entire AI industry’s approach to training data.

Meanwhile, the deletion of ISBNdb’s service pages—while the content remains preserved on the Internet Archive—demonstrates how quickly companies can attempt to distance themselves from controversial practices once public attention arrives. The claim that an extensively documented service was merely a “test of market interest” that was never actually offered strains credulity, particularly given the detailed marketing materials and blog posts the company created. Whether other companies in the AI data supply chain will face similar scrutiny remains to be seen, but the growing awareness of these practices suggests the industry’s reliance on third-party obfuscation may be reaching its limits.

Expert Opinion: This incident reveals the uncomfortable reality that AI development depends on vast quantities of human-created content, and companies are increasingly willing to exploit gray areas in intellectual property law to acquire it. As regulatory frameworks catch up with technological capabilities, we can expect more companies to quietly abandon controversial data-harvesting practices while denying they ever existed. The book scanning controversy may ultimately accelerate legislative action on AI training data, potentially requiring explicit licensing arrangements that could significantly increase development costs for large language models.