A company best known for cataloging books rather than selling them has abruptly reversed course on an offer to broker printed books as training data for AI companies, following press coverage that drew unwanted attention to the practice.
The reversal
On July 30, ISBNdb — a data provider that has spent more than two decades supplying book metadata to bookstores, libraries, and reading apps — removed the landing pages and posts that had promoted “Printed Books Sourcing for Your AI LLMs Dataset Needs.” In a statement posted to its site, the company characterized the offering as “a test of market interest” and said it has never purchased, scanned, or sold books for AI training, nor trained any models of its own. Its business, the company stressed, remains the data about books — not the books themselves.
The takedown came nine days after the outlet 404 Media reported on the arrangement, part of its broader coverage of AI companies’ growing appetite for pre-2022 printed books. Older books are attractive to AI developers because they predate the rise of AI-generated text online, making them free of the synthetic content that can degrade model quality over successive training runs — a phenomenon researchers call model collapse.
What the deleted pages said
Before it disappeared, ISBNdb’s marketing material was notably candid about the mechanics and ethics of the business. The site argued that books bought on the secondary market had already “fully discharged their financial obligation” to the authors who wrote them. It acknowledged, in its own words, that “the optics problem is real” — that a headline along the lines of an AI company destroying millions of books “is not a headline that generates sympathy.” The pages reportedly also raised the idea that authors opposed to AI training could write in ways deliberately meant to undermine the models that consumed their work — a nod to the growing practice of “data poisoning.”
The service, as described before its removal, offered to source anywhere from 1,000 to a million physical books per order, drawing on used bookstores and other catalogs for older, specialized, rare, and out-of-print titles, and promised clients strict non-disclosure agreements shielding their identities and acquisition strategies.
That secrecy is part of why booksellers noticed something was off well before the arrangement became public. Sellers in the U.S. and across Europe — including in the Netherlands, Germany, Switzerland, and Spain — reported unusual spikes in bulk purchases of niche and specialized titles, with one bookseller telling 404 Media that weekly sales had jumped roughly fivefold since April, a pattern more consistent with large-scale data acquisition than ordinary consumer demand.
Why the books get destroyed
The training pipeline behind this demand typically involves destructive scanning: workers slice the binding off a book so its loose pages can be fed through an industrial scanner, after which the original is discarded or recycled. For rare, foreign-language, or low-circulation academic titles, that process can mean the last surviving physical copy of a work ceases to exist outside a corporate training set — a point that has fueled concern among librarians and preservationists about the long-term fate of scarce editions.
The Anthropic backdrop
The episode doesn’t exist in isolation. Anthropic has faced a closely watched copyright case, Bartz v. Anthropic, centered on the company’s acquisition and scanning of millions of print books for model training, with court filings referencing an internal program called Project Panama. In June 2025, U.S. District Judge William Alsup ruled that digitizing legally purchased print books to train large language models qualified as fair use, reasoning that the process replaced physical copies with digital ones rather than creating unauthorized new copies for distribution. Anthropic separately agreed to a $1.5 billion settlement over claims involving pirated digital books. Reporting has also linked some of the underlying book purchases to secondary marketplaces of the kind used by everyday used-book buyers.
The bigger picture
Taken together, the ISBNdb episode is a smaller, more transparent window into a dynamic already playing out at industrial scale elsewhere in the AI industry. As the open web becomes increasingly saturated with AI-generated text, high-quality, verifiably human-authored writing has become a genuinely scarce resource — and physical books, printed before large language models existed, are one of the few remaining sources that can be dated with confidence to before the “slop” era began.
That scarcity is precisely what makes ISBNdb’s reversal notable rather than reassuring. The company says it never actually delivered the service; it doesn’t dispute that demand for it clearly exists, or that the underlying practice — bulk-buying and destructively scanning books to train AI — is already happening elsewhere at a much larger scale. The speed of the walk-back, once the arrangement drew press attention, says as much about how sensitive this territory has become as the original pitch did about how lucrative it looked.
This article incorporates reporting from 404 Media and subsequent coverage of ISBNdb’s marketing materials and their removal.

And staying with all the red meat George Carlin has thrown our way above here, George, who reminds us of…
TWINKLE, TWINKLE, Ms. Little Star! How we out here in the world wonder who you really are, whether Connie West…
Paul, Paul, Paul. Let us not get ahead of ourselves here! And please address me with the more formal Ms.…
I once tried to find out how many separate, individual polling places there are in America on election day, and…
As to the overall problem specific to downtown Cape Charles in the now-gentrified section, as opposed to the overall problem…