The pitch from data broker ISBNdb is blunt: the world’s best AI training data is sitting on a shelf. The catch is that getting at it means slicing the spine off millions of books, then never admitting which lab paid for the job.
AI has a pollution problem of its own making. So much of the web is now machine-written that models risk feeding on their own exhaust. One company thinks the antidote is sitting on a shelf: old, printed books, published before the chatbots arrived.
404 Media reported that ISBNdb now sources physical books in bulk for AI labs to scan into training data. The firm calls itself the world’s largest book database. “The world’s best AI training data is sitting on a shelf,” its site says. Books, it adds, are “dense, edited, authoritative.”
Why old books
The value is in the date. ISBNdb argues that books printed before 2022 predate the large language model era, so they cannot contain AI-generated text. That matters because of model collapse. It names the documented decline that sets in when models train on the synthetic output of earlier models. Each generation ends up a little worse than the last.
The 💜 of EU tech
The latest rumblings from the EU tech scene, a story from our wise ol' founder Boris, and some questionable AI art. It's free, every week, in your inbox. Sign up now!
The open web offers no such guarantee. A fast-growing share of online text is now machine-made, part of the same slop flood the labs helped create. A pre-2022 print run, by contrast, is a fixed, human-authored record that nobody can quietly rewrite.
The poisoning arms race
There is a second reason labs want clean paper. Authors have started to fight back with data poisoning. They borrow tricks from tools like Nightshade, lacing text with characters that a person reads normally but a model cannot. ISBNdb’s own blog cites Anthropic research suggesting that as few as 250 to 500 crafted documents can plant a backdoor in a corpus of trillions of tokens.
Pre-2022 books, written before any of these tools existed, sidestep the problem. ISBNdb sells that as provenance: buy the paper, keep the receipts, and your legal team holds a clean chain of custody.
The optics problem
There is a catch the company states plainly. Scanning at scale usually destroys the book. Workers slice off the spine so the loose pages feed through a machine, which is faster and cheaper than the careful alternative. So ISBNdb offers its clients secrecy. “Strict NDA on every engagement,” its site says. Buyers’ names are “never disclosed.”
The reason sits in its own marketing copy. “The optics problem is real,” the site reads. “‘AI company destroys two million books’ is not a headline that generates sympathy.” One suggested workaround is pure spin: reword the deed as digitally preserving the books.
A familiar fight
All of this runs straight into a copyright battle already raging. A US judge approved Anthropic’s $1.5bn settlement over pirated books, and ruled that training on purchased, scanned books counts as fair use. Part of his reasoning was that the copying destroyed each print original, so one legal copy simply replaced another. Buying and shredding paper, in other words, is the route the courts bless.
So the industry that promised to digitise human knowledge now pays to buy it up and pulp it, one lorry-load at a time. Ingram, the largest book distributor in the US, has already warned publishers and offered them a way to opt out. The scarcest thing in AI, it turns out, is a sentence that no machine ever touched.