This week, heavy readers and casual posters alike were aghast at emerging reports of rare books being mulched in bulk as a sacrifice for LLMs. Word traveled that book suppliers were receiving abnormally large orders, believing that AI firms sought cleaner source material while skirting copyright laws. At the center of the controversy was ISBNdb, a book database who not only encouraged the trend but seemed to facilitate it. The optics bit back, as the site now tries to distance itself from the controversy, scrubbing its own posts on the subject.
“We’ve seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised,” writes ISBNdb in an update. “We don’t train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We’ve taken the page down.”
A few days earlier, 404 Media reported about the concerning trend of suppliers suddenly hit with bulk book orders. With schools and libraries lean on resources, the decreased business has left them vulnerable. It was suspected that AI firms were making these new orders, but putting themselves in front of the crosshairs was ISBNdb, a database who seemingly introduced a service to patch AI firms through to depositories.
“The world’s best AI training data is sitting on a shelf,” read the now deleted landing page. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”
The outrage hit a fever pitch this week but news of the practice broke last January, when court documents exposed “Project Panama,” a program within Anthropic to scan then destroy as many books as they can get a hold of. Anthropic settled with authors for $1.5 billion, but the unsealed filings showed that the practice is considered legal, just bad publicity, and the company was willing to break the bank to keep it under wraps. Now the juice is out of the tube, and suspicious rare book orders are under intense scrutiny.
“It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell,” one anonymous seller told 404‘s Samantha Cole. “On the other hand, I donât like the end-use, and I donât like that uncommon books are being pulped.”
Confronting the hell theyâve rendered for themselves, AI firms are struggling to find “clean” source material to train LLMs on. Whether itâs low quality web content, or material already generated by AI that threatens negative feedback loops, these companies are trying to siphon higher quality stuff without raising too many alarms. Earlier this summer A24 announced a partnership with Google to assist training DeepMind, hoping to help their media generators escape the bog of muddy brown, cursed Ghibli slop.
As for the destruction part, itâs unlikely being done in attempts to cover their tracks or some Brianiac-style rare knowledge obsession. The Project Panama documents didnât specify why they trashed the books, but itâs most likely just trying to save a buck. Archiving and scanning books doesnât have to destroy the source material, but being done cheaply and quickly, it likely will. Ripping apart the spines for clearer scans and disposing of the crumpled heap.