As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing.
“The world's best AI training data is sitting on a shelf,” ISBNdb, a company that produces what it claims is “the world’s largest book database,” and that offers high-volume book acquisition services for AI companies, says on its site. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.”
In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don’t include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in “model collapse,” a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.
“Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] “Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools.”
ISBN stands for International Standard Book Number, the numerical commercial book identifier and barcode on the back of most books. For years, ISBNdb helped book sellers, libraries, and distributors manage their inventory and find and sell books, but the generative AI boom has made it valuable to AI companies. In addition to selling access to book metadata, ISBNdb now helps AI labs source bulk printed book purchases of between 1,000 to 1 million books per order. ISBNdb’s data makes it easier for AI companies to methodically acquire, scan, and turn printed books into training data while avoiding duplication.
AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process. The Washington Post article found that Anthropic was buying books from one company called Better World Books, one of several marketplaces where libraries, retailers, and individuals can sell their books. Google was recently sued by book publishers for similarly training Google Gemini on copyrighted books.
ISBNdb advertises that it can keep the identity of AI companies secret.
“Strict NDA [non-disclosure agreement] on every engagement,” ISBNdb’s site says. “Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed.”
ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process.
“The optics problem is real,” ISBNdb’s site says. “‘AI company destroys two million books’ is not a headline that generates sympathy.”
One professional bookseller who specializes in selling foreign language books on these marketplaces told me that, starting in April, he and other booksellers noticed a historic spike in sales. This bookseller asked to remain anonymous so he can continue to do business on these platforms.
“I personally have mixed feelings about all of this,” the bookseller, who suspects he’s sold hundreds of books to AI companies for training data, told me. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”
This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
The seller told me that, normally, on a good week, he’d sell about 20 books. Since April, he has regularly sold hundreds of books a week. While the seller didn’t have clear evidence that the purchases were being made by AI companies, the purchases made him suspect that they were. First of all, he said, the kind of books he sells are specialized and are usually bought by schools and libraries. Purchases from these organizations have been trending downward because of reduced funding, he said. Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases. I have not seen any evidence that this bookseller’s recent sales were facilitated by ISBNdb or that the client was an Anthropic or another AI company.
“It's not just the quantity, but the weirdness of the orders,” the bookseller told me. “I've had library orders before, and usually they're mostly confined to a single subject or maybe a slightly broader range of subjects. But basically, almost every library in the world has lost their budget. I know all the U.S. college libraries don't buy much anymore. The Australian libraries don't buy much anymore. The type of books [...] there's no rhyme or reason to it. Also, there's a total disregard for the price of the book. I've had some books that sold through this way that were [...] greatly overpriced. That's kind of a tell for AI because they have just so much money.”
“Is it just me, or has there been an uptick in the number of AutoBuy orders since the tail end of last year?” one bookseller wrote on the forums for Alibris, another marketplace for selling books, in February. The AutoBuy function allows a customer to flag books they want to automatically purchase once they become available for sale on Alibris. “Any comment on what is happening? Is an AI going to read every single book? Any insight into how the selections are made? They seem to vary quite a bit in condition, format (hardcover and softcover), price and so on.”
“We have a couple of new bulk buyers that are scooping up trade books so lots of sellers are getting lots of orders,” Mike Feldman, director of client services at Alibris, responded.
One bookseller told me that similarly large orders of books were coming through another marketplace called Biblio. Customers can provide Biblio with a spreadsheet of ISBNs they want to purchase and the company takes it from there.
In June, a publication in the Netherlands talked to several rare booksellers who reported similar large bulk purchases they assumed were coming from AI companies.
It’s hard to say for a fact that the books are being bought for training data and possibly being destroyed by AI companies because ISBNdb and book marketplaces like Biblio and Alibris keep the identity of the buyer hidden. Large bulk purchases of books are first sent to distribution centers where, for example, Alibris checks the quality of the books before sending them off to the client.
Internal Anthropic documents about its plan to scan millions of books, revealed in the copyright lawsuit, don’t make clear why the company wanted to destroy the books in the process. A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both “high volume destructive and non-destructive book scanning” services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.
Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropic’s creation of digital copies of the books was legal specifically because the books were destroyed.
“Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,” Alsup wrote in his ruling. “The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
This, Alsup said, was “clearly transformative” and therefore qualified as fair use under Section 107 of the Copyright Act.
ISBNdb’s site advertises this legal argument to AI companies as well.
“Purchasing paper books at scale from the secondary market does not deprive any creator of income they would otherwise have received,” ISBNdb’s site says. “These are books that have already fully discharged their financial obligation to their creators.”
“Responsible physical sourcing is not book burning,” ISBNdb’s site in a section about why it’s crucial to recycle the destroyed books. “It is the completion of a book’s lifecycle: from tree to knowledge to tree again.”
ISBNdb and Anthropic did not respond to a request for comment.
About the author
Emanuel Maiberg is interested in little known communities and processes that shape technology, troublemakers, and petty beefs. Email him at emanuel@404media.co