Having contributed to the growing shortage of memory and storage, AI companies seemingly have a new target in their sights: humanity's literary history. A recent investigative report from 404 Media reveals that these companies are reportedly purchasing millions of secondhand books through intermediaries to source high-quality training data for their AI models, avoiding public backlash.
AI relies on vast amounts of data to advance, but not just any data. It has to be high-quality data. The problem is that mediocre AI-generated content, commonly referred to as "AI slop," has proliferated across the Internet. This type of content contaminates the data pool and is counterproductive for AI to train on. As a result, leading AI companies have turned to human-authored sources for knowledge, specifically print sources that predate 2022 and are more likely to contain original, uncontaminated content.
There is precedent for AI companies turning to physical books for training AI. For instance, Anthropic, one of the leading AI companies involved in a lawsuit, reportedly invested millions of dollars in extracting information from countless printed books to build its Claude AI models and then destroying them. The company bought books from Better World Books. Although the court decision affirmed that using books for AI training falls under fair use in copyright law, Anthropic faced a staggering $1.5 billion fine for maintaining a repository of seven million pirated books that infringed the copyrights of authors and publishers. Similarly, a coalition of publishers recently filed a lawsuit against Google, accusing the tech giant of allegedly and illegally using millions of copyrighted books to develop its Gemini AI models.
ISBNdb, an online database that reportedly has over 111 million cataloged books, has been a long-favorite platform for booksellers, libraries, and distributors to sell books. With the explosion of the AI industry, ISBNdb has pivoted its business to offer specialized services to bulk-purchase books for AI companies. According to 404 Media, the orders range from 1,000 copies to as many as one million books in a single transaction.
One professional bookseller, who wanted to remain anonymous, purportedly spoke to 404 Media about the unprecedented surge in book sales, which began in April of this year. The seller previously moved around 20 books in a good week, but in recent months, weekly sales have skyrocketed to several hundred books. It represents a fivefold increase over the normal volume. Other booksellers on platforms such as Alibris and Biblio have reported similar spikes in bulk purchases.
While there is no concrete proof that ISBNdb or some other AI company is making the purchase, there are some red flags. Notably, the large-scale purchases only included books with an International Standard Book Number (ISBN), the unique 13-digit code used globally to identify books. There were no patterns in terms of subject, genre, or author. It also appeared that the purchasers disregarded the pricing for the books and snapped up titles at any cost, even if they were overpriced.
During the Anthropic lawsuit, Tom Harvey, who previously participated in the creation of Google Books before leading Anthropic's "Project Panama" digitalization project, confirmed that the AI firm hired several document scanning companies. Datamation Information Services, which offers high-volume, non-destructive, and destructive book scanning services, was one of them. The former method employs different tools, like overhead scanners, flatbed scanners, or V-shaped imaging systems. The latter method, on the other hand, would have personnel gut the books and feed the individual pages into a high-speed industrial scanner. Logically, AI companies opt for the destructive route since it is more efficient and lower-cost. The result is the destruction of millions of books.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
Obviously, printed books represent a treasure trove of information for AI. However, many debate the ethics of removing books from circulation since it is uncertain whether AI companies filter the rare or even out-of-print books from the common titles during digitalization. The other major issue is that scanned books go directly into a private database to train AI, which the general public does not have access to. True, we will have smarter AI, but at the cost of the information not being available to future generations.
Zhiye Liu is a news editor, memory reviewer, and SSD tester at Tom’s Hardware. Although he loves everything that’s hardware, he has a soft spot for CPUs, GPUs, and RAM.