AI Labs Are Buying and Shredding Millions of Rare Books

⚡ TL;DR
AI companies including Anthropic have been buying millions of secondhand and rare books, cutting off their bindings, scanning each page, and discarding the paper originals to build massive training datasets. The practice, revealed through court filings and industry reporting, is drawing criticism from authors, librarians and preservationists even as it helped some labs argue they acquired content lawfully. It follows a landmark $1.5 billion copyright settlement Anthropic reached with authors last year.

Artificial intelligence companies building chatbots like Claude and ChatGPT have been quietly buying up millions of secondhand and rare books, scanning every page, and then destroying the physical copies, according to court records and industry reporting. The practice has emerged as a central and controversial method AI labs use to legally acquire the vast troves of text needed to train large language models.

AI book scanning

The disclosures trace back largely to litigation against Anthropic, the maker of Claude, which admitted in federal court filings that it purchased used print books in bulk, removed their bindings with industrial cutters, scanned each page with high-speed equipment, and then discarded the paper remains once the digital text was captured. The company told the court it pursued this approach specifically because buying a physical copy of a book is legal, and argued that scanning a lawfully purchased book for internal research purposes should qualify as fair use under copyright law.

A Response to Piracy Allegations

The scanning operation came to light as part of Bartz v. Anthropic, a lawsuit brought by a group of authors who accused the company of training its models on pirated copies of their work downloaded from shadow libraries. A federal judge in California found last year that Anthropic’s use of legally purchased, scanned books for AI training could qualify as fair use, but ruled that its earlier use of pirated digital copies was not protected. Anthropic ultimately agreed to pay $1.5 billion to settle claims from authors and publishers over the pirated material, one of the largest copyright settlements in US history.

In the aftermath, Anthropic and other labs leaned harder into the book-buying-and-scanning model as a way to build clean, defensible datasets. Reporting on the practice describes warehouses where workers process books much like a factory line: spines are sliced off, pages are fed through scanners, and the resulting paper is thrown away or pulped, since the physical object has already served its purpose once its contents are digitized.

Rare and Out-of-Print Titles Targeted

What has alarmed librarians, historians and book collectors is that the buying sprees have not been limited to mass-market paperbacks. Out-of-print, rare and secondhand editions have reportedly been swept up as well, some of which exist in limited quantities. Once scanned, those physical copies are gone for good, even though a digital version now exists inside a private company’s training pipeline rather than in a public archive or library catalog.

Preservationists argue this amounts to a one-way transfer of cultural material from the public sphere into proprietary AI systems. A scanned book absorbed into a model’s training data isn’t searchable, browsable or citable the way a library holding would be; its content becomes diffused into an AI’s statistical patterns, useful for generating text but not preserved in any traditional archival sense.

Critics say the process treats books as disposable raw material rather than as objects with independent cultural or historical value, even when a physical copy being destroyed is the only surviving edition in circulation.

Why Buying Beats Licensing

Industry observers note that AI companies have largely avoided negotiating broad licensing deals with publishers for older or out-of-print titles, since many such books have no clear rights holder actively managing digital licensing. Buying a used copy outright sidesteps that complexity: there is no need to track down an author’s estate or a defunct publisher, and the purchase itself establishes a paper trail showing the company paid for the content it later digitized.

That legal logic has pushed several AI labs to treat used bookstores, estate sales and online marketplaces as sourcing pipelines for training data, according to industry reporting. The scale involved is significant, with millions of volumes reportedly processed, spanning genres from literary fiction to technical manuals to decades-old nonfiction.

Publishers and Authors Push Back

Author groups and publishing trade associations have argued that even a lawfully purchased book shouldn’t automatically confer rights to use its contents for training a commercial AI product, since the original purchase price reflects the value of one reader accessing the text, not a company extracting patterns from it to power a product used by millions. That tension is expected to keep fueling copyright litigation even after the Anthropic settlement, as authors in other lawsuits continue pressing labs including OpenAI and Google over how their training datasets were assembled.

The scanning-and-shredding practice sits alongside other flashpoints in the AI industry’s rapid growth, from disputes over intellectual property to broader questions about how AI companies operate at scale. Anthropic itself has separately accused a rival firm of improperly obtaining its technology, as detailed in a recent dispute with Alibaba over Claude’s AI capabilities.

What Happens Next

For now, no federal law explicitly regulates whether a company can destroy a physical book after digitizing its contents for AI training, leaving the practice in a legal gray area shaped mostly by ongoing court cases. Publishers, libraries and author advocacy groups are expected to keep pressing for clearer rules, while AI labs continue to argue that scanning legally acquired books is both efficient and lawful. Until legislation or further court rulings settle the question, the quiet churn of used bookstores into AI training pipelines appears likely to continue.

0
Show Comments (0) Hide Comments (0)
0 0 votes
Article Rating
Subscribe
Notify of
guest
0 Comments
Oldest
Newest Most Voted
0
Would love your thoughts, please comment.x
()
x