Are Rare Books Becoming Raw Material for Artificial Intelligence?
By Wendy Keir | EmpowerAi™
Australian booksellers are raising concerns about what may be happening to some of the books leaving their shelves.
Several antiquarian and second-hand dealers have noticed unusual bulk orders for obscure, out-of-print and apparently unrelated titles. Their concern is that physical books may be purchased, taken apart and scanned so that their contents can be added to the datasets used to train artificial intelligence.
There is no evidence at present that the particular books identified by Australian sellers have been destroyed. The businesses involved deny doing so. The concern has emerged because destructive scanning is no longer a hypothetical practice.
Court documents have already shown that Anthropic, the company behind Claude, bought millions of printed books, removed their bindings, scanned their pages and discarded the physical copies. The digital versions were placed into a permanent research library, parts of which were later used in developing its AI systems.
The legal argument focused mainly on copyright and whether the text could be used for AI training. The latest reaction from booksellers introduces another question.
What is lost when a physical book is treated only as a container for extractable words?
A book contains more than its text
For an ordinary paperback that remains widely available, the destruction of one copy may seem relatively unimportant. Libraries and digitisation projects routinely scan books, and destructive scanning can be a practical way to process large quantities of material efficiently.
Rare and older books are different.
A physical copy may contain handwritten notes, dedications, corrections, ownership marks or evidence of how the book was produced and used. Its paper, binding, illustrations and typography may be historically significant. Even when the underlying text exists elsewhere, the particular object can carry information that disappears once its pages have been separated and scanned.
Australian booksellers interviewed by The Guardian described receiving bulk orders that appeared random or unusually broad. Some were concerned that buyers might not recognise the significance of particular copies before processing them.
That does not prove that the books are destined for AI training. Zoom Books, the Canadian company named in the reporting, says it buys books for resale and reuse rather than digitisation or destruction. Anthropic says it has no relationship with Zoom Books and does not deliberately seek rare or culturally valuable material.
The uncertainty is part of the problem. Sellers often have little visibility over what happens after a large order leaves their shop.
AI companies need material that predates AI
The commercial interest in printed books is understandable from a technical perspective.
The public internet now contains increasing amounts of AI-generated writing. This creates a problem for developers because models trained repeatedly on synthetic material can inherit errors, repetition and increasingly generic patterns.
Books published before the widespread arrival of generative AI offer a large body of human-created language that has generally been through some form of selection, editing and production. Older and out-of-print books may also contain specialist knowledge that is difficult to find online.
That makes physical collections valuable as training material.
The legal position in the United States has also given developers some room to operate. In the Anthropic case, a federal judge ruled that using legally purchased and scanned books to train an AI model was transformative and could qualify as fair use. The court treated the lawfully acquired books differently from millions of unauthorised copies Anthropic had obtained from pirate libraries.
Anthropic agreed to pay $1.5 billion to settle the claims concerning the pirated collection. The settlement did not reverse the finding that training on legally acquired and scanned books could be lawful under US copyright law.
Something can be legally permissible while still raising questions about preservation, consent and cultural value.
Ownership does not settle every ethical question
When someone buys a physical book, they normally have the right to cut it apart, discard it or use it for a private scanning project. Owning that copy does not give them copyright in the text, although it gives them considerable control over the object itself.
At a small scale, this rarely attracts attention.
The position feels different when a technology company purchases and destroys millions of books to create a permanent commercial resource. The individual transactions may be ordinary, but their collective effect can alter the availability of older material and remove copies from circulation.
The concern becomes more serious where books are scarce, privately printed, annotated or absent from institutional collections. A buyer processing material at industrial scale may not notice that one volume carries a history beyond the words on its pages.
This is where the debate moves beyond authors being paid for the use of their work. It also touches librarians, archivists, collectors, historians and the booksellers who often recognise the significance of an unusual copy.
The training dataset gains the text. The wider culture may lose the object and everything about it that the scanning process did not capture.
Provenance is becoming central to the AI and publishing debate
Publishing has spent much of the past two years considering the provenance of new books. Readers increasingly want to know whether a work was written by a person, generated substantially by AI or produced through some combination of the two.
This story brings provenance into the training process itself.
Where did the AI’s knowledge come from? Was the material licensed, pirated, bought or donated? Did the original rights holder know it would be used? Was the physical source preserved? Can the system distinguish between reliable scholarship, an outdated account and an annotated copy containing someone’s private observations?
Large language models tend to flatten those differences. Once material has been converted into training data, its individual history may become difficult to see.
For expert authors, that is significant because intellectual property rarely consists only of sentences. A framework may have emerged from years of client work. A particular term may have a defined meaning. The sequence in which ideas are presented may matter. The author may also have clear limits around what the work should and should not be used to advise upon.
Extracting the words does not necessarily preserve those conditions.
An AI Book Companion™ starts from a different relationship with the book
An AI Book Companion™ also turns a book into something that can be explored through artificial intelligence. The similarity largely ends there.
A properly developed companion begins with the author’s involvement and permission. The author identifies the approved body of work, explains the central frameworks and defines the boundaries within which the companion is allowed to respond.
The system is not simply given a book and told to imitate it.
The build process considers what the author means, where ideas begin and end, how the methodology should be applied and when the companion needs to acknowledge that the source material does not provide an answer.
That difference matters because it preserves provenance.
The reader knows whose work governs the conversation. The author remains responsible for the body of knowledge, while the AI provides a way to navigate and engage with it. The companion does not quietly absorb the book into an anonymous pool of content.
This approach also preserves the original work. The book remains a recognised source with an author, a history and a context. The technology extends access to the thinking without pretending that the source no longer matters.
Books should not become invisible ingredients
AI developers need high-quality data, and books are among the richest sources available. There are legitimate ways to digitise, license and work with published material.
The concern arises when the process treats books as invisible ingredients.
The words are taken, the physical object is discarded and the origin of the knowledge gradually disappears inside a model. Authors may not know that their work was used. Readers cannot see which books shaped a response. Booksellers may not know whether the copies they supplied still exist.
The reports from Australia do not prove that rare collections are currently being destroyed for AI training. They do show that the supply chain has become opaque enough for experienced booksellers to believe that it could be happening.
That should be taken seriously.
As artificial intelligence becomes more involved in publishing, the industry will need to consider the provenance of both its outputs and its inputs. It matters who wrote the new book. It also matters which existing books were used to build the system, how they were obtained and what happened to them afterwards.
Three Key Insights
1. The value of a book is not always contained entirely in its text.
Annotations, bindings, inscriptions and the history of a particular copy can carry cultural information that disappears during destructive scanning.
2. Lawful ownership does not resolve every question surrounding AI training.
A company may be entitled to scan a book it has purchased, while the scale and consequences of destroying physical collections still deserve scrutiny.
3. Provenance needs to cover the material entering an AI system as well as the content it produces.
Authors and readers increasingly need to understand where AI knowledge came from, whether it was used with permission and how its original context has been preserved.
Three Questions I’m Thinking About
1. Should buyers be required to disclose when bulk book purchases are intended for destructive scanning or AI training?
2. Who should decide whether an individual physical copy has cultural or historical value before it is digitised and discarded?
3. Will authors increasingly choose governed, permission-based AI systems as a way of making their work interactive while preserving its ownership and provenance?
Sources
The Guardian — Australian booksellers raise alarm over destruction of rare titles to feed AI
https://www.theguardian.com/technology/2026/aug/02/australian-book-sellers-alarm-destruction-rare-titles-ai-supply-chain
The Washington Post — Inside an AI start-up's plan to scan and dispose of millions of books
https://www.washingtonpost.com/technology/2026/01/27/anthropic-ai-scan-destroy-books/
Reuters — US judge approves Anthropic's $1.5 billion settlement of copyright lawsuit
https://www.reuters.com/world/us-judge-approves-anthropics-15-billion-settlement-copyright-lawsuit-2026-07-20/
EmpowerAi™ | Real AI. Real Impact. Smarter Business Decisions.
AI Book Companion™ helps expert authors transform a completed book into a governed, interactive knowledge asset built around their own ideas, methodology and boundaries.