top of page

Australian Booksellers Warn Rare Titles May Be Destroyed for AI Training

Google News has spotlighted a new conflict in AI development: Australian booksellers fear irreplaceable titles are entering destructive scanning pipelines. The books are cut apart, digitized, and discarded after their text becomes private training data. What looks like an unusual purchasing trend therefore raises a larger question. Who protects the physical record when technology companies value a book mainly as machine-readable text?

The immediate evidence remains incomplete. Australian sellers reportedly cannot identify every buyer or prove that each recently purchased title was destroyed. Yet their suspicions follow documented practices inside Anthropic and new reporting about intermediaries sourcing printed books for AI companies.

That history changes how the latest orders should be interpreted. Anthropic previously bought books in bulk, removed their bindings, scanned their pages, and discarded the originals. A United States court later treated the conversion of lawfully purchased print copies into digital replacements as fair use.

The central conflict is not booksellers against digitization. Libraries have digitized fragile materials for decades while preserving originals and expanding public access. The conflict is preservation against extraction, especially when a physical book disappears into a private dataset that researchers and readers cannot inspect.

Why Australian Booksellers Are Raising the Alarm

The reported sales pattern matters because indiscriminate purchasing can remove obscure books before anyone recognizes their scarcity.

The Guardian reported on August 1 that Australian secondhand booksellers believe they may have entered an AI supply chain. In that suspected pipeline, old books are purchased, scanned, and destroyed after their pages become digital files. The concern includes rare, unusual, and out-of-print editions rather than only common surplus stock.

Melbourne bookseller Tim White, the owner of specialist store Books for Cooks, is among the dealers giving the warning a human face. His business handles culinary books whose value can extend far beyond their printed recipes. An edition can preserve regional terminology, annotations, illustrations, advertisements, production methods, and evidence of how people actually used it.

Those qualities explain the phrase “more than just objects.” A book carries information through its physical construction as well as its words. Binding methods, paper, typography, marginal notes, ownership marks, and inserted documents can help researchers reconstruct a work’s history.

The reported buyers do not always behave like ordinary collectors. Industry accounts describe large orders with little apparent concern for genre, author, or normal retail pricing. An International Standard Book Number, or ISBN, can provide the common selection signal because it lets automated systems identify editions across many catalogs.

That pattern does not prove an AI company ordered a particular Australian book. Bulk buyers can include decorators, resellers, libraries, film productions, exporters, and recycling businesses. A sudden order should not automatically become evidence of destructive scanning.

However, the surrounding market makes the booksellers’ interpretation plausible. Bulk book sourcing has been marketed directly to AI customers seeking printed material. Related reporting describes order sizes ranging from thousands of books to as many as one million copies.

The unusual demand reportedly intensified during 2026. Sellers in several countries described higher sales involving obscure or low-circulation titles. That international pattern makes the Australian concern harder to dismiss as a local misunderstanding.

Google News gave the Australian warning broader visibility, but the aggregation label should not obscure the underlying story. This is a supply-chain investigation, not confirmation that every anonymous order ends inside an AI laboratory. The crucial fact is that a documented destructive process now exists beside a hard-to-audit purchasing market.

That opacity leaves sellers making decisions without meaningful information. A dealer may see a welcome sale after keeping a book for years. The same dealer may reconsider if the buyer plans to eliminate the physical copy and retain its contents inside a private commercial system.

The transaction transfers legal ownership, but it does not settle the cultural question. Booksellers often act as informal custodians for materials that libraries missed. Their inventories can contain the final accessible copy of a local history, technical manual, community cookbook, or short-lived independent publication.

Once such a book leaves the visible market, nobody can reliably establish what happened next. That uncertainty is the first pressure point in this story. It also leads directly to the industrial process that made the warnings credible.

The AI Training Books Pipeline Ends With a Cutter

Destructive scanning turns a slow preservation task into a fast data-extraction process by sacrificing the binding and physical artifact.

A conventional overhead scanner photographs pages while a book remains open. Libraries can use cradles and controlled lighting to reduce stress on fragile bindings. The method protects the original, but page turning and quality checks require time.

Destructive scanning follows a different logic. Workers remove or cut the spine so that individual sheets can pass through an automatic document feeder. Industrial equipment can then process loose pages much faster than a careful operator can photograph a bound volume.

Optical character recognition, or OCR, converts the scanned page images into searchable text. Software can clean obvious errors, identify page structure, and attach catalog metadata. The resulting files can enter a training corpus, which is the collection of material used to shape an AI model’s statistical behavior.

The physical book has little operational value after that conversion. Rebinding millions of volumes would be costly and would undermine the efficiency that justified cutting them. The damaged paper can instead be recycled or discarded.

Anthropic’s book operation showed this workflow at substantial scale. Court records described the company buying used print copies, stripping their bindings, cutting pages, scanning them, and throwing away the paper originals. The project was connected to Claude, Anthropic’s family of language models.

The records emerged through Bartz v. Anthropic, a copyright case brought by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson. The plaintiffs challenged Anthropic’s use of books acquired from both purchased and unauthorized sources.

The court drew a sharp legal distinction between those sources. Judge William Alsup found that using books to train language models was fair use in the circumstances before him. He also found that converting purchased print copies into digital library copies was fair use because Anthropic retained one replacement rather than expanding the number of copies.

The fair-use order did not excuse the acquisition of pirated digital books. Those unauthorized library copies created a separate infringement issue, even when Anthropic later bought physical copies of some titles.

That distinction explains why physical books became strategically useful. A company could legally purchase a copy, replace it with a scan, and dispose of the original. The process created provenance that a downloaded file from a shadow library lacked.

Anthropic later reached a settlement addressing pirated works. A federal judge approved a $1.5 billion agreement covering claims involving nearly half a million books. The copyright settlement did not reverse the court’s favorable treatment of training or purchased-book digitization.

This history matters because it gives other AI developers a recognizable route. Physical acquisition is slower than copying an online collection, but it can reduce one category of legal exposure. Specialized intermediaries can make the route easier by finding, sorting, and delivering books at scale.

The economic incentives also favor older works. Printed books published before the widespread adoption of generative AI offer text created without chatbot assistance. Developers seeking cleaner human-written material may view those collections as an antidote to synthetic content circulating online.

That does not mean older text is automatically better training data. Books contain errors, prejudices, duplicated passages, outdated information, and OCR challenges. Yet their publication dates provide a useful filter against recent machine-generated material.

The irony is difficult to miss. AI systems helped fill the web with inexpensive synthetic text. That contamination increased the appeal of physical writing created before the same systems existed. The industry is now reaching back into print culture to improve models that compete with human authors.

Google News Exposes a Preservation Conflict, Not Just a Copyright Fight

Buying a book can settle ownership of that copy, but it does not preserve the public knowledge, history, or evidence contained in the object.

Copyright dominates debates about AI training because it gives identifiable rights holders a legal claim. Preservation works differently. A book can be out of copyright, lawfully purchased, and still be culturally significant.

The legal analysis in Bartz focused partly on whether Anthropic created an additional library copy. By destroying each purchased original after digitization, the company maintained a one-to-one replacement. That approach helped the digital conversion resemble a change of format rather than straightforward duplication.

For preservation, however, the one-to-one logic creates a loss. A private digital file is not an equivalent public replacement when nobody outside the company can access it. Researchers also lose the physical evidence that OCR cannot capture.

A digital transcript may omit marginalia, damaged pages, foldouts, color distinctions, typography, or handwritten corrections. Even a complete image set cannot preserve paper composition, binding structure, or the relationship between loose materials inserted by previous owners.

This difference is especially important for ephemera and specialized publications. A mass-market novel may exist in thousands of libraries and private collections. A regional cookbook, trade catalog, community history, or discontinued technical manual may survive in only a few known copies.

Rarity is also difficult to measure. Marketplace listings show what is currently offered, not every copy sitting in a home, warehouse, or uncataloged collection. A cheap book can be scarce, while an expensive collectible can be well documented and carefully preserved.

That makes price a poor safeguard. Automated purchasing may capture low-cost titles that attract little collector attention precisely because their importance has not been established. Once destroyed, their rarity becomes even harder to reconstruct.

Australian books face additional exposure because the country’s publishing market is comparatively small. Local editions can differ from British or American versions through covers, illustrations, editing, introductions, recipes, measurements, and regional language. Replacing one with another can erase meaningful context.

The issue also arrives while Australian creative workers are contesting how AI companies use copyrighted material. The Books3 dataset previously brought that concern into public view. It reportedly contained at least 183,000 pirated books, including works by Australian writers.

The Books3 investigation showed how little visibility authors had into training collections. Many learned about inclusion through external analysis rather than notice from developers. Physical sourcing creates a related transparency gap for sellers and cultural institutions.

The two disputes should not be collapsed into one. Books3 concerned unauthorized digital copies. The current bookseller alarm concerns legally purchasable physical copies that may be destroyed after scanning.

Yet both cases reveal the same structural imbalance. AI developers know what enters their datasets, while authors, booksellers, librarians, and readers usually do not. The public cannot assess what was preserved, what was discarded, or what remains available.

Google News can amplify reporting about that imbalance, but search visibility is not an archival response. A news article records concerns at a moment in time. It does not identify every endangered title or recover a destroyed edition.

A practical preservation system would require more than copyright payments. It would need inventory transparency, rarity checks, non-destructive handling rules, and a repository independent from the company training the model. None of those protections automatically follows from purchasing a lawful copy.

The distinction also matters for knowledge workers building private digital collections. Searchable notes and personal archives are most useful when they preserve source context. A personal knowledge base should connect extracted information to its origin instead of treating every document as disposable input.

AI laboratories operate at a different scale, but the principle carries over. Data becomes more trustworthy when users can inspect provenance, verify context, and return to the original source. Destroying that source weakens the evidentiary chain.

Booksellers Face a Market They Cannot See

The new demand gives secondhand sellers revenue while asking them to surrender control over the cultural consequences of each sale.

Booksellers normally expect buyers to read, collect, study, display, resell, or donate a purchase. Destruction has always been possible because ownership includes broad control over a physical copy. The unusual feature here is the potential industrial scale and the buyer’s silence.

A seller cannot protect scarce material without knowing what the purchaser intends. Even a careful dealer may lack the time and records needed to investigate every edition. High-volume online platforms further separate sellers from end users.

This produces an uncomfortable commercial tradeoff. Independent bookshops often work with narrow margins. Refusing a large order can mean rejecting meaningful income based on a suspicion that remains difficult to prove.

Australian book retail already operates under pressure. Separate Guardian reporting found that the country’s number of bookstores fell from 2,879 in 2013 to 1,457 in 2023. It also reported that many independent sellers pay themselves little or no wage.

Those conditions matter because preservation cannot depend entirely on uncompensated individual judgment. Society benefits when sellers identify and retain unusual material, but the seller bears the storage, cataloging, and opportunity costs.

AI buyers can exploit that mismatch without deliberately targeting rare books. A broad order based on publication year, language, or ISBN may sweep up scarce items alongside common ones. The damage can result from indifference rather than a specific plan to eliminate cultural artifacts.

Booksellers could respond with buyer screening, restricted categories, or explicit resale conditions. Each measure has limits. Corporate intermediaries can obscure the final customer, while ordinary ownership rules may make post-sale restrictions difficult to enforce.

They could also flag potentially scarce titles before accepting large orders. Yet reliable scarcity data is fragmented across library catalogs, national bibliographies, specialist databases, and marketplace listings. No single source describes every surviving copy.

Libraries face their own constraints. Acquisition budgets, storage limits, cataloging backlogs, and collection policies prevent them from accepting every unwanted book. Legal-deposit systems preserve many newly published works, but they do not guarantee coverage of every edition or piece of ephemera.

Publishers may hold digital files for recent titles, although those files do not preserve the final manufactured object. Older publishers may have closed, merged, or lost their archives. Small-run books can disappear without leaving a stable production record.

The most credible solution therefore needs shared responsibility. AI companies can fund non-destructive scanning when a book has potential cultural value. Booksellers can pause unusually broad orders. Libraries can help identify priority categories, while governments can strengthen deposit and digitization infrastructure.

A public preservation copy would also reduce the zero-sum character of scanning. The AI company could still use a lawful digital version while researchers gain access under appropriate copyright controls. The original could remain with a library, archive, seller, or collector.

However, such a system conflicts with competitive incentives. A proprietary corpus has greater strategic value when rivals cannot obtain the same material. Publicly depositing scans can weaken the informational advantage that motivated the purchase.

That incentive helps explain the secrecy surrounding some sourcing arrangements. Reporting about ISBNdb described confidentiality as part of the service offered to AI customers. The concern was not merely operational security. Public reaction to destroying books was reportedly treated as an “optics problem.”

The phrase understates the issue. Public criticism is not only emotional discomfort about damaged objects. It reflects a rational fear that private companies can convert shared cultural resources into inaccessible commercial infrastructure.

Still, skepticism is necessary. Australian sellers have raised a warning, not published a forensic chain of custody. No complete list connects their recent orders to particular AI laboratories, scanning vendors, or destroyed copies.

The strongest reporting should preserve that uncertainty. It can establish that destructive AI book scanning has happened and that bulk sourcing services exist. It cannot responsibly state that every suspicious Australian sale followed the same path.

Anthropic’s Precedent Pressures the Rest of the AI Industry

Anthropic demonstrated that legally purchased books can become private AI infrastructure, placing preservation choices outside the normal copyright debate.

Anthropic is the clearest documented example, but the demand for high-quality text extends across the AI sector. Developers need material for initial training, continued pretraining, evaluation, retrieval systems, and specialized models.

OpenAI and Microsoft have pursued a notably different route for one major collection. Their work with Harvard’s Institutional Data Initiative involves nearly one million public-domain books, including material dating back centuries. Those volumes were already digitized, and the physical holdings remain preserved.

Google Books offers another historical comparison. Google developed large-scale scanning systems that photographed bound library books and returned them. The project generated years of litigation, but its physical workflow generally did not require destroying borrowed library collections.

Neither comparison absolves the companies involved from wider copyright or governance questions. They simply show that industrial digitization does not inherently require eliminating the source copy. Destruction is a business and workflow choice.

Non-destructive scanning can be slower, particularly when books resist opening or require careful handling. It also demands more complex equipment and quality control. Those differences become substantial across hundreds of thousands of volumes.

Yet efficiency alone does not determine acceptable practice. Archives routinely impose slower procedures when materials are unique or fragile. The unanswered question is why commercial AI scanning should receive no similar triage requirement.

An AI company could separate ordinary surplus books from editions requiring review. Catalog metadata can provide an initial signal, while library holdings and specialist databases can support further checks. Human experts would still need to resolve uncertain cases.

The difficult category is not the famous first edition. Highly valuable books are less likely to enter an anonymous bulk order because sellers recognize their market value. The vulnerable category includes obscure, inexpensive, and poorly cataloged works whose significance appears only after close study.

This is where the Anthropic precedent creates pressure. Competitors that avoid physical sourcing may possess less varied data. Companies that use careful scanning may face higher costs and longer timelines than companies using industrial cutters.

The result resembles a race in which cultural safeguards become a competitive disadvantage. Voluntary commitments can help, but they remain unstable when model performance and development speed determine investment and market position.

Regulators could change those incentives by requiring preservation assessments for high-volume destructive scanning. Another option would require depositing a preservation copy and detailed metadata with an approved institution.

Copyright law alone is a poor fit for that task. It protects creative rights for a defined period, not every physical trace of publication history. It also gives no general preservation right to a bookseller after a lawful sale.

Cultural-heritage law could address exceptionally important objects, but applying it to ordinary-looking books at scale would be difficult. Export controls and protected-item registers usually cover recognized treasures rather than uncataloged modern publications.

Industry standards may arrive faster than legislation. Scanning vendors could adopt rules that divert signed, annotated, scarce, or institutionally significant material. AI companies could publish acquisition policies and permit independent audits.

Those measures would not eliminate destructive scanning. Many common books exist in abundant copies and have little artifact value. Recycling a damaged surplus copy after creating an accurate scan is different from destroying an edition with uncertain survival.

The present system rarely makes that distinction visible. Buyers can optimize for volume, sellers can optimize for completed sales, and scanning facilities can optimize for throughput. No participant has a clear duty to protect the physical record.

That governance gap is why the Australian warning has traveled beyond specialist bookselling circles. The story is not simply about Anthropic’s past conduct. It asks whether an entire sourcing market will copy the same model without adopting preservation safeguards.

What Google News Readers Should Watch Next

The next phase will be decided by supply-chain transparency, preservation rules, and evidence connecting anonymous orders to actual scanning operations.

The first signal is whether booksellers and marketplaces publish concrete order data. Useful evidence would include purchase dates, title lists, quantities, buyer identities, delivery locations, and repeated patterns across sellers. That information could establish whether the Australian reports reflect coordinated sourcing or unrelated bulk demand.

Privacy and commercial obligations will limit disclosure. Sellers can still provide anonymized information to trade associations, journalists, libraries, or regulators. A shared reporting channel would reveal patterns that no individual shop can see.

If documented orders connect Australian books directly to AI scanning vendors, the preservation concern becomes much stronger. If the evidence instead reveals conventional resellers or collectors, the broadest claims will need revision.

The second signal is whether AI companies adopt public book-acquisition standards. A meaningful policy would distinguish common copies from potentially scarce editions, identify scanning methods, and explain what happens to physical originals.

The standard should also cover intermediaries. A company cannot credibly promise responsible sourcing while contractors operate under weaker rules. Audit rights and retention records would matter more than a general statement about respecting culture.

Non-destructive scanning should become the default for uncertain or scarce material. When destructive scanning remains appropriate, companies should verify that other accessible copies exist. They should preserve complete metadata and offer institutions the original before disposal.

Such commitments would strengthen the claim that AI digitization can coexist with cultural preservation. Continued secrecy would support the opposite conclusion, especially if buyers keep using confidentiality to hide the final destination of books.

The third signal is government or library action in Australia. Policymakers could ask national and state libraries whether current deposit systems cover the editions most exposed to bulk purchasing. They could also examine whether publicly funded digitization can preserve vulnerable collections before commercial demand removes them.

Australian copyright debates will continue, but lawmakers should avoid treating compensation as the only issue. Authors deserve clarity about training uses. Libraries and readers also need protection against the permanent loss of physical and publicly accessible sources.

A targeted response does not require banning book scanning. Governments could fund rarity databases, expand preservation grants, and create voluntary surrender channels for scanning vendors. Trade groups could develop a warning list for vulnerable categories without exposing valuable books to theft.

The public should also resist two easy conclusions. The first is that every discarded book represents a cultural catastrophe. Libraries and sellers already remove duplicate, damaged, or unwanted stock because indefinite storage is impossible.

The second is that digitization automatically equals preservation. A file held behind corporate access controls can disappear after a merger, model change, licensing dispute, or storage decision. Preservation requires durable stewardship, readable formats, metadata, redundancy, and accountable access.

That principle matters beyond publishing. AI systems increasingly absorb documents, images, audio, and institutional records while separating extracted information from its origin. People using AI for research should keep source context and maintain access to important originals.

Google News has made this dispute visible to a large audience, but attention is only the starting point. Readers should ask which titles were purchased, who received them, whether the originals survived, and who can inspect the resulting files.

The most constructive response is neither panic nor indifference. It is a demand for evidence and a preservation rule proportionate to the risk. AI developers can obtain useful text without treating every physical source as disposable.

Over the next three months, watch for verified Australian order records, public acquisition policies from major AI laboratories, and preservation guidance from libraries or government agencies. Each development will show whether the industry recognizes books as historical evidence or merely raw material.

The question is now unavoidable: when a company turns a book into training data, will it leave society with a richer archive, or only a smarter private model?

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page