Mysterious Bulk Book Orders Raise Questions About AI Training
- Martin Chen

- 4 days ago
- 13 min read
Google News has surfaced a striking conflict: secondhand booksellers are receiving hundreds of unrelated orders, yet nobody has identified the customers behind them. Sellers across Britain and Ireland suspect AI companies or their contractors are buying obscure books for digitization. The purchasing patterns resemble known AI data-acquisition projects, but resemblance is not proof.
The orders reportedly span different subjects, languages, editions, and publication dates. They have come from buyers in Britain, Canada, continental Europe, and the United States. Some customers have used different names while directing shipments toward the same area, according to sellers interviewed for the original investigation.
That mystery matters because Anthropic previously bought and destructively scanned millions of books for Claude development. Court records established that physical book acquisition was not a hypothetical strategy. They also showed why AI developers might prefer purchased books over pirated digital libraries.
The central contest is therefore not booksellers against one identified technology company. It is documented purchasing behavior against an opaque supply chain. The transactions are real, while their suspected purpose remains unverified.
What Google News Revealed About the Book Orders
The unusual feature is not simply order size. It is the combination of unrelated titles, repeated buyers, and concentrated delivery patterns.
Stuart Manley, co-owner of Barter Books in Alnwick, told the Guardian that the orders began arriving about three months before the August report. Normal bulk orders often follow recognizable themes, such as transportation, sports, military history, or regional studies.
These requests did not. One reported list included an Estonian translation of John le Carré’s The Mission Song, a particular edition of Anne Brontë’s Agnes Grey, and an October 1983 issue of Warship magazine. Their only obvious shared feature was that each item could be individually cataloged.
Manley said his business sold hundreds of books to three buyers. The selections appeared unrelated by audience, subject, format, or expected resale demand. That randomness led him to wonder whether a machine, rather than a human curator, had assembled the lists.
Jim Shaughnessy, owner of MW Books in Claregalway, described a similar pattern beginning in May. His requests ranged from books about agricultural implements in 18th-century Africa to biographies of racing drivers from the 1950s.
Such combinations would be unusual for a private collector. They would also be inefficient for a conventional reseller trying to build a coherent catalog. However, they make more sense if the buyer is filling gaps in a large digital corpus.
David Gower-Spence, owner of BookLovers of Bath, reported a rush of orders between May and July. Around 200 books were involved, and some buyers reportedly appeared in accounts from other booksellers.
One unnamed seller told the Guardian that separate customer identities sometimes directed shipments to the same address. A delivery postcode examined by the newspaper pointed toward freight warehouses near Heathrow Airport.
That detail suggests consolidation, but it does not establish the final destination. Freight facilities routinely handle shipments for resellers, recyclers, libraries, exporters, and private institutions. An AI contractor is one possible customer among several.
The geographic range deepens the puzzle. Booksellers in Australia, Germany, the Netherlands, Britain, and Ireland have described similarly broad orders. Some included dated manuals, regional histories, specific translations, or out-of-print editions.
A conventional library acquisition could also produce an eccentric list. Institutions sometimes build comprehensive collections, replace missing volumes, or purchase donated-library inventories. Used-book recyclers likewise buy stock in bulk, then sort, resell, export, or pulp it.
The strongest signal is therefore the pattern across multiple sellers. The weakest part is attribution. Google News aggregation spread the Guardian headline quickly, but the available evidence still does not identify an AI laboratory.
That distinction should shape every conclusion. The story documents unusual demand for used books. It does not document who commissioned that demand or how the books will be used.
Why AI Developers Still Want Obscure Books
Books offer structured, edited, and often scarce language that open-web datasets cannot reliably replace.
Large language models learn statistical relationships from extensive collections of text. Training does not work like storing a searchable bookshelf, although models can sometimes reproduce passages encountered in their data.
Developers value books because they contain sustained arguments, edited prose, specialist vocabulary, and knowledge that may never have reached the public web. Older regional histories and technical manuals can cover subjects barely represented online.
Obscure books may be especially attractive because the open internet is increasingly repetitive. Search-optimized pages paraphrase other pages, low-quality sites copy reference material, and AI-generated text is entering new online datasets.
A scanned book can provide a cleaner historical sample. It can also fill a precise catalog gap identified by an International Standard Book Number, or ISBN. An ISBN is a unique commercial identifier assigned to a particular book edition.
If a buyer already has a database of missing ISBNs, automated purchasing becomes straightforward. Software can compare the target list against marketplaces, select available copies, and route them to consolidation facilities.
That mechanism would explain why sellers see strange combinations. The software would optimize for catalog coverage, condition, availability, and shipping cost. It would not care whether two books shared a human-readable theme.
Specific editions can matter because pagination, translation, illustrations, introductions, and supplemental materials vary. A database seeking comprehensive coverage may treat two editions of the same title as distinct objects.
This demand is not limited to model pretraining, the initial process that teaches a model broad language patterns. Digitized books can support evaluation, retrieval systems, specialized datasets, search indexes, and post-training research.
A model developer might not purchase the books directly. Data brokers, scanning vendors, recyclers, logistics companies, or database operators could sit between the bookstore and the final customer.
That possibility makes attribution difficult. A retailer may see only an account name and shipping location. Even the immediate purchaser may serve several industries or resell inventory through multiple channels.
The strongest precedent comes from Anthropic. Unsealed legal filings showed that the company ran a physical-book acquisition effort known internally as Project Panama.
According to a court-record investigation, Anthropic spent tens of millions of dollars acquiring millions of books. Workers removed the spines, scanned the pages, and sent the remaining material for recycling.
An internal planning document described an ambition to scan all the world’s books. The operation sought lawfully acquired physical copies after the company had already collected books from unauthorized online libraries.
That documented project makes the current suspicions plausible. It does not prove that Anthropic, or any other named developer, commissioned the orders now reaching independent sellers.
Other buyers could see the same opportunity. A scanning company can aggregate obscure books and sell digital access, metadata, or processing services. A recycler can extract valuable titles before pulping unwanted inventory.
Google News readers should therefore separate motive from identity. AI developers have a clear reason to seek books, while the identity of these particular buyers remains hidden.
The Real Conflict Is Lawful Purchase Versus Transparent Use
Buying a physical copy can strengthen a company’s legal position without answering whether authors consented to AI training.
The Anthropic litigation established an important distinction between obtaining a book and using its contents. US District Judge William Alsup found that training models on books could qualify as a transformative fair use in the circumstances before him.
The judge treated Anthropic’s lawfully purchased books differently from millions of files obtained through unauthorized shadow libraries. A shadow library distributes copyrighted works without permission from their rights holders.
The June 2025 opinion found that Anthropic could digitize purchased books for its central library and model training. However, the company still faced claims tied to pirated copies retained in its collection.
Anthropic later agreed to a $1.5 billion settlement covering piracy allegations. The settlement did not erase the court’s fair-use finding for training on lawfully acquired material.
That split creates a direct incentive to purchase physical books. A documented transaction offers cleaner provenance than an anonymous download from a piracy site. Provenance records where material came from and how it was acquired.
Physical ownership does not automatically transfer copyright. Purchasing a novel allows the owner to read, lend, resell, or dispose of that copy. It does not grant unrestricted reproduction rights.
Fair use can nevertheless permit some copying under US law. Courts weigh the purpose, nature, amount used, and effect on the market. AI training cases are testing how those factors apply to model development.
The US Copyright Office has warned against universal answers. Its training analysis says outcomes depend on factors including the source material, the purpose of the use, safeguards around outputs, and effects on licensing markets.
The legal picture also changes across borders. A transaction beginning in Britain or Ireland could involve scanning elsewhere, storage in another country, and model training in the United States.
Different jurisdictions offer different exceptions for text and data mining. Contract terms, database rights, copyright duration, and commercial purpose can further complicate the analysis.
For booksellers, the immediate decision is simpler but emotionally difficult. They own inventory and can generally sell it to willing customers. Yet many also consider themselves stewards of cultural material.
The concern becomes sharper when a requested item is scarce. Destroying one common paperback has little effect on public access. Destroying an uncommon regional history or specialized edition can remove a meaningful copy from circulation.
“Rare” also needs careful treatment. Out-of-print does not always mean irreplaceable, valuable, or unique. Thousands of copies might survive in private collections, libraries, and warehouses.
Conversely, a financially inexpensive book can preserve information that is difficult to find anywhere else. Market price is an imperfect measure of historical or research value.
Booksellers therefore face competing signals. Large orders can move stagnant inventory and support thin-margin businesses. The same orders can reduce the availability of works that shops, collectors, and researchers value.
The legal dispute focuses on reproduction and training. The booksellers’ concern adds a preservation question: should a buyer disclose destructive digitization when acquiring culturally important material?
No general rule requires that disclosure. Requiring every customer to explain a purchase would also create privacy and administrative problems.
A narrower approach may be possible. Sellers could apply extra review to signed books, archival items, unique annotations, local histories, and editions with limited institutional holdings.
That still would not solve the broader transparency problem. The companies receiving value from digitized books can remain several contractual layers away from the stores supplying them.
What the Evidence Does Not Establish
The AI explanation fits known behavior, but the public record lacks the documents needed to move from suspicion to attribution.
No named AI company has acknowledged commissioning the reported orders. No bookseller has published a contract connecting a buyer to Anthropic, OpenAI, Google, Meta, or another model developer.
There is also no public shipping record tracing the books from British or Irish stores to a scanning facility. A warehouse near Heathrow can consolidate international freight without knowing its eventual purpose.
Some sellers have identified Zoom Books among their customers. The Canada-based company describes itself as a book recycler, and sellers in several countries have discussed orders associated with it.
A recycler’s involvement would not prove an AI connection. Used-book companies already buy, sort, resell, export, donate, and pulp large quantities of books as part of ordinary operations.
Deleted marketing language, opaque aliases, or unusual purchasing patterns can justify further questions. None independently establishes who funded the orders or what happened after delivery.
The selections could result from automated resale analysis. A buyer might scan marketplace listings for price differences, institutional demand, replacement-copy requests, or export opportunities.
They could also support a private library or digitization project unrelated to generative AI. Universities, genealogy services, preservation groups, and specialist databases all seek obscure material.
Scale alone offers limited proof. Independent sellers notice 200-book or 500-book bursts because those orders can disrupt normal fulfillment. A large reseller may consider the same volume routine.
The timeline is more suggestive. Reports began spreading after court documents exposed Anthropic’s destructive scanning project. That disclosure gave booksellers a concrete framework for interpreting unfamiliar orders.
It may also create confirmation bias. Once sellers know that one AI company scanned millions of books, every irregular order can appear connected to the same industry.
This does not mean the concerns are misplaced. It means reporting should test the hypothesis with records rather than strengthening it through repetition.
One useful test would compare buyer lists across sellers. Shared payment processors, company registrations, telephone numbers, warehouse units, and forwarding accounts could reveal whether apparently different customers belong to one network.
Another test would follow a marked or otherwise traceable shipment, with appropriate consent and legal safeguards. Its route could distinguish domestic resale from international consolidation.
Reporters could also examine customs declarations and corporate procurement records. A scanning contractor handling millions of books would need facilities, equipment, labor, waste processing, and logistics capacity.
Model developers could reduce uncertainty through voluntary disclosures. They could publish acquisition policies, identify major data suppliers, describe preservation safeguards, and distinguish purchased materials from licensed collections.
The current opacity benefits neither side. Booksellers cannot make informed decisions, authors cannot assess potential uses, and AI companies face suspicion even when purchases have unrelated purposes.
Recent litigation against major developers intensifies that distrust. In July, publishers accused Google of using copyrighted books to train Gemini without permission. The allegations remain contested, but the Gemini lawsuit shows that book provenance is now a material legal issue.
That context explains why the Guardian story traveled widely through Google News. It connects a visible retail pattern with a hidden part of the AI supply chain.
However, a responsible conclusion remains narrow. Mysterious buyers are purchasing unusual collections of books. AI training is a credible explanation, not an established fact.
Who Is Pressured by the New Demand
The orders place booksellers at the intersection of commercial survival, cultural preservation, and an AI industry hungry for traceable data.
Independent sellers may welcome a customer who purchases slow-moving stock. Every book sitting on a shelf consumes space, labor, cataloging time, and working capital.
A sudden order can also overwhelm a small operation. Staff must locate each edition, verify its condition, pack it, correct listing errors, and manage marketplace communications.
When orders arrive through separate accounts, sellers may not recognize their combined scale until fulfillment begins. Multiple packages can ultimately converge at a forwarding warehouse.
Marketplaces face a different pressure. Their systems were designed to connect individual buyers with distributed inventory. Automated gap-filling programs can transform those platforms into procurement networks.
Platforms could identify unusual order clusters without revealing customer details publicly. They could warn sellers when several accounts share a delivery point or appear controlled by one organization.
Such monitoring must avoid treating legitimate bulk purchasing as misconduct. Libraries, schools, film productions, interior designers, and exporters routinely place large or eclectic orders.
Publishers and authors confront a more fundamental problem. A secondhand sale normally produces no new royalty, but it does not create a scalable derivative asset either.
Digitization changes that equation. One purchased copy might become part of a dataset used across repeated training runs, evaluations, or commercial model services.
AI developers face pressure to document lawful acquisition. Pirated datasets are cheap and comprehensive, but litigation has made their legal and reputational costs harder to ignore.
Buying physical books offers a more defensible route. It also introduces logistics costs, scanning errors, duplicate detection, storage questions, and scrutiny over destruction.
Libraries and archives may become the quiet counterweight. They can preserve access even when market copies disappear, but their holdings are not universal or permanently guaranteed.
Scarce local works present the highest risk. A regional history might have limited commercial demand while remaining valuable to genealogists, historians, and community researchers.
Digitization can preserve its words while destroying its physical context. Marginal notes, bindings, paper, inscriptions, inserted documents, and printing variations can carry information that optical character recognition misses.
That difference matters when buyers treat every book as interchangeable training material. A text corpus values extractable language, while an archive values the complete object and its history.
The conflict is not solved by declaring all book destruction unacceptable. Libraries deaccession duplicates, publishers pulp unsold stock, and recyclers process damaged volumes every day.
The missing element is informed triage. A bulk buyer optimizing for ISBN coverage may not recognize which copies deserve preservation before destructive scanning.
Booksellers can help identify those cases, but they cannot carry the responsibility alone. Their catalogs may omit provenance, annotations, or edition-specific features until someone physically checks each volume.
Industry standards could require scanning contractors to screen for signed copies, unique manuscripts, annotations, and scarce editions. Valuable discoveries could be diverted to non-destructive scanning or archival placement.
AI companies could also fund preservation copies. If a project destroys one acquired book, it could support the deposit of another copy or a high-quality digital surrogate in an accessible institution.
These safeguards would not settle copyright disputes. They would address a separate concern: preventing a data-acquisition race from quietly removing uncommon physical records.
For knowledge workers, the episode is a reminder that digital answers have supply chains. A chatbot response can draw indirectly on authors, editors, sellers, scanners, catalogers, and logistics workers whom users never see.
Keeping source documents and provenance inside a personal knowledge base offers one practical response. It lets users distinguish a model’s fluent answer from the evidence they can inspect themselves.
Three Signals That Can Confirm or Weaken the AI Theory
The next stage should focus on buyer identity, shipment destinations, and evidence of destructive scanning.
The first signal is a verified contractual link. An invoice, procurement agreement, payment record, or supplier acknowledgment connecting the buyers to an AI developer would strengthen the theory substantially.
A connection to a scanning vendor would also matter, especially if that vendor markets datasets or digitization services to model developers. It would still require proof that the reported books entered an AI project.
If investigations instead connect the purchases to ordinary resale, recycling, or private-library customers, the central claim would weaken. The strange selection pattern might then reflect automated commerce rather than model training.
The second signal is the books’ physical route. Deliveries reaching freight warehouses tell readers little unless investigators can identify the next destination.
Repeated shipments to industrial scanning sites, data contractors, or destructive digitization facilities would support the booksellers’ suspicions. Repeated shipments to resale warehouses would point elsewhere.
Customs records, warehouse tenancy data, and logistics documentation can help establish that path. Sellers may also compare delivery details privately through trade associations.
The third signal is a disclosure from an AI company or intermediary. Developers could state whether they are currently purchasing secondhand books, which vendors they use, and whether scanning destroys the originals.
A denial should be specific enough to evaluate. Saying that a company follows applicable law would not answer whether contractors purchase and scan books on its behalf.
Watch for new court filings as well. Copyright lawsuits have repeatedly revealed internal acquisition practices that companies did not previously discuss publicly.
The Anthropic documents changed how sellers interpret bulk orders because they exposed a complete mechanism. Physical books were acquired, stripped, scanned, digitized, and recycled at scale.
Similar records involving other companies would reinforce the larger conclusion that secondhand marketplaces have become infrastructure for AI data procurement.
Their absence would not prove the opposite. Private contracts can remain hidden, and intermediaries can obscure final customers. Yet evidence must eventually advance beyond pattern matching.
Readers encountering this story through Google News should keep both realities in view. AI companies have demonstrated a strong appetite for books, and one major developer conducted destructive scanning on a vast scale.
At the same time, the current British and Irish orders remain unattributed. Sellers have documented unusual behavior, not the identity or purpose of the ultimate buyer.
That verification gap is the story’s most important feature. It reveals how little transparency exists around the physical supply chains feeding digital models.
Over the coming months, look for shared buyer records, traceable freight destinations, and procurement disclosures. Those signals will determine whether this is an AI acquisition network or an unusual cycle in the global used-book trade.
Until then, treat the bulk orders as a credible lead rather than a solved case. Preserve the source material, follow the logistics, and ask who benefits when a physical book becomes invisible training data.


