Anthropic's Book-Destruction Claim Is Real, but the Rare-Book Story Is Not Proven
- Aisha Washington
- 2 hours ago
- 11 min read
Anthropic bought and destructively scanned millions of physical books, a practice now driving alarming Google News headlines about AI companies erasing rare titles. The documented operation involved removing bindings, scanning pages, and discarding the paper afterward. However, the strongest available evidence does not establish that Anthropic targeted rare or unique books.
That distinction changes the story. The confirmed facts reveal an industrial effort to turn legally purchased books into private AI training data. The broader viral claim adds several unproven elements, including multiple unnamed companies and the destruction of culturally irreplaceable editions.
The controversy also exposes a deeper conflict. AI developers want clean, professionally edited human writing, while authors and preservation advocates want compensation, access, and safeguards. Anthropic's solution reduced one legal risk by destroying each purchased copy, but it created an entirely different cultural and ethical problem.
What Anthropic Actually Bought, Scanned, and Destroyed
The central claim is substantially true for Anthropic, although its most dramatic rare-book detail remains unsupported.
A June 2025 federal court decision established the basic sequence. Anthropic purchased millions of print books, converted each one into a digital file, and discarded the physical copy. The company kept those files in an internal research library.
The process is called destructive scanning. A contractor removes a book's binding or cuts away its spine, leaving separate pages that can pass rapidly through commercial scanners. This method sacrifices the physical object to increase scanning speed.
The resulting files included page images and machine-readable text. Optical character recognition, or OCR, converts photographed text into characters that computers can search and process. Anthropic could then prepare that text for model development.
Judge William Alsup described the operation in his fair-use order. The ruling said Anthropic spent many millions of dollars acquiring millions of print copies, often in used condition.
The order also said one print copy produced one digital copy. Anthropic did not keep both versions or distribute the digital library publicly. It discarded the paper after completing the conversion.
That one-for-one replacement mattered legally. Alsup treated the format change like replacing physical library holdings with space-saving digital copies. He found the internal conversion fair use under the specific facts before the court.
The court did not declare every form of AI book scanning lawful. It evaluated Anthropic's purchased copies, private storage, and destructive replacement. Distribution, piracy, or additional copies would raise separate issues.
Anthropic's project began after its earlier attempts to obtain large digital book collections. The company hired Tom Turvey in February 2024. Turvey had previously worked on partnerships for Google's book-digitization program.
Turvey helped Anthropic obtain books in bulk and arrange industrial scanning. Reporting based on later-unsealed records identified the initiative as Project Panama. Internal language described an ambition to scan an extraordinarily broad portion of the world's books.
The Project Panama records revealed contractors, bulk orders, hydraulic cutting equipment, and high-speed scanners. One vendor proposal covered between 500,000 and two million books during six months.
The final total remains redacted in the released project materials. Still, the court itself used “millions” when describing Anthropic's purchases and scanning. The general scale is therefore not social-media speculation.
The discarded paper was not the project's product. The valuable output was a private, searchable corpus, meaning an organized collection of text prepared for computation. That corpus could support research and model training.
This is the verified foundation beneath the Google News claim. Anthropic really did purchase and destroy millions of physical books during digitization. The remaining questions concern which titles were involved and whether other companies followed the same path.
Why AI Labs Want Books That Never Lived Online
Books offer AI developers edited human language that is harder to find in the increasingly synthetic open web.
Large language models learn statistical patterns from extensive collections of text. Developers seek varied material containing coherent arguments, sustained narratives, specialist terminology, and carefully edited prose. Books combine those qualities across many subjects.
Web pages can also provide useful material. However, online text includes advertising, copied content, search-engine filler, broken formatting, and machine-generated material. Its provenance is often difficult to establish.
Printed books offer another advantage. A physical edition published before the recent generative-AI boom is clearly human-era material. Developers do not need to guess whether a chatbot generated the original paragraphs.
This concern is sometimes called model collapse or synthetic-data contamination. The terms describe different problems, but both involve models learning from outputs created by earlier models. Repeated recycling can amplify errors and reduce variety.
Books also contain material absent from ordinary web crawls. Academic monographs, regional histories, technical manuals, translated literature, and older nonfiction may have little searchable presence. Their scarcity online makes physical copies attractive as data sources.
Scarcity online does not mean rarity in the antiquarian sense. A discontinued textbook can be difficult to locate digitally while thousands of used copies remain in circulation. An obscure title can also be common enough to destroy safely.
That distinction has blurred in social posts. “Rare” can mean expensive, culturally important, unavailable elsewhere, or simply unfamiliar. Those meanings cannot be treated as interchangeable.
Anthropic reportedly bought books from major used-book retailers and distributors. Bulk purchasing favors ordinary copies that sellers can identify, price, and ship through established systems. It does not resemble the careful acquisition of manuscripts or unique annotated editions.
The public record does not provide a title-by-title inventory. Without that inventory, nobody outside the company can determine the scarcity of every scanned edition. Anthropic has not released its internal digital library for independent auditing.
The model developer also had a legal reason to favor purchased print copies. Ownership of a physical copy comes with rights under the first-sale doctrine, including the ability to resell or destroy that particular object.
Ownership does not transfer the underlying copyright. A buyer cannot normally reproduce and distribute an author's text merely by purchasing one book. Anthropic therefore maintained a one-copy relationship between the acquired object and its internal replacement.
The destruction strengthened that argument. If Anthropic had retained both the book and its complete digital duplicate, plaintiffs could characterize the project as expanding the number of usable copies.
This creates a strange incentive. The legally cautious route can favor destroying a physical book, even when a preservation-friendly scanner exists. Copyright strategy and cultural preservation point in opposite directions.
Google Books offers a useful historical contrast. That project scanned library holdings while generally returning the physical books. Its searchable index survived lengthy litigation, but its design and uses differed from Anthropic's private training corpus.
More recently, non-destructive partnerships have shown another route. Libraries can digitize public-domain works while preserving the originals and documenting their provenance. Licensing can also give AI developers access without requiring physical destruction.
Those routes involve coordination, restrictions, and transaction costs. Anthropic's project transformed book purchasing and scanning into an industrial procurement problem. Speed and scale carried more weight than preserving individual objects.
The mechanism explains why the event matters beyond the damaged paper. AI book destruction turns the used-book market into an input pipeline for proprietary model development. Publicly circulating copies become privately controlled data.
The Google News Claim Goes Beyond the Available Evidence
The viral version combines one confirmed industrial project with a much less certain claim about rare books and multiple AI companies.
Three separate propositions often appear together. AI developers value physical books. Anthropic destroyed millions of purchased books during scanning. AI companies are now systematically destroying rare titles.
The first two propositions have substantial documentation. The third does not yet have comparable proof.
Contemporary reports cite booksellers who received unusual orders for obscure, older, or out-of-print works. Some sellers believe intermediaries purchased those books for AI customers. The pattern deserves investigation, but suspicion is not attribution.
A seller can observe unusual demand without knowing the final buyer. An intermediary can source books for several industries. Research institutions, resellers, collectors, libraries, and digitization projects can all purchase unusual titles.
ISBNdb added another layer to the controversy. The book-data company published pages promoting bulk physical-book sourcing for large language model development. Archived promotional language discussed orders ranging from thousands to one million books.
The company also addressed the reputational problem created by destructive scanning. After the pages attracted attention, ISBNdb removed them and reportedly said the service had only tested market interest. It said no such service became operational.
That denial leaves a verification gap. Promotional material demonstrates that a company considered selling the service. It does not prove that an AI laboratory purchased books through that service or destroyed any titles it sourced.
The gap matters because online retellings often convert capability into completed activity. A vendor said it could source books, so posts conclude that unnamed AI companies used it. Booksellers saw unusual orders, so posts assign those purchases to AI labs.
Neither inference is impossible. Neither has been fully demonstrated through contracts, shipping records, customer names, or a verified chain of custody.
Evidence about rare books is weaker still. The 2025 scanning investigation noted that Anthropic's court records did not show rare books being destroyed. The documented purchases came largely through bulk commercial channels.
A rare title can nevertheless enter a bulk lot. Used-book inventories are messy, and some editions become scarce without carrying high prices. The absence of evidence is not proof that every destroyed copy was replaceable.
The responsible conclusion is narrower. Anthropic's AI book destruction is confirmed at a scale of millions. Claims involving rare or unique copies remain unverified because no reliable inventory has surfaced.
The plural phrase “AI companies” also requires care. Meta, OpenAI, Microsoft, and Google have faced allegations concerning copyrighted training material. Those cases often concern digital files, datasets, or publisher-supplied content rather than destructive physical scanning.
Meta has been accused of using pirated book collections to train Llama. Publishers have also sued Google over alleged uses of books in Gemini development. Those disputes strengthen the broader data-acquisition story, but they do not prove identical book-destruction programs.
Google's current Gemini lawsuit concerns alleged copying of works supplied for services such as Google Books and Google Play Books. It is not evidence that Google cut apart rare physical editions.
Combining every copyright lawsuit into one physical-destruction narrative makes the story easier to share and harder to verify. Different companies used different data sources, acquisition methods, and legal theories.
The best rating for the full viral claim is therefore mixed. Millions of purchased books were destructively scanned by Anthropic. Evidence does not establish that millions of rare books disappeared or that every major AI developer runs the same process.
The Real Conflict Is Private Data Versus Public Preservation
The strongest criticism is not that Anthropic burned a lost library, but that it privatized knowledge while destroying the source copies it acquired.
Physical destruction carries emotional weight because books are more than containers. Editions preserve typography, illustrations, marginal notes, bindings, ownership marks, and evidence about how people produced and read texts.
A text-only scan can omit many of those features. Even a complete page image may not preserve paper, ink, construction, or annotations hidden near a binding. Destructive scanning can therefore save words while losing the object.
However, libraries and publishers already discard large numbers of books. Damaged copies, outdated reference works, unsold stock, and low-demand duplicates routinely leave circulation. Destruction alone does not establish cultural loss.
Scarcity determines the stakes. Destroying one recent paperback from a large print run differs from cutting apart the last known copy of a regional publication. The viral claim collapses those cases into one alarming image.
Anthropic has not provided enough transparency to separate them. A public inventory could show titles, editions, publication dates, purchase sources, and the known number of surviving copies. No such comprehensive record is available.
The private destination creates further concern. Anthropic's scanned corpus does not operate like an accessible preservation archive. Researchers, authors, and readers cannot inspect the complete files or retrieve a book because its paper copy disappeared.
The company obtained a durable data asset while the public received no equivalent digital collection. Claude users can ask questions, but a model response is not a preserved text. It can summarize incorrectly, omit context, or generate unsupported details.
A language model is also not a searchable library catalog. Its parameters encode statistical relationships rather than a dependable shelf of complete works. Training does not guarantee faithful retrieval of any particular book.
That distinction undercuts claims that the books' knowledge simply migrated into the model. A private AI system cannot replace access to the original pages. Preservation requires stable copies, metadata, provenance, and future readability.
Authors face a related imbalance. Their books improved commercially valuable models, while they did not negotiate the physical-purchase scanning arrangement. Buying one used copy normally sends no payment to its author or publisher.
Anthropic's earlier acquisition of pirated digital libraries produced a different legal outcome. Alsup separated training from acquisition, finding the training use transformative while allowing claims about pirated copies to proceed.
Anthropic later reached a settlement covering eligible pirated works. A judge has since approved the copyright settlement, which totals $1.5 billion and allocates roughly $3,000 for each covered book.
That settlement does not mean the purchased-and-scanned library was unlawful. It concerned Anthropic's acquisition of pirated files, a legally distinct source. Conflating the two weakens both stories.
The legal contrast remains revealing. Purchased books could be scanned and destroyed under the court's fair-use analysis. Pirated files created major liability even though Anthropic allegedly sought similar underlying text for model training.
This gives AI developers an incentive to build cleaner acquisition records. It does not require them to preserve physical copies, disclose title lists, or make the digital replacements publicly available.
Preservation groups could propose practical safeguards. Scanning contractors could flag signed copies, manuscripts, first editions, unusual bindings, and titles with few known holdings. Specialists could review flagged books before cutting.
Companies could also publish non-sensitive catalogs after scanning. They might deposit preservation copies with trusted libraries under controlled conditions. Procurement agreements could require non-destructive treatment for scarce editions.
Licensing provides another route, although negotiations across millions of works are difficult. Collective licensing organizations could aggregate rights and distribute payments. Publishers could supply verified digital files without sacrificing physical copies.
None of these measures eliminates every conflict. Libraries would still face access restrictions, authors would debate compensation, and developers would protect proprietary datasets. They would at least separate model development from unnecessary cultural loss.
The current uncertainty benefits nobody. It lets exaggerated posts portray ordinary used books as unique treasures. It also lets companies avoid answering whether genuinely scarce material entered their cutters.
For knowledge workers, the lesson extends beyond books. A searchable answer is not the same as a durable source. Good personal knowledge management preserves context, provenance, and access rather than replacing documents with generated recollections.
What to Watch After the Google News Controversy
Three signals will determine whether this remains an Anthropic case or becomes evidence of a wider, continuing industry practice.
The first signal is a verifiable customer trail for book-sourcing intermediaries. Contracts, invoices, shipping records, or named customers would connect unusual bookseller orders to specific AI laboratories.
This evidence would strengthen the broader claim if it showed active, large-scale procurement for destructive scanning. It would weaken the claim if the promoted services never advanced beyond marketing experiments.
The distinction between an advertised service and a completed transaction is fundamental. Screenshots can establish what a vendor offered. They cannot establish who bought it or what happened to the books afterward.
The second signal is a title-level audit from Anthropic or its contractors. Even a limited independent review could classify scanned books by edition, scarcity, condition, and remaining library holdings.
Such an audit would test the rare-book allegation directly. Evidence of unique copies or highly scarce editions would make the preservation concern urgent. Evidence of ordinary commercial copies would narrow the controversy considerably.
Anthropic could disclose metadata without releasing copyrighted text. Titles, ISBNs, editions, purchase channels, and preservation checks would allow outside specialists to assess risk. Redactions could protect confidential vendor details.
The third signal is how courts, lawmakers, and publishers respond to one-for-one destructive conversion. The current ruling addresses a specific record, not every future scanning project.
A later court might distinguish model training from permanent library storage. Legislators could also require disclosure, licensing, or preservation deposits for large commercial corpora. Publishers may develop collective access systems before those rules arrive.
Competitor behavior matters within this third signal. A major laboratory could adopt non-destructive scanning and preservation commitments. That move would turn Anthropic's process from an industry necessity into a contestable business choice.
A company might also license publisher archives at scale. Such an agreement would show that legal access does not require cutting apart purchased books. It would place pressure on rivals to explain more secretive acquisition methods.
The public should watch the language used in future Google News coverage. Headlines that move from “millions of books” to “millions of rare books” are making a significant factual leap.
Readers should also separate physical scanning from digital piracy. Both concern access to copyrighted writing, but they involve different conduct, evidence, and legal consequences. Treating them as identical conceals the real tradeoffs.
The verified story is serious without embellishment. Anthropic converted millions of books into a private digital resource, destroyed the purchased paper copies, and used a process that a federal judge accepted as fair use.
What remains unverified is equally important. No public evidence shows that millions of rare titles disappeared. Nor has the same destructive program been conclusively attributed to every company included in viral posts.
The next credible report should answer concrete questions. Which company placed the order? Which editions were purchased? How many surviving copies exist? Who scanned them, and where did the digital files go?
Until those answers arrive, readers should resist both extremes. The controversy is not a fabricated hoax, but its most alarming version outruns the evidence.
Keep following the source documents behind each Google News headline, especially court filings and named-company records. The decisive issue is no longer whether destructive scanning happened. It is whether AI firms will disclose what they destroyed, preserve scarce works, and compensate the people whose writing made their systems valuable.