top of page

Japanese AI Book Buying Gives Used Stores a 5x Lift, but the Books May Not Return

2 hours ago
11 min read

Japanese AI book buying has reportedly pushed some used bookstores to five times their normal daily sales. The surprise boom carries an uncomfortable possibility. Many purchased books may be scanned, dismantled, and removed from circulation forever.

Stores have received orders for hundreds of books at once, according to Japanese reporting cited by Tom’s Hardware. Several apparently separate buyers reportedly direct their orders to the same logistics center in Okayama Prefecture. One documented shipment involved 50 tons of Japanese books sent to the United States.

The buyers have not been publicly identified, and the logistics operator declined to answer questions. No available evidence proves that every order is connected to an AI company. However, the purchasing pattern resembles a documented process that Anthropic used in the United States.

That process bought physical books, removed their bindings, scanned every page, and retained searchable digital copies. The printed pages were then destroyed. A federal judge later treated that physical scanning route differently from Anthropic’s acquisition of millions of pirated digital books.

The central conflict is therefore larger than an unusual sales spike. Japanese bookstores receive welcome revenue today, while researchers and readers face a possible shortage tomorrow. The transaction preserves information inside a private database but may destroy the public copy that supplied it.

Japan’s Used-Book Boom Has an Unidentified Buyer

The reported sales surge is real at the store level, but the purpose and identity of the bulk buyers remain unconfirmed.

Japan has a deep and varied secondhand book market. It includes nationwide chains, specialist dealers, neighborhood stores, and online sellers with inventories spread across warehouses. That distribution makes the market useful for assembling a broad corpus quickly.

A corpus is a large collection of text used for search, analysis, or AI training. Printed books offer edited, human-written material that is often unavailable through ordinary web crawling. Older books also predate the spread of machine-generated text online.

The latest buying activity appears focused on exactly that kind of material. According to the initial bulk-order account, stores reported days when online sales rose fivefold. Some received orders involving hundreds of books rather than normal consumer-sized baskets.

The subject mix also raises questions. Novels and comics traditionally lead many consumer orders, but the reported purchases reached across philosophy, political history, medicine, law, and social history. One example involved books about life and culture during the Edo period.

That breadth looks less like personal collecting and more like dataset construction. A reader usually follows authors, genres, periods, or specific research questions. A data buyer benefits from diversity, clean metadata, and enough volume to improve statistical coverage.

The orders reportedly came through multiple accounts. Yet many of those accounts requested delivery to one logistics center in Okayama Prefecture. That common destination is an important clue, although it does not identify the final customer.

A logistics center could consolidate books for export, resale, storage, library processing, or digitization. The operator reportedly declined to explain its role. Without invoices, contracts, or named clients, assigning the orders to a specific AI developer would be speculation.

The 50-ton shipment provides stronger historical context. Japanese reporting described records of that volume being sent to the United States for scanning associated with Anthropic. It does not establish that Anthropic placed the current Okayama orders.

Likewise, an unnamed foreign company reportedly contacted a western Japanese store about acquiring tens of thousands of books. That inquiry shows overseas demand at an industrial scale. It still leaves the buyer’s purpose and identity unresolved.

This distinction matters. The evidence supports reporting that suspicious bulk purchasing is occurring. It also supports concern that some books are being gathered for destructive scanning. It does not support describing every sale as a confirmed Anthropic purchase.

The more accurate conclusion is narrower. Japanese AI book buying appears to have reached a scale that ordinary retail behavior cannot easily explain. The routing and subject selection warrant scrutiny before the books disappear into an opaque supply chain.

Why AI Companies Still Want Physical Books

Physical books solve a data-quality problem that web-scale AI developers helped create, while giving buyers a clearer acquisition trail than pirate libraries.

Large language models learn statistical patterns from enormous text collections. The quality of that text affects how well a model handles reasoning, style, specialized vocabulary, and factual relationships. Books offer long, structured arguments that short web pages rarely match.

Books also cover knowledge that never became freely accessible online. Local histories, specialist manuals, academic works, older translations, and out-of-print titles may exist only on paper. Japanese-language collections are especially valuable for improving models beyond English.

The timing also reflects growing concern about synthetic data contamination. Since generative AI became widely available, machine-written pages have multiplied across the web. Crawlers can now ingest text produced by earlier models, including errors and repetitive stylistic patterns.

Human-authored books provide a cleaner alternative. Works published before the generative AI boom are unlikely to contain modern chatbot output. Their publication process often included editing, fact-checking, and consistent organization.

That does not mean every printed book is accurate or useful. It means books offer a relatively dense source of deliberate human language. For AI developers competing on multilingual performance, such material can be hard to replace.

Court records show how strongly Anthropic valued books. In the 2025 fair-use ruling, Judge William Alsup wrote that the company viewed books as a cost-effective route to a top-tier language model.

Anthropic initially obtained digital copies from several sources. The court found that it downloaded 196,640 books from Books3, at least five million from Library Genesis, and at least two million from PiLiMi.

Those collections contained unauthorized copies, according to the ruling. Anthropic kept a permanent central library even when particular books were not selected for a training mixture. That decision created legal exposure separate from model training itself.

The company later pursued a physical acquisition route. In 2024, it hired Tom Turvey, who previously led partnerships for Google’s book-scanning project. His task involved obtaining an extremely broad book collection without repeating the earlier legal problems.

Anthropic bought millions of physical books, many in used condition. Contractors removed the bindings, separated the pages, scanned them, and converted the images into machine-readable text. The company retained the resulting digital files.

This workflow explains why secondhand markets are attractive. Used copies can be sourced from many sellers, and purchasing them creates a conventional transaction record. Buyers can also target editions using International Standard Book Numbers, or ISBNs.

An ISBN identifies a particular book edition and format. It helps buyers avoid duplicate purchases, connect scans with metadata, and track language or publication details. Those capabilities matter when the order contains thousands of unrelated titles.

Destructive scanning is faster than photographing bound pages individually. Once the spine is removed, loose sheets can pass through high-speed scanners. Optical character recognition then converts the page images into searchable text.

The process trades preservation for throughput. A rare volume and a common paperback can enter the same mechanical workflow unless someone separates them first. The final dataset survives, but the purchased object does not.

This mechanism makes the Japanese orders plausible as AI-related acquisitions. It does not prove the connection. Still, the demand for clean Japanese text, combined with a documented industrial workflow, gives bookstore owners reasonable grounds for concern.

Japanese AI Book Buying Turns Revenue Into a Preservation Risk

The core reversal is simple: a purchase that helps a bookstore can still weaken the book market that made the sale possible.

For a seller, a fivefold sales day is normally good news. Used bookstores operate with limited space, uneven demand, and inventory that can remain unsold for years. Bulk orders release storage capacity and convert slow stock into revenue.

Many acquired books may also be ordinary surplus copies. Destroying one common paperback does not erase its contents from the world. Stores routinely discard damaged or unwanted inventory when preservation costs exceed likely demand.

The concern grows when buyers purchase by metadata rather than scarcity. A large acquisition system may treat an out-of-print local study like any other ISBN. It may not recognize that the listed copy is one of very few still available.

Once scanned, the physical book could be pulped. Its text might remain accessible only inside a private corporate repository. Libraries, researchers, collectors, and future readers would not automatically gain access to that digital version.

That outcome differs from a preservation project. Libraries digitize fragile materials to extend access while retaining originals when possible. A private AI corpus is primarily designed to improve commercial models, not to serve as a public archive.

The distinction also affects verification. Researchers need stable editions, page images, notes, illustrations, and physical context. Plain extracted text can omit marginalia, layout, typography, charts, and signs of an edition’s history.

Japanese publishing makes the issue particularly sensitive. Many works have limited print runs and never receive digital editions. Regional histories and older technical texts can disappear from the market without attracting national attention.

Used bookstores function as an informal preservation network. Their shelves keep unwanted private collections available until another reader finds them. Removing books in industrial quantities changes that network, even when every individual purchase is legal.

The pressure is unevenly distributed. Large chains can replenish popular titles from many branches. Small specialist sellers may hold distinctive inventories that took decades to assemble. A single broad order can permanently change what those shops offer.

This does not make bulk buying inherently improper. Owners have the right to sell their stock, and many need the revenue. Buyers also have legitimate reasons to digitize books, including accessibility, search, scholarship, and language technology.

The problem is the absence of informed choice. A store may believe it is supplying readers, libraries, or collectors. If the buyer instead intends to dismantle every volume, the seller cannot apply its own preservation standards.

Some owners might decline such orders for scarce material. Others might charge differently, retain one copy, or ask that rare books receive nondestructive scanning. They cannot make those decisions when intermediaries conceal the final use.

Japanese AI book buying therefore creates a market information failure. Sellers know the titles and quantities, but they may not know the destination or purpose. Buyers know the intended process, but they can spread orders across multiple accounts.

The result is a short-term transfer of value. Bookstores receive cash, while the buyer receives both a physical object and exclusive control over its digitized contents. The public may lose access without realizing a loss occurred.

This is why the story is not simply about AI companies paying for training data. Payment resolves one question about acquisition. It does not resolve preservation, transparency, licensing, or the social value of continued circulation.

Buying the Book Does Not Settle the Copyright Question

Ownership of a physical copy and permission to reproduce its text are separate rights, while AI training remains legally contested across jurisdictions.

A buyer can usually resell, lend, alter, or destroy a lawfully purchased book. Copyright law still limits making and using copies of the protected expression. AI scanning places those two principles in direct tension.

In Anthropic’s case, the federal court separated acquisition from training. Judge Alsup found the company’s use of lawfully obtained books for training transformative under the facts before him. The ruling did not create a universal exemption for every AI system.

The judge also distinguished Anthropic’s physical scans from its pirate-library downloads. Buying print copies and converting them for internal use received more favorable treatment. Keeping millions of knowingly pirated files did not.

That difference is one reason physical books now have strategic value. A documented purchase may reduce one category of legal risk. It does not eliminate claims involving reproduction, market harm, data transparency, or different national laws.

Anthropic later agreed to resolve claims involving its pirated copies. The proposed settlement covered an estimated 500,000 works and required destruction of original files obtained from the challenged digital libraries.

The copyright settlement did not establish that destructive scanning itself was unlawful. It highlighted how acquisition methods can matter even when courts view model training as transformative.

Authors and publishers remain skeptical of broad training claims. The Authors Guild has tracked roughly a dozen American copyright cases involving books and other protected material. Defendants include Anthropic, OpenAI, Microsoft, Meta, Nvidia, Google, Bloomberg, and Databricks.

Those pending AI cases address different datasets, models, plaintiffs, and legal theories. One judgment cannot be applied automatically to all of them. Appeals may also alter current interpretations.

Japan adds another legal system to the analysis. Japanese copyright law contains provisions that can permit information analysis under defined conditions. Their application depends on purpose, enjoyment of the work, and potential harm to rights holders.

An overseas shipment can involve several stages and jurisdictions. A book may be purchased in Japan, consolidated in Okayama, scanned in the United States, and used by a company operating globally. Each step creates different factual questions.

The current reporting does not reveal those contracts. It also does not identify who selects the titles, who owns the scans, or whether publishers receive licensing payments. Those gaps prevent firm conclusions about legality.

The same caution applies to destruction. Reports about Anthropic’s earlier workflow show that large-scale destructive scanning exists. They do not prove that all books routed through the Okayama center will be pulped.

Some could be resold after nondestructive imaging. Others could be stored or processed for ordinary export. Until a buyer, logistics provider, scanner, or shipping record supplies more detail, the final disposition remains an allegation.

That uncertainty should not end the inquiry. It defines the inquiry. The strongest article is not one that names an unsupported culprit, but one that identifies the records needed to test the claim.

Bookstores can retain account information, order lists, shipping labels, and unusual correspondence. Export records can establish destinations and weights. Scanning contractors can disclose preservation policies without revealing protected customer data.

Publishers and authors can also ask whether newly acquired Japanese-language datasets are licensed. Model developers increasingly describe safety methods and computing resources. Training-data acquisition remains far less transparent.

The legal battle will continue, but transparency can improve before courts deliver final answers. Buyers can disclose their purpose, protect scarce copies, and offer rights holders meaningful information. None of those steps requires exposing model architecture.

The Next Evidence Must Come From the Supply Chain

Three signals will determine whether this is a temporary retail anomaly or an industrial transfer of Japanese publishing into private AI archives.

The first signal is identification of the final buyer. Multiple purchasing accounts do not establish multiple customers when all shipments converge on one facility. Contracts or export documentation could connect that facility to a scanning company or AI developer.

A confirmed customer would strengthen the AI acquisition theory. Evidence of ordinary export, library resale, or commercial archiving would weaken it. Either result would replace suspicion with an accountable chain of custody.

The second signal is the treatment of scarce books. Sellers should track whether buyers select common duplicates or sweep up every matching ISBN. Orders containing out-of-print scholarship and regional history deserve particular attention.

A buyer that excludes scarce copies would reduce the preservation concern. Nondestructive scanning would reduce it further. A process that destroys books without rarity screening would support fears about permanent cultural loss.

The third signal is disclosure from AI developers. Companies can state whether they are purchasing Japanese print collections, whether contractors destroy originals, and how they handle copyright. Silence will leave intermediaries and bookstores carrying the reputational burden.

Developers could also explain whether scans remain private or support a public preservation copy. A private corpus offers commercial value to one company. An accessible archive can preserve at least some benefit for readers and researchers.

These signals matter beyond Japan. European and American booksellers have also reported unusual bulk orders associated with demand for AI training data. A repeatable international pattern would suggest a coordinated data-sourcing industry rather than isolated collecting.

It would also change the competitive picture. Model developers once competed mainly for computing capacity, engineering talent, and internet-scale datasets. They now compete for curated human knowledge that has not already circulated freely online.

That competition puts booksellers in an unfamiliar position. They are no longer only retailers serving readers. Their catalogs can become upstream data infrastructure for companies with much larger budgets and little public visibility.

Owners do not need to reject every bulk sale. They can introduce safeguards without abandoning new revenue. Those safeguards could include quantity reviews, rarity checks, buyer declarations, and separate approval for destructive use.

Marketplaces can also flag coordinated accounts that share destinations. A high-volume order across unrelated academic subjects is not necessarily abusive. It is still useful information for sellers deciding whether to release scarce inventory.

Libraries, publishers, and universities should watch the same patterns. A sudden disappearance of specialized titles can affect research long before official circulation statistics reveal a shortage. Shared watchlists could identify vulnerable works earlier.

Readers have a role as well. Anyone selling an inherited collection should ask whether unique material belongs in a library or archive before choosing the fastest bulk buyer. Once a rare copy enters destructive processing, the decision cannot be reversed.

The immediate Japanese sales surge may fade when the current acquisition target is complete. If so, stores could experience the reversal owners fear. Revenue rises during collection, then falls because desirable inventory is harder to replace.

The more durable question concerns who controls the resulting knowledge. A scanned book can improve language tools, search systems, and access for people with disabilities. Those benefits do not require secrecy or unnecessary destruction.

Japanese AI book buying should therefore be judged by what happens after purchase. Are scarce works protected? Are creators informed? Can researchers access preserved copies? Does the buyer accept public accountability?

Until those answers emerge, the responsible conclusion remains cautious. The bulk orders are documented through bookstore accounts, while the AI connection is strongly suggested but incompletely verified. The risk becomes irreversible only when the books enter the scanner.

Watch the Okayama supply chain, the fate of out-of-print titles, and disclosures from major model developers. Those three signals will show whether a retail boom became a hidden transfer of cultural memory.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page