top of page

Hacker News Spots an AI Book Flood, and Human Authors Are Losing Ground

Aug 3
12 min read

Hacker News surfaced a study of 14,419 books that challenges the easiest defense of AI-generated fiction: readers will simply ignore the bad material. The researchers found that books containing substantial detected AI text sold less effectively on average. Yet their sheer volume still captured sales, revenue, and scarce ranking positions from other books.

The working paper, submitted in July 2026 and still under review, examines self-published genre fiction sold primarily through Amazon. Its authors come from Stony Brook University, Columbia Law School, the University of Michigan, and the MIT Initiative on the Digital Economy.

Their central claim is not that AI writes better novels. It is that generative AI can reshape a creative market without winning a quality contest. Cheap production lets publishers release enough books to divide a limited pool of reader attention among far more titles.

That distinction matters for authors, publishing platforms, and AI companies. It also gives copyright courts new evidence for a disputed theory called market dilution. Under that theory, AI output can damage a market through scale, even when individual outputs perform poorly.

The study offers unusually detailed evidence, but it does not settle every question. Its AI classifications depend on a text detector, while its market data cover a selected segment of Amazon fiction. The analysis is also observational, so its strongest patterns do not prove that AI alone caused every decline.

Still, the results make one conclusion difficult to dismiss. Low average quality does not prevent automated content from becoming economically significant when production approaches industrial volume.

What the Hacker News Paper Actually Found

The study found a supply shock, not an AI bestseller takeover.

The researchers analyzed 14,419 titles released from January 2023 through March 2026. They matched full-text AI detection with daily sales observations extending through June 2026. The sample covered eight clusters of self-published genre fiction.

Each title entered one of three groups. Books with no detected AI text received a zero score. Light-AI books contained detected AI text up to a 25 percent threshold. Substantial-AI books crossed that threshold.

These labels describe detector output, not legal authorship. They do not establish that a named writer used a specific model. They also do not prove infringement or identify who generated a passage.

Within the sample, substantial-AI books represented 20 percent of titles. However, they produced only 12.1 percent of observed launch-window sales and 11.3 percent of revenue. Books with no detected AI text represented 62.9 percent of titles but earned 71.7 percent of sales.

That result supports the familiar criticism that AI-heavy books underperform individually. Yet the market-level trend points in another direction. Their observed sales share climbed from almost nothing in early 2023 to roughly one-fifth by the second quarter of 2026.

Their presence also reached the top of the market. The researchers constructed weekly Top-25 groups within each genre, using median Amazon sales rank. New entrants with substantial AI text eventually represented 31 percent of additions to those groups.

Books with no detected AI text remained the majority among leading titles. Even so, the combined share of light-AI and substantial-AI books approached 40 percent of constructed Top-25 slots by the second quarter of 2026.

The study therefore separates average performance from aggregate pressure. One AI-heavy book can sell poorly while thousands of similar releases collectively take meaningful attention. That is the mechanism behind the paper’s title and its strongest finding.

The contrast becomes clearer in the researchers’ growth measures. The cumulative released catalog grew 38.3 times its early-2023 level. Quarterly selling titles grew 19.2 times, while unit sales grew 7.3 times and revenue grew 8.9 times.

Those figures do not mean the entire Amazon catalog expanded at those rates. They describe the study’s observed release cohorts and indexed sample. However, the internal gap is striking because supply grew much faster than demand.

More titles were competing for each sale. Revenue per selling book consequently fell in six of eight genre groups between the 2023 and 2025 cohorts. For books with no detected AI text, it fell in seven of eight.

That final result anchors the dilution argument. The average did not decline only because weak AI titles entered the denominator. Books classified as having no AI text also earned less in later cohorts.

AI Book Market Dilution Works Through Volume

Generative AI changes publishing economics by making repeated production easier than repeated discovery.

Traditional book production contains several natural limits. Drafting takes time, editing requires attention, and every release carries an opportunity cost. Generative AI reduces the drafting constraint, especially in formula-driven categories built around recurring plots and familiar character types.

The study does not claim that every AI user becomes highly prolific. It examined 824 byline identities around their first substantial-AI release. Among the 385 that later released more substantial-AI books, 287 increased their monthly output.

That 74.5 percent share shows why average quality is an incomplete measure. An author or producer does not need every title to succeed. A large portfolio can spread risk across genres, pen names, covers, keywords, and release dates.

The resulting competition happens at several levels. Books compete for search placement, recommendation slots, category rankings, subscription reading, advertising exposure, and a reader’s limited browsing time. Each additional title becomes another candidate for those constrained positions.

This makes publishing resemble other algorithmic markets flooded by inexpensive content. The scarce resource is no longer the ability to produce an item. It is the ability to place that item before a buyer at the correct moment.

The researchers found stronger pressure where substantial-AI titles occupied more of the catalog. Books with no detected AI text held about 88 percent of constructed Top-25 positions in low-exposure groups. Their share dropped to roughly 63 percent in the highest-exposure groups.

The relationship grew stronger in genres with greater Kindle Unlimited participation. Kindle Unlimited gives readers access to a shared subscription catalog. That structure intensifies competition because new books enter the same pool of reading choices.

In Kindle Unlimited-heavy genres, the advantage held by no-AI books was 8.5 percentage points smaller for sales share. It was 8.4 points smaller for revenue share. These are adjusted associations, not experimental estimates.

The system also rewards frequent release schedules. A producer can test more titles, observe which combinations attract readers, and repeat successful patterns. Human authors working on slower cycles face a different risk profile.

A novelist might spend months refining one release. An AI-centered operation can distribute comparable production effort across many books. Even if most fail, the operation receives more chances to enter recommendations or find a responsive niche.

This is the paper’s real reversal. AI fiction does not need to persuade every reader that it is good. It needs to make publishing cheap enough that a minority of successful releases compensates for a large unsuccessful tail.

The study’s top performers demonstrate that this tail can contain commercially meaningful exceptions. Its leading substantial-AI titles reached significant unit volume. The top 15 titles also captured a sizable share of substantial-AI revenue within the eligible sample.

Those winners do not prove that automated production guarantees success. They show that the market contains a viable path from high-volume experimentation to meaningful sales. That path changes incentives for every competing producer.

Human Authors Face an Attention Problem, Not Just a Quality Problem

Human writers can produce better books and still lose visibility when platforms receive far more inventory than readers can absorb.

Publishing debates often assume that reader judgment will solve the AI-content problem. Under that view, weak books receive poor reviews, disappear from rankings, and leave serious authors largely unharmed.

The study shows why that filter is incomplete. Discovery happens before a reader can judge a full book. Buyers usually encounter a cover, title, description, short preview, category placement, recommendation, or advertisement.

A low-quality book can therefore consume attention without completing a sale. It can occupy a search result, trigger an impression, or enter a recommendation set. None of those events requires the reader to finish the book.

Ranking positions create an even tighter bottleneck. A category can display many books, but only a small number receive prominent placement. Every successful AI-heavy entrant displaces another title from that visible group.

The pressure extends beyond direct competition between individual novels. A growing catalog raises the cost of distinguishing legitimate authors, trusted publishers, and carefully edited work. Readers must evaluate more uncertain signals before making a choice.

Public labeling could reduce that uncertainty, but Amazon currently draws a limited distinction. Its KDP content rules require publishers to inform Amazon about AI-generated text, images, or translations.

Amazon defines AI-generated content as material created by an AI tool, even when a person later edits it substantially. It does not require disclosure for brainstorming, error correction, or other AI-assisted work when the creator writes the final text.

The disclosure goes to the platform during publication. Amazon’s rules do not promise a visible customer label for every disclosed title. This creates a gap between platform knowledge and reader knowledge.

That gap is central to the research. The authors report that none of the books in their dataset publicly disclosed whether they contained AI-produced content. Readers therefore could not reliably separate the study’s detected groups before purchasing.

Amazon introduced its disclosure requirement in 2023 after authors and advocacy groups raised concerns. At the time, the Authors Guild called the change a useful first step but pressed for public transparency, according to an Associated Press report.

The platform faces a difficult moderation problem. A strict label requires reliable information about production methods. Self-reporting can fail, while automated detection can misclassify edited work, translated prose, or distinctive human styles.

A looser policy reduces false accusations but leaves readers with less information. It also allows producers to benefit from ambiguity. A polished cover and plausible description can hide how quickly the underlying text was created.

Human authors must respond to this uncertainty even when they avoid AI. They may need stronger communities, direct reader relationships, recognizable series identities, or more consistent proof of authorship. These activities consume time that might otherwise support writing.

Knowledge management also becomes more important when creators must document drafts, sources, revisions, and editorial decisions. A searchable personal knowledge base can preserve that creative trail without forcing writers into a separate administrative system.

The deeper issue is not whether AI belongs in a writer’s workflow. Editing assistance and automated mass production create very different market effects. Current storefronts do not always make that difference visible to readers.

The Copyright Fight Now Has a Market Mechanism

The paper connects AI training disputes with evidence that machine-produced substitutes can crowd the markets supporting human work.

Copyright law examines several factors when deciding whether an unauthorized use qualifies as fair use. One important factor concerns the effect on the potential market for the copyrighted work.

Generative AI complicates that analysis. A model might train on books without reproducing an entire title in any single output. Yet it can support the creation of many new works that compete with books in the same genres.

The paper links its findings to Kadrey v. Meta, a copyright case involving authors and model training. In that litigation, the court considered a market-dilution theory but found that the plaintiffs had not supplied enough evidence for it.

The researchers address part of that evidentiary gap. Their study examines whether books containing substantial detected AI text enter the same market as other books. It then measures changes in sales, revenue, and rank as AI exposure rises.

Their results align with dilution, but the paper cannot decide liability. Copyright infringement depends on legal questions about copying, protected expression, fair use, and causation. A market correlation does not resolve those issues by itself.

The upstream and downstream acts also require separation. Training a model on copyrighted books is one act. A user generating and publishing a competing novel is another. A platform ranking that novel beside human-written books adds a third layer.

The U.S. Copyright Office has treated market effects as an important part of the training debate. Its 2025 analysis argued that commercial training can receive weaker fair-use protection when outputs compete in relevant creative markets. That assessment remains policy guidance, not a universal court ruling.

The current study introduces a possible bridge between abstract legal concern and observed market behavior. It suggests that substitution need not resemble a copied edition replacing its source. Thousands of loosely substitutable genre titles can distribute the harm across many writers.

The paper also analyzes distinctive language shared with existing books. It defines rare expressions as sequences of at least five words found in no more than five Google Books volumes. Those expressions must also be absent from a massive general-web snapshot.

Among the 50 highest-earning books in each comparison group, substantial-AI books showed greater rare-expression coverage. The average reached 45 percent, compared with 37.7 percent for books with no detected AI text.

Across the top 200 substantial-AI books, rare-expression coverage also rose with revenue. The increase was 7.6 percentage points for every tenfold revenue difference. The corresponding relationship for no-AI books was not statistically distinguishable from zero.

This pattern is provocative, but it requires careful wording. The measure detects aggregate textual overlap. It does not identify which model generated a passage, prove access to a specific book, or establish copying from an individual title.

The authors explicitly acknowledge that limit. Distinctive overlap can result from several processes, including common genre influences, training-data recall, editing choices, or repeated source material. Legal conclusions require evidence beyond statistical similarity.

Even so, the combination matters. High-volume production can pressure human authors economically, while successful AI-heavy titles show more overlap with existing language. Together, those findings strengthen questions about who supplies value and who captures it.

What the Numbers Do Not Prove

The findings deserve attention, but they remain evidence from one observed market and one classification system.

The paper is a working draft under review, not a completed peer-reviewed publication. Its methods and conclusions can change after criticism, replication, or additional analysis. Readers should treat its claims as serious preliminary research.

The first uncertainty concerns AI detection. The researchers used Pangram’s full-text detector and assigned every book a percentage score. Detector output can contain false positives and false negatives, especially across edited prose and changing model families.

Full-text analysis is stronger than judging a short preview. It gives the classifier more material and lets the researchers test multiple thresholds. Yet a larger sample does not eliminate systematic detector error.

The “no AI text” label also means no AI text was detected. It does not guarantee a completely human process. Likewise, a score above 25 percent does not prove that every flagged passage came directly from a generative model.

The study reports robustness checks for threshold changes and possible detector drift. Those checks strengthen the analysis. They cannot replace independent validation against production records from authors or platforms.

The second uncertainty concerns sample boundaries. The dataset focuses on self-published genre-fiction ebooks released during a specific period. Romance, mystery, fantasy, horror, science fiction, and related categories dominate the analysis.

These markets are suitable for studying volume because readers often seek familiar tropes and frequent releases. They are not representative of every publishing segment. Academic books, literary fiction, children’s books, nonfiction, and traditionally published titles operate differently.

The sales panel is also proprietary and maintained by a major publisher. According to the paper, it tracks a large portion of daily ebook unit volume. Outside researchers cannot fully audit that system using the public paper alone.

The third uncertainty is causation. Revenue per book fell while AI exposure increased, and declines were larger in more exposed genres. Those patterns support dilution, but other forces changed during the same period.

Advertising conditions, subscription behavior, genre popularity, pricing, seasonal releases, and platform algorithms can all affect sales. The authors use fixed effects and comparisons to reduce alternative explanations. Observational data still cannot recreate a randomized market.

Readers may also benefit from a larger catalog. More books can create greater variety, faster release schedules, and additional niche material. The researchers state that their evidence does not determine the market’s total welfare effect.

Human authors can lose average revenue while some readers gain more choice. AI-assisted writers can also use tools to improve accessibility or complete projects that otherwise would not exist. Those benefits do not erase dilution, but they complicate policy decisions.

The Hacker News framing introduces another limitation. A short discussion thread can draw attention to a paper without producing a representative public verdict. Four comments and 15 points show early interest, not a broad technical consensus.

The strongest responsible conclusion is narrower. This dataset contains a clear association between rising detected-AI supply and weaker per-title outcomes. It offers a plausible mechanism, but it does not prove universal harm across publishing.

Three Signals Will Show Whether the Flood Keeps Rising

The next phase depends on platform transparency, independent replication, and legal treatment of market dilution.

The first signal is Amazon’s treatment of AI disclosure. Its current rules collect information from publishers, but readers do not receive a universal public label. A visible disclosure system would create new evidence about purchasing behavior.

If Amazon exposes reliable labels, researchers can test whether readers avoid AI-generated books or accept them in specific genres. Stable sales after labeling would weaken claims that success depends on hidden production methods.

A public label would not solve every problem. Publishers might misreport their process, while AI-assisted and AI-generated content often sit on a continuum. However, platform data could become more useful than detector-based inference.

The second signal is independent replication. Researchers need comparable studies using different detectors, markets, periods, and sales datasets. Similar results across platforms would strengthen the paper’s claim that dilution reflects a general creative-market mechanism.

Contradictory results would also be valuable. They could reveal that Kindle Unlimited, genre fiction, or Amazon’s ranking system produces unusually strong effects. That would shift attention from generative AI alone toward platform design.

Replication should also separate quantity from visibility. Release counts show supply, while impressions, searches, recommendation placement, and advertising auctions reveal attention. Those data would clarify where displacement actually occurs.

The third signal is how courts handle market evidence in AI copyright cases. Judges must decide whether diffuse competition from generated works belongs within traditional market-effect analysis. They must also connect any harm to legally relevant copying.

The copyrightability guidance already distinguishes human creative control from purely generated output. Training disputes ask a different question about the source material used to build models.

Future rulings can strengthen the paper’s significance by accepting market dilution as relevant evidence. They can weaken it by demanding tighter proof that particular training acts caused substitution involving protected works.

Platforms do not need to wait for a final legal answer. Amazon can study release velocity, customer complaints, refund behavior, repetitive metadata, and undisclosed generation. Those operational signals may matter sooner than courtroom doctrine.

Authors should watch these same indicators. A growing catalog is manageable when reader demand grows alongside it. The danger increases when output expands faster than sales, leaving every title to compete for a shrinking share of attention.

The Hacker News discussion began with a paper, but the next evidence must come from markets and platforms. Watch whether Amazon makes AI labels visible, whether independent datasets reproduce dilution, and whether courts accept that mechanism.

For readers, the immediate action is simpler. Look beyond covers, rankings, and release frequency when choosing books. Follow trusted authors, sample prose carefully, and ask platforms for clearer production labels. Better discovery signals will not stop automated publishing, but they can decide whether scale alone keeps winning.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page