Verge Artists Investigation: Creators Are Turning AI Slop Into Legal Leverage
- Ethan Carter

- 6 hours ago
- 12 min read
The Verge artists report captures a first for creators fighting generative AI companies: one group secured a landmark settlement after years of uncertain litigation.
Authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson sued Anthropic in 2024. Their case grew into a class action covering hundreds of thousands of books. A federal judge approved its $1.5 billion settlement on July 20, 2026.
That result sounds like a decisive victory over unauthorized AI training. It is not. The court accepted model training as fair use while separating it from Anthropic’s acquisition and retention of pirated books.
That distinction now defines the conflict. Creators have found legal leverage, but the strongest leverage concerns how training material was obtained. It does not necessarily give them control over training itself.
The settlement also arrived as publishers opened new fronts against Meta and Google. Those cases test whether the Anthropic result provides a repeatable strategy or an outcome tied to unusually damaging evidence.
The Authors Who Turned Discovery Into a Class Action
The Anthropic case began with individual writers tracing their books into datasets assembled for AI development.
Johnson writes deeply reported nonfiction, including The Feather Thief and The Fishermen and the Dragon. His third book, To Be a Friend Is Fatal, also became part of the litigation record.
The discovery was personal, but the problem was industrial. Authors had spent years hearing that large language models required enormous text collections. Few companies disclosed complete training inventories.
The Atlantic’s 2023 Books3 search project gave writers a rare window into one disputed collection. Books3 was an unauthorized corpus of nearly 200,000 books distributed within a larger dataset called the Pile.
Johnson found his work among those books. Other writers made similar discoveries, turning an abstract concern about machine learning into evidence attached to identifiable copyrighted works.
Anthropic developed Claude using large quantities of text. The company did not publish a complete list showing every work used in every model. That opacity made external dataset research especially important.
Bartz, Graeber, and Johnson filed their original complaint in August 2024. They alleged that Anthropic had copied pirated books while building and training Claude.
The plaintiffs did not need to prove that Claude reproduced an entire novel on demand. Their complaint focused on copies Anthropic allegedly downloaded and kept before model training occurred.
That choice proved important. Earlier AI copyright complaints often emphasized generated outputs, derivative works, or an author’s distinctive style. Those claims can encounter difficult similarity and ownership questions.
A writing style alone generally receives limited copyright protection. Plaintiffs also face challenges when they cannot connect a specific work to a specific model or infringing output.
The Anthropic plaintiffs followed a more concrete trail. They identified registered books, unauthorized digital copies, shadow libraries, and Anthropic’s internal collection practices.
Court records later described millions of books obtained from LibGen and Pirate Library Mirror. These services are shadow libraries, meaning repositories that distribute books without normal publisher authorization.
The company also acquired millions of print books, removed their bindings, scanned them, and discarded the physical copies. Judge William Alsup treated those purchased copies differently from the pirated downloads.
The case therefore separated three actions that public arguments often blend together. Anthropic acquired books, built a permanent digital library, and used selected material for model training.
Each action presented a different legal question. That separation created the opening that eventually produced a settlement.
Readers can trace the original allegations in the authors’ class action complaint. The filing accused Anthropic of building commercial AI systems with copied books obtained without permission.
The complaint was an allegation, not a judgment. Yet discovery produced enough evidence to move the piracy dispute beyond a generalized grievance about AI.
That changed the bargaining position. Creators were no longer asking a court to treat all machine learning as inherently unlawful. They were challenging specific copies and acquisition decisions.
Why the Verge Artists Victory Is Narrower Than It Looks
The authors won meaningful compensation without establishing that training a model on copyrighted books always requires permission.
Judge Alsup issued a mixed summary judgment in June 2025. He found that Anthropic’s use of books to train its language models qualified as fair use in the circumstances before him.
Fair use is a legal defense that permits certain unauthorized uses of copyrighted material. Courts assess purpose, the original work, the amount copied, and market effects.
Alsup viewed the training process as transformative because Claude did not simply provide readers with stored copies of the books. The model learned statistical relationships used to generate new text.
His fair use order called the training use “quintessentially transformative.” It compared model training with people learning from books before producing new writing.
The judge rejected the argument that copyright protects authors from all new competition. A system capable of producing competing text does not automatically infringe every work used in its development.
That conclusion handed Anthropic a significant legal victory. It also gave other AI developers language supporting a fair use defense for model training.
However, the ruling treated Anthropic’s permanent library as a separate matter. Downloading pirated books to create that library did not become lawful simply because some copies later supported transformative training.
Anthropic could not retroactively convert unlawful acquisition into fair use by pointing to a later technological purpose. The source and handling of the copies still mattered.
This distinction is the case’s central reversal. The plaintiffs lost the broad argument most associated with AI training, yet retained a claim carrying enormous potential exposure.
A trial over the pirated library could have examined individual works and statutory damages. Copyright law can impose damages for registered works without requiring proof of every lost sale.
Anthropic settled before that trial. The agreement created a $1.5 billion fund covering more than 482,000 works, according to the final approval order.
The estimated base payment is approximately $3,000 per covered work before approved deductions and claim adjustments. That figure reflects the settlement structure, not a universal valuation for training rights.
On July 20, 2026, District Judge Araceli Martínez-Olguín granted final approval. Her approval order described the relief as meaningful and entered judgment.
More than 91 percent of the covered works had valid claims submitted by the approval stage. Only a small share of class members chose to exclude themselves.
Those numbers demonstrate unusually broad participation. They also show why class actions can alter the economics of a dispute that no single author could afford to pursue alone.
The settlement does not require Anthropic to admit that model training infringed copyright. It also preserves certain future claims concerning later models and uses.
Anthropic has emphasized its fair use victory. Deputy general counsel Aparna Sridhar said the ruling established that training AI on books was fair use under current law.
The plaintiffs emphasize a different part of the result. Their attorneys describe the agreement as the largest known copyright recovery in history.
Both descriptions can be true. Training received legal protection, while the method used to assemble a training library generated historic liability.
That tension explains why the Verge artists story cannot be reduced to creators defeating AI. The result establishes leverage around provenance, not complete control over learning systems.
Pirated Libraries Became the Creators’ Strongest Leverage
The emerging litigation strategy targets provable copying before training, where technical complexity offers AI companies less protection.
Model training is difficult to explain in court. It involves tokenization, optimization, and repeated adjustments to numerical parameters. Those processes do not resemble a public library handing readers complete books.
Training-data acquisition is more familiar. A company either possessed an authorized copy, purchased one, received a license, or obtained material from another source.
That factual trail creates records. Download logs, internal messages, dataset manifests, file paths, and storage decisions can establish what employees knew and did.
The Anthropic litigation exposed a difference between buying print books for scanning and downloading them from pirate sites. The court accepted the former process as fair use within its analysis.
It did not extend that protection to the shadow-library copies. Keeping an unauthorized library for broad future purposes remained distinct from using a work during a transformative training process.
This gives creators and publishers a narrower but clearer route. They can investigate acquisition, storage, and distribution without first proving that a chatbot copied their prose in an output.
The theory also scales. A single unauthorized dataset can contain works belonging to thousands of rights holders. A class action can aggregate those claims around common collection practices.
Registration still matters. Settlement eligibility depended on criteria involving copyright registration, identifiers, ownership, and inclusion in the covered pirate collections.
Writers who never registered their work can face weaker remedies. People who publish online, create informal art, or cannot trace a dataset may have fewer practical options.
Visual artists confront additional difficulties. Image training collections can contain captions, thumbnails, transformed files, or links collected across many platforms.
An illustrator might recognize a stylistic resemblance without receiving proof that a particular image entered a model. Recognition alone does not establish copying under copyright law.
The lawsuits involving Sarah Andersen, Kelly McKernan, Karla Ortiz, Midjourney, Stability AI, and DeviantArt continue testing those boundaries. Their allegations concern training data, model behavior, and related branding claims.
These cases have survived some dismissal efforts, but they have not produced a comparable final resolution. Their outcomes will depend on specific models, datasets, works, and legal theories.
That variability matters. “Artists versus AI” describes a political conflict, not a single lawsuit with one controlling answer.
Book cases currently offer cleaner evidence because investigators can match titles, ISBNs, and digital files. A book in LibGen is easier to identify than an influence dispersed across generated images.
Creators are responding by becoming more organized. Industry groups maintain dataset searches, litigation trackers, model contract language, and collective licensing proposals.
Better recordkeeping strengthens that work. Authors can preserve copyright registrations, drafts, contracts, publication dates, and evidence showing where unauthorized copies appeared.
A personal knowledge system can help creators organize those records. It cannot establish infringement, but it can reduce the scramble when evidence surfaces.
Legal representation remains the decisive resource. Class counsel can fund discovery and expert work that individual writers usually cannot support.
The Anthropic settlement makes that investment easier to justify. It provides a visible example of a case where a narrow theory produced a large recovery.
It also warns AI companies that provenance is not administrative housekeeping. Training-data lineage can become a major balance-sheet and product risk.
A company that cannot explain where its corpus came from faces more than reputational criticism. It can lose the ability to separate legitimate sources from unauthorized collections.
Meta Shows Why One Win Does Not Settle Fair Use
Meta’s 2025 victory demonstrates that creators can lose even when a judge remains deeply skeptical about unlicensed AI training.
Thirteen authors, including Sarah Silverman, Richard Kadrey, and Ta-Nehisi Coates, accused Meta of using copyrighted books to train its Llama models.
District Judge Vince Chhabria granted summary judgment to Meta in June 2025. He found that the plaintiffs had not built the necessary record to prove market harm.
Yet Chhabria carefully limited his decision. His ruling did not declare Meta’s use of copyrighted works lawful for every author or every model.
The judge wrote that the plaintiffs had made the wrong arguments. He suggested that a stronger case could focus on AI-generated material flooding markets and reducing demand for human work.
That theory differs from straightforward substitution. A chatbot does not need to reproduce a particular novel to weaken the economic market supporting novelists.
It can generate vast quantities of competing material. If those outputs dilute attention, prices, or commissions, creators might argue that training harms the potential market for their work.
Proving that effect remains difficult. Plaintiffs need evidence connecting the challenged copying, the model’s capabilities, and measurable damage to a legally recognized market.
General anxiety about AI slop is not enough. AI slop refers to cheaply generated, repetitive content distributed at high volume with limited editorial effort.
Creators experience that volume directly on publishing platforms, social networks, and image marketplaces. Courts still require more than a recognizable cultural trend.
The Meta plaintiffs did not provide sufficient evidence for the theory Chhabria considered potentially persuasive. Meta therefore won the claims presented in that case.
The Meta ruling offers AI developers comfort, but only within limits. Different plaintiffs can bring different evidence against later models.
Meta also faces a separate 2026 lawsuit from five publishers and author Scott Turow. That complaint alleges copying and distribution involving millions of books used for Llama.
The new plaintiffs include organizations with extensive catalogs and detailed commercial records. They can present market evidence that individual authors may struggle to assemble.
They also allege direct involvement by senior leadership. Those allegations remain unproven, and Meta can challenge both the facts and their legal significance.
The contrast with Anthropic remains instructive. Anthropic received a favorable ruling on training, then settled claims tied to pirated acquisition.
Meta secured judgment because the plaintiffs’ market case was insufficient. The judge still left open the possibility that better evidence might produce another result.
No national appellate court has yet converted these district court decisions into a universal rule for generative AI. Other judges can apply fair use factors differently.
The disputed uses also vary. Training a model, storing a permanent library, reproducing passages, and distributing dataset files are not interchangeable actions.
Creators therefore need claims tied to identifiable conduct. AI companies need provenance controls covering every stage from collection through deployment.
The Verge artists narrative reveals progress, but not doctrinal closure. Litigation is becoming more precise because broad moral arguments have not delivered predictable outcomes.
Google Is Now Testing the Next Version of the Playbook
The newest Google litigation asks whether material supplied for one service can be repurposed for AI training without fresh authorization.
Hachette Book Group, Cengage Learning, Elsevier, Scott Turow, and S.C.R.I.B.E. filed a proposed class action against Google in July 2026.
The plaintiffs allege Google copied millions of books and other textual works to develop Gemini. Google had not filed a full response when the initial reports appeared.
The complaint targets several potential sources, including Google Books, Google Play Books, Google Scholar, web scraping, and disputed pirate collections.
This creates a different provenance problem. Publishers gave Google some books for specific search, retail, or academic services. That did not necessarily authorize every later AI use.
The plaintiffs argue that Google repurposed those works for Gemini without permission. They also allege that copyright-management information was removed or altered.
Those are allegations requiring evidence and judicial review. Google can dispute the factual account, invoke fair use, and argue that its agreements permitted relevant processing.
Google enters the fight with a major historical precedent. Its earlier Google Books project survived years of litigation over the scanning of millions of books.
Courts found that searchable snippets and text analysis served transformative purposes without offering full books as market substitutes. The Supreme Court declined to revisit that outcome in 2016.
Generative AI changes the factual picture. Gemini can produce new passages, explanations, summaries, and other material instead of returning limited search snippets.
Publishers argue that those outputs can compete with authors and licensed reference products. Google can respond that the models transform inputs and do not store accessible copies of entire works.
The acquisition question may again become decisive. A book delivered for Google Books carries a clearer contractual history than a file downloaded from LibGen.
The dispute will test whether permission for one digital service constrains later internal use. It will also examine how fair use interacts with existing commercial relationships.
A publisher announcement claims internal Google discussions identified serious legal risks. Those quoted materials currently reflect the plaintiffs’ account.
Google’s defense will matter because the company can distinguish its practices from Anthropic’s shadow-library downloads. It can also rely on the Google Books precedent.
The case nevertheless shows that the Anthropic settlement changed expectations. Publishers now have a concrete example of provenance claims surviving after training itself received fair use protection.
They can structure complaints around copying, authorization, retention, and market competition. They can also coordinate catalogs large enough to support class-wide analysis.
Google, Meta, and Anthropic face different allegations. Treating their cases as one referendum on AI would erase the distinctions most likely to determine each outcome.
The stronger comparison concerns operational discipline. Every developer needs to know what it copied, why it copied it, and which rights governed that copy.
A model’s technical sophistication does not answer those questions. Neither does a general claim that more data improves performance.
Three Signals Will Show Whether Artists Keep Winning
The next phase depends on claim payments, stronger market evidence, and judicial treatment of data obtained through different channels.
The first signal is the Anthropic settlement’s distribution process. Final approval closes a major litigation stage, but creators will judge the result through actual claims and payments.
More than 91 percent of covered works had been claimed by the approval stage. Administrators must still resolve competing ownership claims, documentation problems, and individual eligibility questions.
A smooth distribution would strengthen class actions as a practical creator strategy. Prolonged disputes or unexpectedly low net recoveries would weaken that lesson.
The result will also influence lawyers evaluating smaller creator groups. A headline settlement matters less if only large publishers can navigate its documentation requirements.
The second signal is evidence of market dilution in the Meta cases. Judge Chhabria effectively provided future plaintiffs with a map of what the earlier record lacked.
Authors and publishers must connect model-generated supply to recognized markets for books, licenses, commissions, or related creative work.
That will require more than counts of AI-generated titles. Plaintiffs need reliable evidence showing substitution, reduced demand, suppressed compensation, or damage to licensing opportunities.
If they build that record, Meta’s earlier victory will look narrow. If they cannot, fair use defenses will remain difficult to overcome through market-harm theories.
The third signal is how courts classify Google’s different data sources. Books obtained through partner programs, public websites, and unauthorized repositories do not share one legal status.
A ruling that separates those sources would reinforce the Anthropic framework. Provenance would become as important as the later training process.
A ruling treating all sources as part of one transformative system would favor developers. It would reduce the practical value of acquisition-focused claims outside clear piracy.
Watch procedural decisions before any trial. Motions to dismiss, discovery disputes, and class certification can reveal which theories a judge considers viable.
Also watch licensing behavior outside court. Developers increasingly sign content agreements with publishers, platforms, and media companies while defending unlicensed training elsewhere.
Those deals do not prove that licenses are legally required. They show that reliable access, current material, and reduced litigation risk possess commercial value.
For creators, the Verge artists lesson is neither surrender nor guaranteed victory. The strongest cases replace generalized outrage with registered works, traceable copies, and specific acquisition evidence.
For AI companies, fair use is not a substitute for data governance. A favorable training ruling cannot erase liability created while assembling a permanent library.
The larger question is now concrete: will courts extend the acquisition-versus-training distinction across Google, Meta, image generators, and other AI developers?
Creators should watch the evidence, not only the verdict headlines. Documenting ownership, dataset appearances, platform distribution, and market effects can determine whether the next complaint survives.
The first landmark settlement changed the bargaining landscape. The next three signals will show whether it created a durable legal playbook or one exceptional result.


