Anthropic Claude Faces Sony and Warner’s Sweeping Music Copyright Lawsuit
Anthropic Claude faces a sweeping copyright lawsuit after Sony and Warner publishers accused its developer of deliberately acquiring protected music without permission. The August 28 complaint targets Anthropic, CEO Dario Amodei, and co-founder Benjamin Mann. It alleges that they helped build Claude with tens of thousands of copyrighted compositions, primarily lyrics.
The dispute reaches beyond familiar arguments about whether AI training qualifies as fair use. Sony Music Publishing, Warner Chappell Music, and affiliated publishers say Anthropic obtained source material through piracy, scraping, and unauthorized copying. Anthropic told TechCrunch that it disagrees with the claims and plans to defend itself in court.
That distinction matters because Anthropic has already confronted similar allegations involving books. In 2025, it agreed to a $1.5 billion settlement covering approximately 500,000 works allegedly downloaded from pirate libraries. The new case asks whether that history also exposes Anthropic Claude to music claims carrying potentially enormous statutory damages.
The Lawsuit Targets Anthropic and Its Founders
The publishers are treating data acquisition as intentional corporate conduct, not an accidental flaw in a training dataset.
The plaintiffs filed their 48-page complaint in the U.S. District Court for the Northern District of California. Dario Amodei and Benjamin Mann are named individually alongside Anthropic. That choice raises the personal stakes and distinguishes the case from a routine corporate copyright complaint.
The publishers allege a coordinated campaign involving torrenting, web scraping, downloading, storage, model training, and infringing Claude outputs. Their copyright complaint reportedly identifies tens of thousands of musical compositions. Discovery would determine the precise number and how Anthropic allegedly handled each work.
The works cited in reporting include “Ain’t No Mountain High Enough,” “Eye of the Tiger,” “Livin’ on a Prayer,” and “Uptown Funk.” The complaint also names “All I Want for Christmas Is You,” Leonard Cohen’s “Hallelujah,” and Taylor Swift’s “Paper Rings.”
These examples matter because music copyright involves several distinct rights. A composition covers the underlying music and lyrics, while a sound-recording copyright covers a particular recorded performance. The plaintiffs in this case are music publishers asserting rights in compositions, especially the lyrics allegedly copied into text datasets.
According to the filing, Anthropic acquired protected material through several channels. The alleged sources include pirate book libraries, lyric websites, web crawls, third-party datasets, and scanned physical books. The publishers claim those materials were stored and copied again when Anthropic prepared training corpora for Claude models.
A training corpus is the selected collection of text and other data used to adjust a model’s parameters. Creating one ordinarily requires developers to collect, clean, filter, duplicate, and transform source files. Each intermediate copy can become relevant when a copyright owner challenges how the original material was obtained.
The plaintiffs seek statutory damages reaching $150,000 for each work they establish was willfully infringed. Statutory damages allow qualifying copyright owners to pursue an amount set by law without proving an exact financial loss. The maximum is an available remedy, not an automatic award.
They also seek up to $25,000 for each alleged removal or alteration of copyright management information. That information can include names, ownership details, or other data identifying a protected work. The complaint further requests a jury trial, destruction of infringing copies, and an accounting of Claude’s training data.
Multiplying a maximum damage figure by tens of thousands of compositions produces a theoretical exposure measured in billions. A court would still need to decide liability, eligible works, willfulness, duplication, and the appropriate remedy. Early headline calculations therefore describe litigation stakes, not a predicted judgment.
The complaint states four claims. They cover alleged direct infringement through torrenting, contributory infringement by Amodei and Mann, direct infringement by Anthropic, and removal or alteration of copyright information. Every central allegation remains contested until established through evidence or resolved by agreement.
Anthropic’s immediate response was brief. A spokesperson said the company disagrees with the publishers’ claims and intends to defend itself “robustly in court,” according to its initial response. That leaves its detailed legal defenses for later filings.
Why Anthropic Claude Faces a Different Copyright Test
The central question is no longer simply whether model training transforms creative work, but whether Anthropic lawfully acquired the copies it used.
AI companies have often argued that training is transformative. Their models analyze patterns across large collections instead of distributing ordinary substitutes for every source. Under that theory, training can qualify as fair use even when developers do not license each work.
The music publishers are trying to move the dispute onto less favorable ground. They claim Anthropic deliberately downloaded pirated libraries and kept the resulting files in a permanent central repository. If proven, that conduct presents a different question from training on lawfully purchased material.
The distinction emerged clearly in Bartz v. Anthropic, the authors’ case that preceded the new lawsuit. U.S. District Judge William Alsup found Anthropic’s training use transformative in the circumstances before him. However, he separated that use from Anthropic’s alleged acquisition and retention of pirated books.
The judge characterized the creation of a permanent library from pirated copies as unlawful. That ruling did not establish a universal rule covering every AI model or dataset. It nevertheless gave copyright plaintiffs a roadmap that does not depend entirely on defeating a broad fair-use defense.
Sony and Warner’s publishers are following that roadmap. They allege that Mann personally used BitTorrent in 2021 to download at least five million books from Library Genesis. BitTorrent distributes files among participants, meaning a downloader can also upload pieces to other users.
The complaint further alleges that Anthropic employees downloaded at least two million books from Pirate Library Mirror in 2022. It cites internal communications and evidence disclosed during the Bartz litigation. The publishers say books in these collections included lyrics, sheet music, and song anthologies.
Those allegations do not prove every identified composition entered a Claude training run. Anthropic reportedly maintained a much larger central library and selected subsets for individual models. The plaintiffs must connect the protected works, relevant copies, and challenged conduct with sufficient evidence.
However, the publishers also allege other acquisition paths. They claim Anthropic scraped lyrics from Musixmatch and LyricFind, services that license lyrics for authorized uses. The complaint says Anthropic collected the text without securing comparable permission from the publishers.
The source question therefore spans more than pirate libraries. It includes web pages, compiled datasets, scanned books, and model-training copies. Different sources can create different defenses because public availability does not automatically determine copyright status or authorization.
The plaintiffs also challenge Claude’s outputs. They allege the models sometimes reproduce protected lyrics when users request songs, chord progressions, or related information. In some examples, they claim lyrics appear even when the user did not directly request them.
Model output introduces a separate infringement theory. Training litigation examines copies used while building a system, while output litigation asks what the deployed system produces for users. A court can treat those actions differently even when both involve the same composition.
Anthropic has implemented safeguards intended to limit requests for copyrighted text. The publishers argue those controls are insufficient and can be bypassed through revised prompts. That allegation will require testing across model versions, prompts, settings, and the specific works asserted in court.
The outcome will not determine whether all Anthropic Claude training is lawful. It will focus on the works, acquisition methods, models, and outputs supported by the evidence. Even so, its reasoning can influence how other developers document and defend their data pipelines.
The Book Settlement Changed the Publishers’ Leverage
Anthropic’s earlier settlement turned alleged data piracy from a theoretical AI risk into a measurable financial and operational liability.
In September 2025, Anthropic agreed to establish a settlement fund of at least $1.5 billion in the Bartz litigation. The proposed arrangement allocated approximately $3,000 for each of an estimated 500,000 covered books. It did not include an admission of liability.
The agreement also required Anthropic to destroy certain downloaded copies identified in the case. Claims based on Claude’s outputs could remain outside its limited release. Those boundaries are important because a settlement resolves specified claims without deciding every disputed legal question.
The author settlement followed a mixed ruling for Anthropic. Training itself received favorable fair-use treatment, but the permanent pirate library did not. A scheduled damages trial created significant downside if a jury found willful infringement across many works.
Sony and Warner now argue that some of the same downloads contained protected music. A pirated songbook can implicate book authors, music publishers, composers, and lyricists through different copyrights. Settling one group’s claims does not necessarily release claims owned by another group.
That creates the primary conflict in the new case. Anthropic presents Claude as a system built through transformative machine learning. The publishers present the underlying data pipeline as an industrial copying operation that ignored established licensing markets.
The music industry enters that conflict with extensive licensing infrastructure. Publishers routinely license compositions for recordings, performances, films, advertisements, digital services, and lyric displays. They can argue that AI developers had identifiable counterparties and established methods for seeking permission.
Anthropic can respond that model training differs from those conventional uses. A model does not operate like a lyric database simply because protected text influenced its parameters. That argument becomes harder when the dispute involves verbatim output or copies allegedly obtained through pirate sites.
The new case also joins a growing sequence of music claims against Anthropic. Universal Music Publishing Group, Concord, and ABKCO first sued over roughly 500 works in 2023. A later action expanded claims to more than 20,000 compositions and newer Claude models.
BMG filed another case involving 493 compositions in March 2026. Round Hill Music followed with its own action in August. Sony Music Publishing and Warner Chappell now bring the publishing arms of all three major music groups into litigation involving Claude.
The lawsuits are not identical. They involve different plaintiffs, works, time periods, acquisition theories, model versions, and alleged outputs. Courts must also guard against duplicate recovery when separate claims overlap around the same conduct.
Still, the expanding plaintiff group increases pressure on Anthropic. Each publisher can seek discovery into training sources, internal communications, dataset governance, and model behavior. Parallel cases also raise legal costs and increase the possibility of inconsistent rulings.
The industry comparison extends beyond Anthropic. OpenAI, Meta, and other developers face copyright litigation from authors, news organizations, visual artists, and rights holders. Their cases examine fair use, market harm, access, output similarity, and the provenance of training data.
Anthropic’s position is unusually exposed because Bartz produced judicial findings about pirated acquisition. Competing AI companies may face uncertain training claims, but they do not all share the same public factual record. Sony and Warner’s publishers are using that record as the foundation for a broader music case.
This pressure can affect enterprise customers even before any verdict. Businesses increasingly ask AI vendors about dataset provenance, indemnification, retention policies, and output safeguards. Legal disputes can influence procurement because customers want predictable ownership and manageable downstream risk.
Teams evaluating AI-generated research should maintain their own source records. A searchable personal knowledge base helps users separate quoted material, generated analysis, and original documents. That practice does not resolve vendor liability, but it improves internal accountability.
The Largest Claims Still Need Evidence
The complaint is serious, but its broadest conclusions depend on facts that discovery and model testing have not yet established publicly.
The publishers describe tens of thousands of infringed compositions. Public reporting does not yet provide a final court-accepted list connecting every claimed work to a specific file, dataset, training run, or output. The complaint begins that process rather than completing it.
A book’s presence in a pirate library does not automatically prove Anthropic downloaded that edition. A downloaded anthology’s presence in a central repository does not automatically prove its lyrics trained every Claude model. Each link in that chain matters for liability and damages.
The plaintiffs will likely seek download logs, dataset manifests, filtering records, storage information, training documentation, and internal messages. They may also examine which employees approved particular acquisition methods. Anthropic can challenge the authenticity, meaning, completeness, or legal significance of that evidence.
Naming Amodei and Mann individually raises another evidentiary burden. The publishers allege Amodei authorized the torrenting operation and Mann directly participated in it. Personal liability will depend on their actions, knowledge, control, and connection to the infringements ultimately proven.
The complaint reportedly quotes internal language portraying pirate libraries as questionable or unlawful. Such statements can support an inference of knowledge. Anthropic can still argue that isolated phrases do not establish every defendant’s intent for every work or later model.
Willfulness also matters because it affects potential statutory damages. Plaintiffs must do more than identify copying if they want maximum awards. They must support the required mental state while addressing defenses, time limits, registration rules, and possible overlap among asserted works.
Output claims require equally careful testing. Large language models generate responses probabilistically, meaning wording can vary with prompts, context, system instructions, and model versions. A documented output under one prompt does not establish identical behavior for every user or deployment.
The plaintiffs allege that Claude’s copyright guardrails are easily bypassed through repeated or reformulated requests. Anthropic can present evidence about refusal rates, mitigations, updates, and the rarity of challenged outputs. Courts may examine whether the product materially contributes to infringement or primarily supports lawful uses.
Similarity alone is not always enough. Copyright generally protects original expression, not ideas, titles, facts, or commonplace phrases. Song lyrics can receive substantial protection, but any output comparison must still identify protectable expression and meaningful similarity.
Market harm presents another contested issue. Publishers say unauthorized training and reproduced lyrics compete with licensed uses and undermine songwriters. Anthropic can argue that Claude serves broader analytical and conversational functions rather than replacing licensed recordings, sheet music, or lyric services.
The fair-use analysis could therefore split again. A court might view certain training uses as transformative while rejecting unlawful acquisition. It might also treat verbatim output, retained source files, or stripped ownership information as separate conduct requiring separate remedies.
The book piracy ruling does not decide those music questions automatically. Different plaintiffs own different rights, and the new complaint alleges additional conduct. Its reasoning is influential context, not a substitute for evidence in this case.
There is also uncertainty around remedies. An injunction could require Anthropic to delete identified source files, disclose additional training information, or strengthen output controls. A court must consider scope, feasibility, ownership, and whether the requested restrictions match proven violations.
Retraining a frontier model is costly and operationally complex, but the complaint does not guarantee that outcome. Courts can craft narrower relief aimed at files, processes, or outputs. Settlement negotiations could produce licensing, deletion, auditing, or payment terms without a final liability ruling.
The music publishers’ strongest narrative is deliberate acquisition from known pirate sources. Anthropic’s strongest distinction is that data storage, model training, and generated output are legally separate acts. The litigation will test how far courts preserve those boundaries.
What the Next Filings Will Reveal
Three signals will show whether this case becomes a broad test of AI training or a narrower dispute about documented piracy.
The first signal is Anthropic’s formal response. A motion to dismiss or answer should reveal which allegations it disputes and which legal defenses it prioritizes. The company must decide whether to focus on fair use, causation, ownership, personal liability, procedural defects, or limits on damages.
A response centered on lawful acquisition would weaken the publishers’ core narrative. A response that accepts certain downloads but disputes their use could narrow the case toward causation and remedies. Either direction would clarify what Anthropic believes the existing record can support.
The second signal is how the court manages discovery across related music cases. Coordinated discovery could expose common records once while reducing duplication. Separate proceedings could produce competing schedules, evidentiary disputes, and rulings involving overlapping Anthropic systems.
Training-data disclosure will be especially important. The publishers want an accounting of the material used to build Claude. Anthropic will likely seek protection for trade secrets, security-sensitive systems, and confidential information supplied by third parties.
Courts can address that conflict through protective orders and restricted access. The important question is whether plaintiffs obtain records detailed enough to connect particular compositions to Anthropic’s acquisition, storage, training, and output processes.
The third signal is whether licensing or settlement discussions begin before major merits rulings. The music industry already operates large rights-management systems. A negotiated framework could influence how AI developers obtain lyrics and compositions for future models.
A settlement would not necessarily answer whether training is fair use. It could instead reflect litigation risk, business priorities, or the cost of reconstructing historical datasets. Any agreement must therefore be read through its exact scope and released claims.
The parallel lawsuits also deserve attention. The earlier Concord and Universal actions cover specific compositions and model generations. BMG, Round Hill, Sony, and Warner can test similar theories through different catalogs and evidence.
A major win for one publisher could strengthen coordinated licensing demands. A ruling that separates transformative training from unlawful acquisition would reinforce the Bartz distinction. A dismissal based on missing links between files and models would weaken the broadest version of the publishers’ case.
For AI developers, the immediate lesson concerns provenance rather than a final ban on copyrighted training. Teams need auditable records showing where data came from, what permissions applied, which models used it, and when copies were deleted.
For enterprise buyers, the practical questions concern contracts and controls. Buyers should examine vendor representations about training data, copyright safeguards, output filtering, indemnification, and incident response. They should not assume that a model’s general safety positioning answers intellectual-property questions.
Knowledge workers also need realistic expectations. A Claude response can sound authoritative while reproducing or closely tracking protected material. Users should verify quotations, preserve source links, and distinguish model-generated summaries from text cleared for publication.
The broader stakes reach beyond one AI company. If courts impose large liabilities for poorly documented acquisition, model developers will have stronger reasons to license data or use traceable collections. If plaintiffs cannot connect works to specific conduct, sweeping catalog claims will become harder to sustain.
The music lawsuit overview captures why this filing stands out. It combines an unusually broad catalog with individual claims against company founders. It also builds on evidence already exposed through the earlier book litigation.
Anthropic Claude is not facing a final judgment today. It is facing a detailed allegation that its training-data strategy crossed a line courts have already treated differently from transformative model use. The next filings will show whether Anthropic can break that connection.
Readers should watch the formal defense, the scope of dataset discovery, and any licensing negotiations. Together, those signals will reveal whether this becomes a precedent for AI training or a case centered on piracy. The distinction will shape how developers build models and how customers evaluate them.



