Sony Music and Warner publishers sued Anthropic over alleged piracy and unauthorized AI training, widening copyright risks around Claude and the sourcing of data.

Sony Music Publishing, Warner Chappell and other music publishers have sued Anthropic and its senior leaders in federal court, alleging the AI company illegally obtained and used copyrighted songs, lyrics and sheet music to develop Claude. The complaint frames the dispute not simply as a question of whether copyrighted material can be used for AI training, but how Anthropic allegedly acquired it.
Filed in the U.S. District Court for the Northern District of California, the case names Anthropic, CEO Dario Amodei and co-founder Benjamin Mann as defendants. The publishers accuse the company of torrenting, scraping and downloading large volumes of protected works without permission. Anthropic told TechCrunch that it disagrees with the claims and plans to defend itself.
The lawsuit adds a major music-industry challenge to an already significant legal problem for Anthropic. According to reporting by TechCrunch and The Decoder, the plaintiffs allege that tens of thousands of musical compositions were used in the development of Claude models, primarily in the form of song lyrics, alongside sheet music and related works.
The publishers describe the alleged conduct in exceptionally severe terms, characterizing it as a large-scale and continuing theft of intellectual property. Those descriptions are allegations in the complaint, not findings by the court.
The action follows Anthropic’s earlier dispute with authors and publishers in Bartz v. Anthropic. In that case, the company agreed to pay $1.5 billion after a court found that acquiring copyrighted books through piracy was unlawful, even while treating the separate question of using copyrighted material for training differently. TechCrunch reported that the court ordered the payment; The Decoder described it as a settlement.
That distinction is central to the new case. The music publishers are attempting to build on the earlier ruling by arguing that unauthorized acquisition can create liability independently of whether the files were later included directly in a commercial model.
The complaint allegedly focuses on Anthropic’s use of torrent networks and so-called pirate libraries, including LibGen and PiLiMi. The Decoder reported that the publishers claim Anthropic torrented at least seven million books from those sources, including books containing lyrics and musical material. Whether those books were used directly in commercial Claude models is expected to become a major issue in discovery.
The plaintiffs also reportedly challenge the collection of lyrics from licensed services such as MusixMatch and LyricFind, arguing that the scraping violated platform terms and publisher rights. Other datasets named in the complaint include Books3, The Pile and Common Crawl, which the publishers claim contained unauthorized material.
The allegations extend beyond digital files. The Decoder reported that the complaint accuses Anthropic of scanning and destroying used songbooks and sheet-music collections. It also alleges that a non-commercial model trained on material from LibGen and PiLiMi generated synthetic data later used in the development or feedback process for a commercial Claude model.
Anthropic has previously denied using books from LibGen and PiLiMi directly to train commercial products. The publishers’ argument, according to The Decoder, is that such a denial may depend on Anthropic’s definition of “training” and may not address the use of derivative or synthetic data.
The strongest facts currently available are that the publishers filed the case, that Anthropic and its executives are named defendants, and that Anthropic has rejected the allegations. The more detailed claims about torrenting, datasets, model-development pathways and executive direction come from the plaintiffs’ complaint as reported by TechCrunch and The Decoder. They remain contested.
The Decoder reported that the complaint seeks as much as $150,000 for each infringed work and up to $25,000 per violation involving the removal of copyright management information, such as notices or identifying metadata. The eventual exposure would depend on which claims survive, how many works are found to be covered, and whether the court accepts the publishers’ theories of liability.
The case also raises a difficult evidentiary question for AI companies: what counts as use of protected content when it passes through pretraining datasets, synthetic-data pipelines, evaluation systems or reinforcement processes? A company may distinguish between a source file and the later model, while rights holders may argue that each stage reflects the original unauthorized acquisition or reproduction.
A November 2025 ruling from the Munich Regional Court, cited by The Decoder, adds an international data point. That court held that protected lyrics can count as reproductions within model parameters and that generated lyric outputs can amount to unlawful public disclosure. The ruling is not a decision in the U.S. case, but it illustrates how courts are examining both model internals and generated content.
For model developers, the lawsuit increases pressure to document data provenance at a level that goes beyond broad dataset names. Teams may need records showing where files came from, what licenses applied, how terms of service were interpreted, and whether restricted material entered synthetic-data or reinforcement-learning pipelines.
For enterprise buyers, the dispute adds another layer to vendor diligence. Questions about a model’s copyright policy now include not only whether it reproduces lyrics or books, but also whether the provider can demonstrate lawful acquisition and governance of training data. Contractual indemnity, audit rights, model-version controls and restrictions on high-risk outputs may become more important in procurement.
The case could also affect the economics of AI training. If courts treat pirated acquisition as a separate source of liability, companies may have to spend more on licensed corpora, filtering, provenance systems and legal review before data reaches a training run. That would not resolve every copyright question, but it could make data sourcing a more visible competitive factor among foundation-model providers.
The first signals will come from Anthropic’s formal response and any motions challenging the complaint. Discovery should show whether the alleged files entered Claude’s commercial training pipeline, were used indirectly through synthetic data, or were retained for other development purposes.
Builders should watch for court rulings on executive liability, the treatment of torrenting as a standalone infringement, and the publishers’ ability to connect specific works to particular models or training stages. Licensing agreements between AI companies and music-rights holders would also indicate whether the lawsuit is pushing the market toward negotiated access.
The case’s interaction with Bartz v. Anthropic will be especially important. If the earlier decision becomes a foundation for liability over acquisition methods, other rights holders could pursue similar claims against model developers whose datasets contain disputed material.
This lawsuit shifts attention from the familiar question of what an AI model produces to the less visible systems that assemble its training data. For Anthropic, the legal risk may turn on procurement, scraping controls and internal data lineage as much as on Claude’s behavior at the prompt level.
For the broader market, the practical lesson is narrower but consequential: “not directly trained on” may not be a sufficient compliance answer if courts examine intermediate models, synthetic data and feedback pipelines. AI companies that can prove where their data came from—and what they did when rights were unclear—will be better positioned as copyright litigation moves deeper into the infrastructure of model development.