AI News

The legality of training AI models on copyrighted books remains unresolved, despite a major court ruling that treated one company’s model training as lawful. A TechCrunch analysis of recent copyright litigation found that the decisive issue may not be whether developers used protected works, but how they obtained them and whether their AI products compete with the original markets.

The uncertainty affects companies building systems such as ChatGPT, Gemini, and Claude, as well as authors, publishers, enterprise buyers, and developers integrating generative AI into products. Courts are applying copyright rules written decades before large-scale machine learning existed, producing early decisions that may influence the industry without settling the broader dispute.

Anthropic ruling separates training from book acquisition

One of the most consequential cases involved Anthropic and a group of writers whose books were used in training the company’s AI models. According to TechCrunch, Judge William Alsup ruled that Anthropic’s AI training was lawful but ordered the company to pay a $1.5 billion copyright settlement because it had obtained some books from unauthorized online “shadow libraries.”

That distinction is central. The ruling treated ingesting copyrighted works to train a language model as closer to reading and learning from literature than reproducing a book for distribution. The penalty instead focused on the allegedly illegal acquisition of the training material.

Cathy Gellis, an attorney specializing in intellectual property, copyright, and technology, told TechCrunch that the reasoning could be favorable to AI companies because copyright law generally focuses on copying rather than simply using or experiencing a work. However, the financial consequences and the ruling’s treatment of data sourcing still create substantial exposure for model developers.

The case also does not guarantee that other courts will reach the same conclusion. The Anthropic decision remains an influential opening ruling within a much wider body of pending litigation.

Fair use turns on purpose and market competition

Much of the legal debate centers on fair use, a doctrine that can permit unlicensed use of copyrighted material for purposes such as criticism, education, parody, or other transformative activities. Courts commonly examine the purpose of the use, the amount of material involved, and its effect on the market for the original work.

Jason Henderson, an intellectual-property attorney quoted by TechCrunch, said courts appear especially concerned when a company trains on copyrighted material to build a product that directly competes with the source material. That concern shaped a separate case involving Thomson Reuters and Ross Intelligence, a legal research company that used Reuters content to develop a competing AI-based platform.

In that dispute, Judge Stephanos Bibas found Ross’s use was not transformative because it lacked a sufficiently different purpose or character from Thomson Reuters’s product. The decision gives AI companies a warning: training data and model outputs may receive different treatment depending on whether the resulting system substitutes for the original market.

That logic could become important for publishers and authors arguing that chatbots can generate synthetic books, summaries, or other content that competes with their work. TechCrunch reported that such an argument has not yet prevailed in court, leaving its eventual strength uncertain.

Evidence is influential, but not conclusive

The available evidence does not establish a single legal rule for all AI training. The Anthropic ruling supports the possibility that training itself can qualify as lawful use, while the Ross Intelligence case demonstrates that a directly competing product can weaken a fair-use defense. Different facts about the data source, product purpose, output behavior, and market effect could lead to different results.

Both cases are also narrower than the industry-wide question. A decision involving books acquired from illegal sources may not resolve the status of books obtained through licensed databases, web scraping, or other forms of access. Likewise, a ruling about a legal research product may not map neatly onto a general-purpose chatbot.

Gellis told TechCrunch that the initial decisions are shaping company behavior but could later be overturned or limited as litigation proceeds. Most major AI companies remain involved in pending copyright disputes, so current rulings should be treated as signals rather than a final settlement of the law.

The article also distinguishes training disputes from questions about copyright in AI-generated works. In Thaler v. Perlmutter, a court held that a work generated entirely by AI is not copyrightable. That separate rule raises practical questions about how much human contribution is needed for protection and how anyone would verify whether a work was created or assisted by AI.

What the uncertainty means for AI builders

For model developers, data provenance is becoming a product and risk-management issue rather than a back-office legal detail. The Anthropic case suggests that obtaining training material from unauthorized repositories can create liability even when the underlying model-training theory survives judicial review. Companies may therefore face pressure to document where books and other protected works came from, how they were processed, and what controls govern future data intake.

Product teams also need to assess whether their systems are positioned as general-purpose tools or as substitutes for a specific content market. The Ross Intelligence ruling indicates that a narrowly targeted AI product competing with an established information service may face a different fair-use analysis from a system with a broader, more transformative purpose.

Enterprise buyers should expect copyright assurances to remain qualified. Vendors may be able to describe their data practices and indemnification policies, but no contract can eliminate uncertainty while courts are still deciding whether training, retrieval, output generation, and market substitution should be analyzed separately.

Researchers and founders face a similar trade-off. Using high-quality copyrighted material may improve a system, but the legal value of that data depends on the acquisition method and intended use. Until more cases mature, conservative data governance and clear records may matter nearly as much as model performance.

What to watch next

The most important signals will come from later rulings in pending copyright cases involving AI companies, especially decisions that address books obtained through lawful channels rather than shadow libraries. Courts’ treatment of chatbot outputs that resemble or compete with authors’ works will also clarify how market harm is evaluated.

Builders should watch for more precise rules on licensed training data, opt-out systems, collective licensing, and the evidentiary standards used to prove what data entered a model. Another key question is whether appellate courts preserve the distinction between lawful training and unlawful acquisition established in the Anthropic dispute.

The industry should also track how courts apply copyrightability standards to AI-assisted work. The line between human authorship and machine assistance will affect publishers, software companies, and users who rely on generative tools to create commercial content.

Creati.ai perspective

The immediate lesson is not that AI training on copyrighted books is legal or illegal. It is that legality depends on a chain of facts: how the material was acquired, what the model or product does with it, whether it competes with the source market, and how much human authorship appears in the result.

For AI companies, the safest response is to treat legal uncertainty as an operational constraint. Documented data provenance, product-specific market analysis, and adaptable training strategies will be more durable than assuming one favorable ruling settles the question. The court decisions are beginning to define the boundaries, but they have not yet drawn a stable line.

Featured

AI Training on Copyrighted Books Remains a Legal Gamble for Model Builders

A TechCrunch legal analysis finds AI training on copyrighted books remains unsettled, with fair-use rulings creating risks for model builders.