- 1There is no single answer: US courts have ruled three different ways in three major 2025 cases. Thomson Reuters v. Ross Intelligence found no fair use, Bartz v. Anthropic found training on purchased books transformative fair use while treating piracy separately, and Kadrey v. Meta found fair use on that specific record while the judge explicitly warned it does not make AI training broadly lawful.
- 2How the data was acquired, not just how it was used, is the fault line every ruling has turned on. Anthropic's record $1.5 billion settlement, given final court approval in July 2026, was over training on books pulled from pirate libraries, not over the training itself.
- 3The most consequential US cases are still unresolved as of mid-2026: New York Times v. OpenAI and the consolidated Authors Guild MDL remain in discovery with no trial date set, and Getty Images v. Stability AI is active in both a US court and, after a mixed November 2025 UK High Court ruling, on appeal-adjacent footing in the UK.
- 4The EU does not run this case by case. The DSM Copyright Directive's Article 4 permits text-and-data-mining unless a rightsholder opts out via a machine-readable reservation, and the EU AI Act's Article 53 now layers copyright-policy and public training-data-summary duties on top for general-purpose AI providers.
- 5For a startup licensing or buying training data today, the practical rule holds regardless of how the litigation resolves: data you have a license or documented consent for is safe to train on, while scraped copyrighted data carries real, still-unresolved legal risk that no amount of technical cleverness removes.
Is it legal to train an AI model on copyrighted data? As of 2026, the honest answer is that it depends on facts a court has not yet settled for most scenarios, and it depends differently in the US and the EU. US judges apply a fact-specific four-factor fair-use test, and three major 2025 rulings, Thomson Reuters v. Ross Intelligence, Bartz v. Anthropic, and Kadrey v. Meta, reached three different outcomes on three different fact patterns, with the biggest cases (New York Times v. OpenAI, the Authors Guild MDL, Getty v. Stability AI) still unresolved. The EU instead uses a statutory text-and-data-mining exception with an opt-out, now paired with new AI Act transparency duties. None of that uncertainty touches licensed or consented data, which is safe under any theory. It is scraped, unlicensed, and especially pirated data that carries the risk this guide maps out case by case.
On this page ▾
- Why there is no single yes-or-no answer
- The legal core: the four-factor fair-use test
- What US courts have actually ruled, case by case
- The landmark cases at a glance
- What the US Copyright Office has actually said
- How the EU's approach differs: statute first, opt-out, and now transparency
- What this means if you are deciding whether to use a dataset
- The honest bottom line
Why there is no single yes-or-no answer
Anyone searching for a clean answer to whether AI training on copyrighted material is legal will not find one, and a source that hands you one anyway is oversimplifying. As of 2026, the honest state of the law is that it depends on the jurisdiction, on how the data was acquired, on what the model outputs, and on facts that courts are still actively working through case by case. The United States decides this through litigation and a flexible statutory test applied to specific records. The European Union instead built a statutory exception with an opt-out, then layered new transparency duties on top. Neither regime has produced a general rule that says training on copyrighted works is always permitted or always forbidden.
What has changed since the first wave of AI copyright suits landed in 2023 is that courts have now actually ruled, repeatedly, and the rulings do not point the same direction. Three major US decisions came down in 2025: one rejected an AI company's fair-use defense outright, one split the difference between how a company trained versus how it acquired its data, and one sided with the AI company on a specific, thin record while its own judge warned against reading it broadly. That pattern, uncertainty with real edges, is the actual landscape, and it is what the rest of this guide walks through.
The legal core: the four-factor fair-use test
In the US, the question of whether copying a work to train an AI model infringes copyright almost always comes down to the fair-use doctrine in Section 107 of the Copyright Act. Fair use is not a bright-line rule. It is a balancing test across four factors that a judge weighs together: the purpose and character of the use (including whether it is commercial and, critically, whether it is transformative), the nature of the copyrighted work, the amount and substantiality of the portion used, and the effect of the use on the potential market for the original work.
The factor doing the most work in every AI training case so far is the first one, specifically the question of transformativeness. Copying a work to build something that serves an entirely different purpose than the original, the argument goes, weighs in favor of fair use even if the underlying copying is total. Copying a work to build a substitute product that competes with the original in the same market weighs against it. That single question, is training transformative enough, or does it just build a competing product, is exactly where US courts have split.
What US courts have actually ruled, case by case
Here is the current, verified status of every landmark US AI-copyright matter as of August 2026. Treat this as the anchor set; new rulings will supersede parts of it, which is exactly the point of not treating any of this as settled.
Thomson Reuters v. Ross Intelligence
Thomson Reuters sued Ross Intelligence, a legal-research AI startup, for training its product on Westlaw's headnotes, the short editorial summaries Thomson Reuters' attorneys write to accompany judicial opinions. In February 2025, Judge Stephanos Bibas of the District of Delaware rejected Ross's fair-use defense, finding no fair use after weighing all four factors together. [2] The court's reasoning centered on the fact that Ross was not building something transformative. It was building a directly competing legal-research product that reproduced the functional value of Thomson Reuters' own editorial work, which cut hard against transformativeness and market effect at once. This was the first major US ruling to address AI training and fair use head-on, and it went against the AI company.
The case is not over. Ross pursued an interlocutory appeal, the Third Circuit agreed to hear it, and oral argument was held in June 2026. As of the most recent public filings, both sides were still briefing how a separate, related Third Circuit ruling on incorporated technical standards should apply, and no appellate decision had issued. [2] Treat the February 2025 outcome as the trial-court answer on this specific fact pattern, not as a final, appeal-proof precedent.
Bartz v. Anthropic
In June 2025, Judge William Alsup in the Northern District of California split Anthropic's case into two distinct questions, and answered them differently. Training Claude on books Anthropic had legally purchased and then digitized, he ruled, was fair use, describing it as quintessentially transformative. Separately, he found that downloading and retaining millions of pirated books from shadow libraries to build a permanent research library was not itself protected by fair use, regardless of whether some of those books were later used in training, and he sent that piracy question to trial. [2]
“Pirating copies to build a research library without paying for it, and to retain copies should they prove useful for one thing or another, was its own use, and not a transformative one.”
That piracy trial never happened. In August 2025, Anthropic agreed to a $1.5 billion settlement covering the pirated-book claims, and the court granted final approval on July 20, 2026, calling the deal fair, reasonable, and adequate and overruling all objections filed against it. [2] Class members are receiving roughly $3,000 per work, about four times the statutory minimum for willful infringement. [4] Because this resolved as a settlement rather than a trial verdict, there is no appellate ruling on the merits of the piracy claim, and Judge Alsup's fair-use finding on purchased books stands as the operative ruling on that narrower question.
Kadrey v. Meta
Also in the Northern District of California, and also decided in June 2025, Judge Vince Chhabria granted summary judgment to Meta on the fair-use question in Kadrey v. Meta, a case brought by book authors over training the Llama models. On the specific record before him, the plaintiffs failed to present evidence of concrete market harm, which the court treated as effectively dispositive. [2] Meta won.
Judge Chhabria immediately narrowed his own ruling in language that has become the most quoted line in the entire body of 2025 AI-copyright case law. [2]
“This order does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful. It stands only for the proposition that these plaintiffs made the wrong arguments and failed to develop a record in support of the right one.”
In other words, a different set of plaintiffs, arguing market dilution with real evidence instead of the arguments this particular group chose, could plausibly win a similar case. The plaintiffs asked the court to certify that ruling for an immediate interlocutory appeal to the Ninth Circuit; Judge Chhabria denied that request in July 2026, preferring to let a full final judgment develop first so the appellate court gets a complete record rather than a piecemeal one. [2] The case remains open.
New York Times v. OpenAI and Microsoft, and the Authors Guild MDL
The New York Times sued OpenAI and Microsoft in December 2023 over training ChatGPT on Times journalism, alleging the model can reproduce Times articles nearly verbatim in ways that substitute for a subscription. In March 2025, Judge Sidney Stein denied most of OpenAI's motion to dismiss, letting the central infringement claims proceed while narrowing some peripheral ones. [2] As of mid-2026, the case is deep in contested discovery, including a January 2026 ruling compelling OpenAI to produce 20 million de-identified ChatGPT logs, and no trial date has been set. [4]
Running in parallel, a wave of author lawsuits against OpenAI, including the Authors Guild's, has been consolidated into In re OpenAI, Inc. Copyright Infringement Litigation in the Southern District of New York. In October 2025, Judge Stein denied OpenAI summary judgment on an output-based theory, holding that ChatGPT-generated plot summaries of plaintiffs' novels could themselves constitute infringement, separate from any claim about the training inputs. [2] That matters beyond this one case: it signals that generation-time outputs, not just the training copy, are an independent source of liability exposure. The consolidated case remains in discovery with no trial date set. [4]
Getty Images v. Stability AI
Getty Images sued Stability AI in both the US and the UK over scraping millions of Getty and iStock photos to train Stable Diffusion. The two cases have produced very different results so far. The UK claim reached judgment first: in November 2025, the England and Wales High Court rejected Getty's core copyright theory, holding that Stable Diffusion's model weights are not an infringing copy under UK copyright law because they consist of statistically trained parameters rather than a stored or reconstructable reproduction of any specific photograph. [2] Getty had already narrowed its own case before judgment, dropping its primary training claim because the training itself occurred outside the UK's jurisdiction. The court did find limited, historical trademark infringement from early Stable Diffusion versions that reproduced Getty's watermark in generated images, but rejected the broader trademark claims. [4]
The US case tells a different story procedurally, not yet substantively. The original Delaware suit was voluntarily dismissed in August 2025 after a jurisdictional dispute over Stability AI's UK parent company, and Getty immediately refiled in the Northern District of California, where the case remains active with copyright, DMCA, and Lanham Act trademark claims still pending. [2] No US court has yet ruled on the merits of Getty's copyright theory, so the UK outcome should not be read as a preview of the US result. US and UK copyright law define infringing copying differently, and a US court is free to reach a different conclusion on essentially the same technical facts.
The landmark cases at a glance
This table summarizes verified status as of August 2026. Every unresolved row is genuinely unresolved, not a placeholder, and should be checked again before anyone relies on it for a specific legal decision.
| Case / jurisdiction | Core question | Outcome / status (Aug. 2026) |
|---|---|---|
| Thomson Reuters v. Ross Intelligence (D. Del.) | Is training a competing legal-research AI on copyrighted headnotes fair use? | No fair use found (Feb. 2025). On interlocutory appeal to the Third Circuit; argued June 2026, decision pending. [2][4] |
| Bartz v. Anthropic (N.D. Cal.) | Is training Claude on purchased books fair use? What about pirated shadow-library copies? | Purchased-book training ruled transformative fair use (June 2025). Piracy claims settled for $1.5B, approved July 2026; no merits ruling on piracy. [2][4] |
| Kadrey v. Meta (N.D. Cal.) | Is training Llama on books, including pirated copies, fair use absent proof of market harm? | Summary judgment for Meta on this record (June 2025); judge explicitly declined to bless AI training broadly. Interlocutory appeal request denied July 2026. [2][4] |
| NYT / Authors Guild v. OpenAI & Microsoft (S.D.N.Y., consolidated) | Does training on, and generating near-verbatim outputs from, news articles and books infringe? | Unresolved. Core claims survived dismissal; discovery ongoing (20M logs ordered produced); no trial date set. [2][4] |
| Getty Images v. Stability AI (UK High Court) | Are AI model weights an infringing 'copy' under UK copyright law? | Getty's core copyright claim rejected (Nov. 2025); only a narrow, historical trademark finding upheld. [2] |
| Getty Images v. Stability AI (N.D. Cal., US) | Does scraping images to train Stable Diffusion infringe US copyright, DMCA, and trademark law? | Unresolved. Refiled in California Aug. 2025 after a Delaware jurisdictional dismissal; still active, no merits ruling. [2] |
What the US Copyright Office has actually said
The US Copyright Office ran its own multi-part study of AI and copyright, and the third and most relevant installment, on generative AI training, has an unusually tangled publication history. The Office released a pre-publication version of Part 3 on May 9, 2025, and it concluded that some AI training uses will qualify as fair use (particularly noncommercial research that does not let a model reproduce protected expression) while others will not, especially training on pirated sources to build a product that competes with the original works in their own market where licensing was reasonably available. [2]
The report's release was immediately entangled in a separation-of-powers fight: the head of the Copyright Office was removed from her position the day after the report went out, an action a federal appeals court later found was likely unlawful, and the litigation over that firing worked its way up before the Register was restored to her post. None of that changed the substance of the report. As of this writing, the Office has still not published a final, non-pre-publication version of Part 3, though it maintains that the final text will not differ substantively from the draft already public. [2] Treat Part 3 as the Office's considered, current position on the law, useful as persuasive analysis, but keep in mind it is not itself binding on any court, and several of the courts above have reached their own independent conclusions without deferring to it.
How the EU's approach differs: statute first, opt-out, and now transparency
The US resolves this question through litigation applied to specific facts, which is exactly why the answer above is a list of ongoing cases rather than a rule. The EU instead legislated a specific answer years before generative AI became a mainstream product. Article 4 of the 2019 Digital Single Market Copyright Directive creates a text-and-data-mining exception: reproductions and extractions of lawfully accessed copyrighted works are permitted for TDM, for any purpose including commercial AI training, unless the rightsholder has expressly reserved those rights. [2] For content published online, that reservation has to be made through machine-readable means, such as a robots.txt-style signal or website terms, which puts the opt-out burden on rightsholders rather than requiring AI developers to seek permission for every source up front. [4]
The 2024 EU AI Act layered a second obligation on top of that exception rather than replacing it. Article 53 requires providers of general-purpose AI models to maintain a copyright policy that identifies and honors Article 4 reservations, and to publish a sufficiently detailed public summary of the content used to train the model, covering pre-training, fine-tuning, and alignment data alike, using a template the European Commission published in July 2025. [2] The practical effect is a two-layer regime: a substantive rule about which training uses are permitted (governed by whether rightsholders opted out), plus a disclosure rule that makes it possible for rightsholders to find out whether their opt-out was actually honored. Nothing analogous to either layer exists at the federal level in the US, where the substantive question is still being litigated case by case and there is no general public training-data disclosure requirement.
| United States | European Union | |
|---|---|---|
| Governing mechanism | Case law applying the Section 107 fair-use test | Statutory exception (DSM Directive Art. 4) with an opt-out |
| Default rule | No default; a court decides fair use case by case | TDM is permitted by default unless the rightsholder reserves rights |
| Who bears the burden | The AI developer, to prove fair use if sued | The rightsholder, to make a machine-readable reservation |
| Transparency duty | None at the federal level as of 2026 | AI Act Art. 53 requires a public training-content summary for general-purpose models |
| Current certainty | Low; three 2025 rulings split three ways, biggest cases pending | Higher on the TDM question itself, though enforcement and Art. 53 compliance are still maturing |
What this means if you are deciding whether to use a dataset
Founders and buyers do not need to resolve a circuit split to make a defensible decision today. The rulings above converge on one operational lesson even though they disagree on almost everything else: acquisition method is what courts scrutinize hardest, and a license or documented consent removes the entire fair-use question rather than betting on how it resolves.
- 1Start with acquisition, not use. For every source in a dataset, determine whether it was purchased, licensed, collected with consent, or scraped without permission. This single fact predicts legal exposure better than anything about the model architecture or training purpose.
- 2Treat pirated or shadow-library sourcing as a hard stop. Every court to squarely address piracy, Ross, Anthropic, and Meta alike, has treated the pirated acquisition itself as a serious problem independent of how the data was later used. This is the one fact pattern with no favorable ruling anywhere in the current case law.
- 3Do not assume scraping public web content is safe just because it is public. Ross lost specifically because its use competed directly with the copyright holder's product; NYT v. OpenAI turns on similar market-substitution facts. Scraped copyrighted data can be technically accessible and still legally exposed.
- 4Check for an EU nexus separately from the US analysis. If any content, users, or model deployment touches the EU, confirm whether sources carry a DSM Article 4 reservation (robots.txt-style signals, site terms) and, for general-purpose models, whether the Article 53 training-content summary and copyright-policy obligations apply. [4][6]
- 5Ask whether the training use would compete with, or substitute for, the original market. Judge Chhabria's ruling in Kadrey flagged market dilution as the argument that could have changed the outcome even on a favorable record; that theory has not gone away, it was simply undeveloped in that case. [4]
- 6When the data is licensed or consented, stop worrying about fair use. A license is a contractual grant of rights that exists independently of how any court eventually rules on the fair-use doctrine. It is the only path in this entire landscape that is safe under every scenario above.
The honest bottom line
Training AI on copyrighted material is not categorically legal and not categorically illegal in the United States. It is a fact-specific fair-use question that has produced three different outcomes in three major 2025 cases, with the largest and most consequential disputes, against OpenAI and Getty, still working through discovery with no trial date in sight. The EU has taken a different path entirely, answering the core question by statute while adding a new transparency layer through the AI Act. What is settled, across every ruling and every regime, is that acquisition method matters more than almost anything else a court examines, and that piracy has not won a single round.
For a startup deciding whether a dataset is safe to buy, sell, or train on, that gives a workable answer even inside genuine legal uncertainty. Data you have documented rights to, whether through direct licensing, consented collection, or purchase, is a defensible asset regardless of how any pending appeal resolves. Data you cannot trace to a lawful source is a liability wearing the shape of an asset, and no ruling on the horizon looks likely to change that.
Trace your data's rights before you list or license it.
Get a free, no-obligation provenance review from the team that vets these deals every day. We will help you document exactly how each source was acquired, and whether it is ready for a buyer's legal team.
List your data on Dayda→