The Reckoning That Was Always Coming
For the first half of this decade, the dominant assumption in AI development was deceptively simple: if data was publicly accessible on the internet, it was fair game for training. Billions of web pages, millions of books, decades of journalism, centuries of creative work — all of it ingested, compressed, and encoded into the weights of large language models with little legal scrutiny and even less compensation to the people who created it.
That assumption is now being systematically dismantled. In courtrooms in Delaware and Hamburg, in regulatory offices in Brussels and Sacramento, and in licensing negotiations between AI companies and the world's largest media conglomerates, the rules governing AI training data are being rewritten in real time. The era of unchecked scraping is ending. What replaces it will define the economics of artificial intelligence for the next decade.
This explainer maps the legal terrain as it stands in September 2026: the landmark cases, the regulatory frameworks, the emerging licensing market, and the structural questions that remain unresolved.
The US Litigation Landscape: Fair Use Under Pressure
The United States has no federal statute specifically governing AI training data. Instead, the question of whether training a model on copyrighted material constitutes infringement has been left to the courts, which must apply the four-factor fair use test under 17 U.S.C. § 107 — a doctrine developed long before neural networks existed.
Thomson Reuters v. ROSS Intelligence: The Pivotal Ruling
The most consequential US decision to date came in February 2025, when Judge Stephanos Bibas of the US District Court for the District of Delaware issued a ruling in Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc. that sent shockwaves through the AI industry. ROSS had used Westlaw headnotes — editorial summaries of judicial opinions created by Thomson Reuters lawyers — to train a competing AI-powered legal research tool. The court found that at least 2,243 of those headnotes were original and protected by copyright.
The court found that ROSS's use was not transformative because it was intended to create a direct, competing legal research tool — a ruling that sent shockwaves through every AI company training on proprietary datasets.
The fair use defence failed on all four factors. Critically, the court found that ROSS's use was not transformative — it was not creating something new and different, but rather building a market substitute for the very product it had copied from. The ruling also recognised that AI training creates a potential future market for licensing legal content specifically for AI purposes, and that ROSS's actions harmed that market.
As of June 2026, the case is before the US Court of Appeals for the Third Circuit on interlocutory appeal. Oral arguments in June 2026 saw a three-judge panel express scepticism about ROSS's arguments, questioning whether its tool was meaningfully different from existing legal search engines. A ruling affirming the lower court would significantly strengthen the position of copyright holders across every sector currently litigating against AI developers.
The Broader Litigation Wave
The court found that ROSS's use was not transformative because it was intended to create a direct, competing legal research tool — a ruling that sent shockwaves through every AI company training on proprietary datasets.
The Thomson Reuters case is not an outlier — it is the leading edge of a wave. By mid-2026, nearly 70 lawsuits had been filed against AI companies over training data practices, spanning news publishers, visual artists, authors, musicians, and software developers. Cases such as Andersen v. Stability AI continue to test whether model training constitutes direct infringement, while the music industry has seen a "domino effect" of settlements and licensing agreements following initial litigation against generative audio services.
The US Copyright Office, for its part, has recommended against new statutory exceptions for AI training, favouring the development of licensing markets without government intervention. This stance effectively places the burden of establishing legal clarity on the courts — a slow and expensive process that creates significant uncertainty for AI developers and rights holders alike.
California has moved faster than the federal government. The California Generative AI Training Data Transparency Act, effective January 2026, mandates high-level disclosures regarding training datasets, creating a fragmented state-level compliance environment in the absence of federal preemption.
The European Framework: Structured Rights, Mandatory Compliance
The European Union has taken a fundamentally different approach. Rather than leaving the question to litigation, the EU has constructed a layered regulatory framework that explicitly addresses AI training data — and its territorial reach extends far beyond European borders.
The DSM Directive and the Text and Data Mining Exception
The EU Digital Single Market Directive (2019/790) established a text and data mining (TDM) exception in Articles 3 and 4. Article 4 permits the reproduction and extraction of lawfully accessible works for TDM purposes — including AI training — provided that the rightsholder has not expressly reserved these rights in a machine-readable manner. This opt-out mechanism is the cornerstone of the EU's approach: training on publicly accessible data is permitted by default, but rights holders can withdraw that permission through technical signals such as robots.txt entries, asset-level metadata, or dedicated AI exclusion protocols.
The legal standard for opt-outs has been clarified through litigation. The Higher Regional Court of Hamburg, in Kneschke v. LAION, held that an opt-out must be "interpreted by machines" — written in natural language alone is insufficient. Dutch courts, in DPG Media v. HowardsHome, have further indicated that opt-outs must be clear and specific; generic or overly broad reservations may be deemed ineffective.
The EU AI Act: Binding Obligations for GPAI Providers
The EU AI Act (Regulation (EU) 2024/1689), now in full enforcement, adds a further layer of obligation for providers of general-purpose AI (GPAI) models. Under Article 53, GPAI providers must publish a "sufficiently detailed" summary of training content — including dataset modalities, sizes, languages, and the top 10 per cent of crawled domains by volume — and must implement a policy to comply with EU copyright law, specifically respecting machine-readable opt-outs under the DSM Directive.
The EU AI Act's territorial reach is extraordinary: copyright obligations apply to models offered in the EU regardless of where they were trained, forcing US-based companies to reconcile fair use assumptions with European opt-out mandates.
The territorial implications are profound. Recital 106 of the AI Act asserts that these obligations apply to models offered in the EU regardless of where training occurred. A US-based AI company that trains on data without respecting machine-readable opt-outs — a practice that may be legally defensible under US fair use doctrine — faces potential fines of up to 3 per cent of global annual turnover when that model is deployed in the EU. This creates what legal scholars are calling "compliance dilemmas": training practices lawful in one jurisdiction may constitute regulatory violations in another.
The Court of Justice of the EU is currently considering a preliminary ruling in Like Company v. Google Ireland Limited (Case C-250/25), which is expected to provide definitive guidance on whether training large language models constitutes reproduction under the DSM Directive and the extent to which the TDM exception applies to generative outputs. The distinction between "input" (training) and "output" (generation) is emerging as a critical legal boundary: while the analytical process of training may be covered by the TDM exception, the unauthorised reproduction of protected works in AI outputs may constitute separate infringement.
The EU AI Act's territorial reach is extraordinary: copyright obligations apply to models offered in the EU regardless of where they were trained, forcing US-based companies to reconcile fair use assumptions with European opt-out mandates.
Legislative Signals: Towards Mandatory Compensation?
In March 2026, the European Parliament adopted a non-binding resolution proposing a potential shift towards mandatory compensation schemes — including a proposed 5–7 per cent flat-rate licensing fee for AI training on copyrighted content. While non-binding, this resolution signals the direction of future legislative intent and has already influenced negotiating dynamics between AI companies and rights holders across the continent.
The Emerging Licensing Economy
Litigation and regulation are not the only forces reshaping the landscape. A structured commercial licensing market for AI training data has emerged with remarkable speed, driven by AI companies' desire for "clean provenance" data and rights holders' recognition that licensing — rather than litigation alone — offers a more sustainable path to compensation.
From Scraping to Live Access
The market has evolved rapidly. In 2023, there were approximately two publicly documented "live access" licensing deals — arrangements where AI companies pay for real-time, attributed access to content rather than one-time training data dumps. By mid-2026, that number had grown to an estimated 34. The shift reflects AI companies' recognition that real-time feeds, attribution, and grounding capabilities improve model accuracy and reduce hallucinations — making licensed content not merely a legal necessity but a technical advantage.
From two live-access licensing deals in 2023 to an estimated 34 in 2026 — the AI content licensing market has not merely grown; it has structurally transformed.
News and journalism remain the primary drivers of the licensing economy, accounting for 48 of the 91 public deals tracked by mid-2026. AI firms value the constantly refreshed, real-time nature of news content, and major publishers have recognised that their archives and live feeds represent a significant commercial asset in the AI era.
Music, Images, and Multimodal Expansion
Licensing has expanded well beyond text. The music industry has seen a cascade of settlements and agreements following initial litigation against generative audio services. Major record labels — Universal Music Group, Sony Music Entertainment, and Warner Music Group — have struck deals with generative AI services including Klay, Suno, and Udio. These agreements typically include opt-in frameworks giving artists control over whether their compositions are used for training, and commitments by AI companies to retrain models on exclusively licensed content.
In the visual domain, Getty Images' litigation against Stability AI has driven a market-wide shift towards "clean provenance" models, where images from licensed sources command a premium. The case has provided copyright holders with significant leverage in negotiations, and the broader visual arts sector has moved towards structured licensing frameworks that distinguish between training use, output generation, and commercial deployment.
The Collective Rights Management Question
One of the most significant structural questions in the emerging licensing market is whether individual licensing — negotiated deal by deal between AI companies and rights holders — is scalable, or whether collective rights management organisations (CMOs) will need to play a central role. The French music rights organisation SACEM has already exercised opt-out rights over its entire repertoire, requiring explicit authorisation and financial negotiation for any AI training use. Similar moves by other CMOs could fundamentally reshape the economics of AI training data, particularly for smaller AI developers who lack the resources to negotiate thousands of individual agreements.
From two live-access licensing deals in 2023 to an estimated 34 in 2026 — the AI content licensing market has not merely grown; it has structurally transformed.
Cross-Jurisdictional Tensions and the Path Forward
The interaction between the US litigation-driven model and the EU's structured regulatory framework creates significant friction for global AI providers. A company that trains a model in the United States under fair use assumptions, then deploys it in the EU, may find itself simultaneously compliant with US law and in violation of EU regulation. The EU's extraterritorial reach — asserting copyright obligations over any model offered in the EU market — effectively exports European standards to global AI development.
This dynamic is accelerating convergence. US-based AI companies are increasingly adopting EU-compliant practices not because they are legally required to do so in the United States, but because the cost of maintaining separate training pipelines for different jurisdictions is prohibitive. The practical effect is that the EU's more stringent framework is becoming the de facto global standard for AI training data governance.
What Rights Holders Must Do Now
For rights holders — whether individual creators, publishers, or collective management organisations — the current moment demands active engagement rather than passive reliance on legal protections. Machine-readable opt-outs must be implemented and maintained; generic natural-language reservations are legally insufficient under EU standards. Rights holders should audit their digital presence to ensure that opt-out signals are correctly implemented across all platforms and content repositories.
At the same time, the emerging licensing market offers genuine commercial opportunity. Rights holders who engage proactively with AI companies — offering structured access to high-quality, attributed content — are better positioned to capture value from the AI training economy than those who rely solely on litigation or opt-out mechanisms.
What AI Developers Must Do Now
For AI developers, the compliance landscape demands a fundamental shift in data governance practices. Training data provenance must be documented with sufficient detail to satisfy EU AI Act transparency requirements. Machine-readable opt-outs must be respected systematically, not selectively. And the economics of licensing — once viewed as an unnecessary cost — must be integrated into AI development budgets as a standard line item.
The companies that will thrive in this environment are those that treat copyright compliance not as a legal constraint to be minimised, but as a quality signal: licensed, attributed, high-provenance training data produces better models, reduces legal risk, and builds the trust relationships with rights holders that will be essential for sustained access to the content that powers AI.
Conclusion: The Architecture of a New Settlement
The legal reckoning over AI training data is not a temporary disruption — it is the construction of a new institutional settlement between the AI industry and the creative economy. That settlement is being built simultaneously in courtrooms, regulatory offices, and commercial negotiating rooms, and its contours are becoming clearer with each ruling, each regulation, and each licensing deal.
The settlement that is emerging is not a simple victory for either side. Rights holders are gaining recognition, compensation mechanisms, and regulatory leverage they did not have three years ago. AI developers are gaining legal clarity, access to high-quality licensed content, and a more sustainable foundation for long-term development. The transition is painful and expensive, but the direction is clear: the era of unchecked scraping is ending, and a structured, rights-respecting AI training economy is taking its place.
The question is no longer whether AI companies will pay for the data that powers their models. The question is how much, to whom, and through what mechanisms. Those questions are being answered now — and the answers will shape the economics of artificial intelligence for a generation.





