Subscribe on LinkedIn

Brussels Revisits Copyright –  Part 2: Why Mandatory AI Licensing Is Not a Silver Bullet And Could Ricochet

This is the second installment in a three-part series analysing the European Commission’s current call for evidence on EU copyright rules. Part 1 examined live content piracy and the fundamental rights implications of enforcement without safeguards. This post focuses on text and data mining (TDM) and the push for mandatory AI training licences; the third piece will explore the research exception.

The European Commission’s call for evidence to support its review of the Directive on Copyright in the Digital Single Market (DCDSM) and to lay the foundations for a targeted legislative proposal aimed at strengthening copyright in light of AI and other market developments, which closes on 25 June, asks whether the existing copyright framework adequately addresses AI and TDM, and whether new measures are needed.

Proposals for mandatory licensing of AI training data, compulsory collective licensing, and new remuneration rights for rightsholders are being presented to the Commission as measures to protect creators. They rest on two factual claims that are both testable and have now been tested: that restricting training data to licensed material is technically manageable without significant capability loss, and that a remuneration mechanism would reach the individual creators named as its intended beneficiaries. Peer-reviewed research published in 2025 has addressed the first claim. Five years of the Article 15 DCDSM press publishers’ right data have addressed the second.

The framework already answers this question

Articles 3 and 4 of the DCDSM already provide the legal basis for TDM, including AI model training on lawfully accessed content. The AI Act confirmed this interpretation in Recital 105 and Article 53(1)(c). The General-Purpose AI Code of Practice process presupposes it. The Commission’s own legislative architecture has already answered the question that mandatory licensing proposals purport to raise.

As with the live piracy enforcement examined in Part 1, the argument being put to the Commission is not that existing tools have demonstrably failed. It is that the existing framework should be replaced with a new mechanism whose consequences, on the available evidence, are likely to be worse than the problem it claims to solve.

What the data shows

Mandatory licensing advocates tend to frame data restriction as a toll road: the training run continues, a licence fee changes hands, and the only difference is that rightsholders receive compensation. The peer-reviewed evidence suggests the metaphor is wrong. Restricting training data does not just change who pays. It changes what the model can do.

Research by the Apertus consortium found that applying licence filtering to post-training fine-tuning data caused a 51% decrease in Massive Multitask Language Understanding (MMLU) chain-of-thought evaluation performance on the same underlying base model, dropping the score from 0.513 to 0.253. A separate Norwegian language model study found a 21% drop in language test performance and a 23% drop in world knowledge tests when models were trained on smaller datasets that excluded content accessible under current TDM rules. Both studies control for the relevant variables. Both reach the same conclusion.

The performance losses are concentrated in reasoning capabilities: multi-step inference, contextual judgment, and domain generalisation. Those are precisely the capabilities that matter most for professional and research applications. A model that performs at half its previous level on complex reasoning is not a more limited version of the same tool. It is a qualitatively different, less useful tool.

Cross-country analysis by Peukert (2025) connects this to innovation outcomes. Jurisdictions with broader copyright exceptions show consistently higher levels of AI innovation across publications, code, patents, and ventures, with the gap widening through the late 2010s and 2020s. The Commission is pushing an AI competitiveness agenda. Narrowing the TDM exception would work against it.

Why remuneration will not reach creators

The performance argument is, in principle, separable from the question of who benefits from a mandatory licensing regime. A proposal could have negative affects on AI capability and still be defensible if the remuneration it generated genuinely reached individual creators. The EU has a recent and detailed record that bears directly on that question.

Article 15 of the DCDSM introduced a press publishers’ right whose stated purpose was to direct revenue from large platforms to publishers and, through them, to the journalists and other creators who produce the content those platforms aggregate. Five years of operation have produced no verifiable benefit for individual journalists or creators. The money collected has not demonstrably reached the people the right was designed to protect. The mechanism collected. It did not meaningfully distribute.

The structural conditions for an AI training remuneration mechanism are worse than those that produced the Article 15 failure, not better:

  • Associate Professor Quintais of IViR has noted that title-by-title attribution across datasets of 12 to 36 trillion tokens is computationally infeasible.
  • When a dataset contains billions of works, the economic contribution of any individual work to any individual model is analytically unquantifiable at the individual level.
  • Administrative overhead would consume a pool that is already divided into fractions too small to be meaningful to any single creator.

The mandatory licensing argument is, at its core, a claim that creators deserve compensation for a new form of use of their work. That proposition is worth taking seriously. What it does not demonstrate is that a mandatory licensing mechanism is the instrument that achieves it. Nor that copyright law is the avenue to pursue.

What the Commission should now decide

The call for evidence asked whether the existing framework adequately addresses AI and TDM. The evidence says yes, provided it is clarified, applied, and defended. Five conclusions follow.

  1. Defend the existing framework: Articles 3 and 4 of the DCDSM, confirmed by the AI Act, provide the legal basis for TDM, including AI model training on lawfully accessed content. The Commission has already settled this question in its own legislation. No new licensing regime is needed or justified on the available evidence.
  2. Reject mandatory licensing: The empirical record on both model performance and creator benefit supports this conclusion. A mechanism that damages AI capability and fails to reach individual creators does not warrant legislative adoption, regardless of the strength of the lobbying case made for it.
  3. Keep transparency requirements proportionate: Aggregate-level summaries of training data types and sources, as envisaged in the AI Act, are achievable and appropriate. Title-by-title disclosure is computationally infeasible and functions in practice as a de facto licensing mandate. Any transparency obligation that can only be complied with under a licensing regime is a licensing regime.
  4. Protect Article 3 from the opt-out process: The machine-readable opt-out discussion must not function as a back-channel route to restricting research TDM. The Commission must treat the opt-out process and the DCDSM review as connected questions. Strengthening transparency requirements on the opt-out side while purporting to defend Article 3 on the review side would produce a framework that says one thing and does another.

The available evidence

In Part 1 of this series, the structural problem was urgency without oversight: enforcement tools justified by the time-sensitivity of a live event, applied without the safeguards the law requires, producing harm that the Commission itself documented and Italy ignored. The problem here is different in mechanism but parallel in structure.

A proposal presented as a silver bullet for creator compensation that the evidence suggests will ricochet: damaging the AI capability of European models, failing to reach the individual creators it names as beneficiaries, and reproducing at AI training scale the democratic failure that five years of Article 15 press publishers’ right data have already put on the record. Whether data restriction is technically manageable and whether a remuneration mechanism would reach creators are factual questions. Both have been answered. Legislating on the basis that they have not been is a choice, not an oversight, and the Commission has enough evidence to make a different one.

Written by Caroline De Cock, LL.M., Head of Research