AI Training and EU Copyright: Is It Legal? A Deep Dive into the TDM Exception
A central legal question in the discourse surrounding GenAI is whether training models on copyrighted data is lawful. At EU level, certain stakeholders are pushing a narrative whereby Text and Data Mining (TDM) is “something different from AI,” and therefore the TDM exceptions in the EU’s Copyright Directive do not (fully) apply to AI training.
In reality, AI model training is fundamentally a form of TDM, and a careful analysis shows the law covers such innovative uses.
The Law, The Tech, and The Inarguable Link
The argument is built on a direct mapping of the technical process of AI training onto the legal definition of TDM.
1. The Legal Definition of TDM
The EU’s Copyright Directive (2019/790) defines TDM broadly in Article 2(2) as:
“…any automated analytical technique aimed at analysing text and data in digital form in order to generate information which includes but is not limited to patterns, trends and correlations.”
This definition is technology-neutral and focuses on the process of automated analysis to generate information like patterns. The Directive then creates exceptions for this activity in Articles 3 (for scientific purposes) and 4 (for any purpose, subject to an opt-out for rightsholders).
2. The Technical Process of AI Training
To determine if AI training falls under this definition, it is necessary to understand the technical process. Training a large AI model, such as an LLM, is a multi-stage computational process designed to enable the model to learn from vast quantities of data.
- Data Ingestion and Preprocessing: The process begins with the collection of massive digital datasets (e.g., text, images). This raw data is unstructured and must be prepared for analysis. This involves a series of automated preprocessing steps, such as cleaning (removing inconsistencies), normalization (standardizing formats), and tokenization (breaking text into words or sub-words). This stage transforms unstructured text and data into a structured format suitable for computational analysis.
- Learning Patterns and Correlations: The preprocessed data is then fed into a neural network architecture (like the Transformer). The core of the training process involves the model’s algorithms analyzing this data to identify statistical relationships, patterns, and correlations. For a language model, this means learning the probabilities of how words and concepts typically appear together in context. For an image model, it involves learning the relationships between pixels and associated text descriptions.
- Generating Information (The Trained Model): The direct output of this analytical process is not a copy or a database of the original training data. Instead, the process “generates information” in the form of the trained model itself, which consists of a complex set of billions or trillions of numerical values known as “weights” or “parameters”. This set of numbers is a highly compressed, mathematical representation of the patterns, concepts, and correlations learned from the input data. The model does not store the original works; it adjusts its internal parameters to reflect the patterns it has identified.
3. Mapping the Technical to the Legal: Why AI Training IS TDM
When the technical process of AI training is mapped onto the legal definition of TDM, the alignment is direct and unambiguous.
- AI training is an “automated analytical technique”: It uses complex algorithms to computationally analyze data without direct human intervention at each step.
- It is “aimed at analysing text and data in digital form”: This is precisely what the ingestion and preprocessing stages accomplish.
- It is done “in order to generate information”: The output is the trained model’s set of weights and parameters.
- This information “includes but is not limited to patterns, trends and correlations”: The model’s weights are the very embodiment of the statistical patterns and correlations discovered in the training data.
The counterargument that TDM is for analysis while AI training is for synthesis (i.e., creating new content) is legally and technically flawed. This argument incorrectly conflates the process of training with the potential use of the trained model.
The TDM exceptions in the Copyright Directive apply to the acts of reproduction and extraction undertaken for the purpose of mining—that is, the training process itself. As Knowledge Rights 21 explains, “as long as the computer model analyses copyright works (as an automated analytical technique) to generate an output the exceptions will apply”. Whether the resulting model is later used to classify data (a traditional AI task) or to generate new text (a generative AI task) is irrelevant to the legality of the training process under the TDM exceptions. The law concerns itself with the act of mining, not the future applications of the knowledge gained from it.
Corroborating Evidence: Legislative Intent and Judicial Rulings
This interpretation is not a loophole; it is confirmed by the EU’s own legislative record and emerging case law.
- Legislative Intent: The idea that the Copyright Directive’s TDM exceptions were not devised with AI in mind is plainly false. Statements from the European Parliament at the time highlighted that the provisions were introduced “in order to contribute to the development of data analytics and artificial intelligence”. Some Member States in the Council of the EU similarly framed the Article 4 exception as a tool to “unleash the potential of artificial intelligence”.
- The EU AI Act: The AI Act creates an undeniable link. Recital 105 acknowledges that AI training uses TDM techniques and directly references the exceptions in the Copyright Directive. Article 53 goes further, explicitly requiring AI providers to respect the opt-out mechanism from Article 4 of the Copyright Directive. This provision would be meaningless if the legislature didn’t believe the TDM exception was the relevant framework for AI training.
- Court Rulings: This view has been endorsed by the judiciary. In a landmark 2023 decision, the Regional Court of Hamburg (Germany), ruling on a case involving the LAION-5B dataset used for training image models, affirmed this interpretation. The court dismissed a copyright claim from a photographer and held that the creation of datasets for AI training is an activity subject to the TDM exemption, considering that the creation of data sets, which enable AI training, is TDM under the law. The court explicitly rejected the argument that lawmakers didn’t have GenAI in mind, reasoning that they aimed to include fields of application that they were not aware of at the time.
The EU framework strikes a deliberate balance: it permits the innovation inherent in training AI models while retaining traditional copyright remedies against infringing outputs. Viewing the TDM exceptions as a “loophole” being exploited by AI developers ignores the deliberate policy choices behind their creation. The exceptions in Articles 3 and 4 of the Copyright Directive were not accidental. They were crafted to foster research and innovation in an increasingly data-driven European economy, recognizing that analyzing large datasets is a prerequisite for progress. AI training is the quintessential example of the large-scale data analysis these exceptions were designed to enable.
To argue that “learning” from data through computational analysis is not covered by the TDM exceptions is to argue for the creation of a new, de facto exclusive right for creators: the right to control the analysis of their publicly available works. This would be a radical expansion of copyright that runs contrary to the Directive’s stated goals and would have a profound chilling effect on the entire European data economy, far beyond the realm of GenAI.
Written by Caroline De Cock, LL.M. , Head of Research.
