The AI Training Debate: Why Focusing on Copyright Is Missing the Point
As public debates rage over AI’s impact on creative industries, one critical misunderstanding keeps surfacing: the assumption that AI models consume and store entire copyrighted works like a digital hoarder, waiting to reproduce them on command. This is not how AI training is supposed to work, not even close.
In reality, AI systems do not memorise and regurgitate full songs, novels, or scientific papers. They break down large volumes of information into patterns, structures, and statistical relationships using tiny fragments of data called tokens. The purpose of training is not to copy, but to learn how information is structured, so that the model can generate new, original content, or analyse and summarise facts more efficiently.
AI Is Learning Patterns, Not Hoarding Copyrighted Works
Let’s be clear: when AI models are trained on text, they are not normally storing entire copyrighted works in a library-like memory. Instead, they convert text into numerical representations, called tokens, that capture patterns in language, concepts, and factual relationships. These patterns are used to predict what comes next in a sentence, how ideas connect, or how complex systems behave.
This is fundamentally different from copying or reproducing a work. The process is closer to how a scientist reads hundreds of research papers over a career, internalises patterns in experimental results, and then applies that understanding to generate a new hypothesis. No one would claim that a scientist is infringing copyright simply by learning from what they’ve read.
AI operates on the same principle but at a vastly accelerated scale.
Copyright Infringement by AI Outputs Is Already Covered by Existing Laws
It’s important to recognise that when generative AI outputs actually cross the line into copyright infringement, existing intellectual property laws already apply. There is no legal vacuum here. If an AI system produces a work that substantially copies or replicates a copyrighted song, film scene, book passage, or any other protected material, that output can be challenged under current copyright frameworks.
This is the same legal standard applied to human creators. Whether infringement is the result of a person or a machine, the decisive factor is whether the final output unlawfully reproduces protected content, not how the learning or creative process occurred.
What this means is clear:
- If a generative AI tool outputs a song that mimics a specific copyrighted melody, the rights holder can pursue legal remedies.
- If an AI-generated image is substantially similar to a copyrighted photograph or artwork, infringement laws apply.
- If AI produces a near-verbatim excerpt from a protected book or article, that too is covered under existing statutes.
In short, we don’t need to reinvent copyright law to address infringement by AI outputs: the legal tools already exist. What’s needed is to apply these laws sensibly to ensure that genuine cases of infringement are addressed, while avoiding blanket restrictions that stifle legitimate, transformative uses of AI.
Facts and Data Are Not Copyrighted
This is a critical legal and conceptual point: copyright does not, and has never, protected facts, ideas, or data. Copyright only covers the specific expression of those ideas, not the underlying truths themselves.
When AI models analyse scientific literature, they aren’t “stealing” the expression of those works. They’re extracting factual information such as chemical structures, statistical relationships, historical events, and biological processes, all of which exist independently of how any one author describes them.
This distinction is why Text and Data Mining (TDM) has been an established, legitimate practice in science for decades. Researchers routinely analyse vast corpora of academic papers to uncover hidden patterns, trends, and insights, without seeking individual permissions from every publisher or author. It’s how major scientific breakthroughs happen, and it’s widely accepted as a necessary part of research.
Generative AI Is Simply the Next Iteration of TDM
What’s changed is not the principle, but the scale and efficiency of this analysis. Generative AI has supercharged TDM, enabling models to process orders of magnitude more data, faster than any human team ever could. But the underlying activity, namely analysing patterns across large datasets, is nothing new.
The current calls to clamp down on AI training miss this critical continuity. To restrict AI training is, in effect, to reverse decades of progress in data-driven science. And in doing so, we risk:
- Slowing critical advancements in medicine, climate science, and technology
- Erecting legal barriers to open scientific inquiry
- Concentrating AI innovation in the hands of a few corporations able to afford costly data licensing deals.
This is not a theoretical risk; it’s a practical, immediate threat to the future of global research.
Smarter Regulation Starts with Understanding the Basics
Before imposing sweeping restrictions on AI training, policymakers must ask: Do these rules reflect how AI actually works? Or are they based on outdated assumptions about copying and copyright?
If we legislate based on the false premise that AI models are “hoarding” copyrighted works, we risk creating laws that choke off innovation while failing to address real concerns about harmful outputs, such as misinformation or deepfakes.The focus of regulation should be on how AI outputs are used, not on preventing AI systems from engaging in the same pattern analysis that humans, and especially researchers, have always done.
Written by Caroline De Cock, LL.M. , Head of Research.
May 19, 2025
