A Tale of Monkeys, Memorisation, and Misunderstandings
There is a famous thought experiment known as the Infinite Monkey Theorem. It suggests that if you give an infinite number of monkeys an infinite number of typewriters and an infinite amount of time, eventually—purely by chance—one of them will type out the complete works of William Shakespeare.
Now, imagine that the second the monkey types “To be, or not to be,” a lawyer kicks down the door and serves the primate with a subpoena.
“Aha!” the lawyer says. “You have memorised the Bard! You are holding a copy of the First Folio in your brain! Pay up, bubbles.”
This sounds absurd, but legal history shows us that dragging primates into court isn’t as far-fetched as it seems. We have already spent years debating whether a macaque named Naruto owned the copyright to a selfie he snapped on a stolen camera (spoiler: the courts ruled that animals can’t own copyright, much to the disappointment of monkeys everywhere looking to monetize their Wikipedia pages).
But while the “monkey selfie” case was a legal curiosity, the current attempt to regulate Artificial Intelligence carries much greater legal and policy implications. By anthropomorphizing software, we are effectively suing the mechanical monkey for “knowing” Shakespeare, when all it did was hit keys at random.
A new paper by legal scholar Aline Larroyed, The Fallacy of the File, exposes exactly how this metaphorical confusion is poisoning EU copyright law right now.
The Monkey in the Machine: It’s a Bug, Not a Feature
As Larroyed’s research highlights, the term “memorisation” has a very specific, boring definition in machine learning. It refers to a rare technical occurrence—a statistical artifact—where a model overfits and accidentally reproduces verbatim fragments of training data.
In computer science, this isn’t the goal; it is a failure of generalisation. When a model “memorises,” it fails to learn the underlying pattern and instead rote-repeats a specific string, usually because that data appeared too frequently in the training set. It is the digital equivalent of our monkey getting stuck on the ‘E’ key; it isn’t artistic expression or intentional copying, it is a glitch.
However, when that word travels from the computer lab to the courtroom, the nuance is stripped away. Suddenly, “memorisation” can be misread as equivalent to copyright reproduction (though not always).
Policymakers and judges assume that if an AI “memorises,” it must “contain” a copy of the work. They view Large Language Models (LLMs) as high-tech filing cabinets or databases full of stored JPEGs and PDFs.
Larroyed dismantles this understanding. She argues that treating an LLM as a database of works is a “category error”. There are no files in the computer. There are only parameters—billions of optimized numbers that represent statistical probabilities.
The Cost of Metaphorical Laws
Why does this linguistic mix-up matter? Because relying on the wrong metaphors creates dangerous legal precedents.
Larroyed warns that this “doctrinal drift” threatens to destabilize the balance of copyright law. The European Court of Justice (CJEU) has established that reproduction requires a work to be identifiable with “precision and objectivity”—in other words, it must be fixed in a medium.
AI training does not do this. It optimizes parameters to enable statistical generalisation; it does not retain identifiable copies of the works. By forcing the square peg of “statistical optimization” into the round hole of “copyright reproduction,” we risk breaking the regulatory framework.
The consequences, as outlined in The Fallacy of the File, are severe:
- Stifling Innovation: We risk creating compliance costs that crush SMEs and open science initiatives, handing the keys to the future to the few massive incumbents who can afford to pay for “licensing” data they shouldn’t need to license.
- Policy Misalignment: We are regulating the wrong layer. We are trying to police the training process (the math) instead of the outputs (where the actual harm might occur).
- Undermining TDM Exceptions: This drift threatens to undo the Text and Data Mining (TDM) exceptions carefully carved out in the EU’s Digital Single Market Directive, which were designed to allow exactly this kind of computational analysis.
The Way Out: Mechanism-Aware Regulation
So, how do we fix this? Larroyed suggests we need to move away from “metaphorical shortcuts” and toward “mechanism-aware regulation”:
- Abandon the Filing Cabinet: We must accept that models do not “contain” works. Training an AI is not storage; it is parameter optimization.
- Focus on Outputs, Not Weights: If the monkey types out a defamatory statement or a copyright-infringing passage, address that harm at the point of generation. Regulating the training data assumes a “leakage” that is technically rare and statistically avoidable.
- ata Governance, Not Bans: Instead of banning the act of learning, we should focus on robust dataset governance and deduplication to prevent the “overfitting” that leads to memorisation in the first place.
Conclusion: Stop Making a Monkey Out of Copyright Law
The “memorisation” debate is a perfect example of what happens when we let metaphors drive policy. We take a technical term, wrap it in a human narrative, and turn it into a legal weapon. It is the entire storyline of my forthcoming book AI Tools, Not Gods.
It is time to unenchant our view of AI. These models are not filing cabinets. They are not digital thieves. They are statistical tools.
We denied Naruto the macaque copyright for his selfie because he lacked human agency. It would be a strange irony if we now decided that AI models do possess enough agency to be treated as willful copyright infringers for “memorising” books they cannot read.
Let’s leave the Infinite Monkeys alone and focus on building governance and case law that reflects technical reality.
Get the article’s core concepts in a unique auditory summary—listen to the AI-generated track below!
This song was created entirely using artificial intelligence tools. If you enjoy this experiment, keep an eye out for each song illustrating our December articles.
Written by Caroline De Cock, LL.M., Head of Research
