Subscribe on LinkedIn

From Monkeys to Dolphins: 2025, The Year Courts Hallucinated Copyright Infringement

As 2025 draws to a close, it is the perfect time to look back at the year’s defining legal copyright obsession: the struggle to define exactly what an AI “knows.”

We previously published A Tale of Monkeys, Memorisation, and Misunderstandings, inspired by a paper by legal scholar Aline Larroyed, The Fallacy of the File. We argued that if a monkey accidentally types Shakespeare, suing the primate for copyright infringement is a category error. We warned that treating statistical probability as “theft” and technical glitches as “memorisation” would lead to bad legal interpretations.

This is sadly not a hypothetical scenario. Back in spring, a case surfaced in Hungary that perfectly encapsulated this confusion. It involved a pop star, a freshwater lake, and a bold plan to import marine mammals.

The case was Like Company v. Google Ireland Limited (Case C-250/25). When the Budapest Regional Court referred this dispute to the Court of Justice of the European Union (CJEU), it marked the moment the “misunderstanding” of how AI works stopped being a Twitter debate and started threatening the legal infrastructure of the European internet.

The Dolphin in the Data Lake

To understand the confusion of 2025, you have to understand the story that started it.

On July 21, 2023, the Hungarian news portal balatonkornyeke.hu published a story that was strange even by the standards of celebrity gossip. The article, titled “Kozsó doesn’t give up…”, detailed the eccentric plans of singer Zsolt Kocsor to introduce dolphins into Lake Balaton—a freshwater lake ecologically unsuitable for them.

The publisher, Like Company, discovered that Google’s Gemini chatbot could answer questions about facts included in the article, including a detailed summary of the dolphin plans.

Like Company sued. Their argument was seductive in its simplicity:

  1. Ingestion: The AI “read” our article.
  2. Storage: The AI “memorised” our article (reproduced it).
  3. Rebroadcast: The AI “communicated” our article to the public.

Therefore, the AI is a pirate.

But, just like the lawyer serving a subpoena to a monkey, this lawsuit relies on a “technically fallacious anthropomorphism“. It assumes the AI is a digital library storing a copy of the dolphin story.

The “Library” Fallacy: Why Vectors Aren’t Books

The central error of 2025 was the belief that training an AI is the same as saving a PDF to a hard drive. As the analysis of Case C-250/25 showed us, the plaintiff conflated Information Processing with Reproduction.

We spent much of this year explaining that when an AI “trains” on a text, it doesn’t store the text. It destroys it.

  • Destruction of Syntax: The process begins with tokenization. The AI breaks the text into integers, discarding the font, layout, and “expression” immediately.
  • Ontology of the Copy: These integers are mapped to embeddings—vectors in a high-dimensional space. The model does not store the sentence “Kozsó loves dolphins.” It stores the statistical probability that the concept of “Kozsó” is semantically close to “dolphins” and “Balaton”.

To claim that these mathematical weights are a “reproduction” of the original article is like claiming a weather forecast is a “reproduction” of last year’s rain. It is a model of the phenomenon, not the phenomenon itself.

The “RAG” Smoking Gun

Looking back, the most ironic twist of the Like Company case was the factual anomaly that blew the “illegal training” argument out of the water.

The lawsuit alleged the AI “learned” the article during training. However, the article was published on July 21, 2023. The lawsuit claimed infringement began in June 2023—more than a month before the article was written.

If Gemini didn’t train on the article, how did it know about the dolphins?

It almost certainly used Retrieval Augmented Generation (RAG). Put simply: it Googled it. It acted as a search engine, found the live article, read it in real-time, and distilled the facts.

This destroyed the “Training is Reproduction” argument for this specific case. The case is actually about whether a machine is allowed to read a webpage and summarise the facts found therein. And if the CJEU says no, then the Internet and our society as a whole  is in trouble.

The Three Legal Battlegrounds

While the “dolphin” story was the catalyst, the legal war of 2025 was fought on three specific fronts.

1. The “Reproduction” Trap (Article 2)

The plaintiff argued that loading a file into GPU memory (VRAM) for processing constitutes an infringing copy. The defense argued these are “transient” acts essential to a technological process.

The Danger: If the CJEU agrees that “learning” (weight adjustment) is reproduction, it implies the internal state of a machine is an infringing copy, effectively regulating the act of “reading” itself.

2. The “Snippet Tax” (Article 15)

Article 15 was designed to tax search engines for displaying “snippets.” But Gemini wrote a comprehensive summary.

The Value Gap: Copyright protects expression, not facts. If the court rules that summarising facts is infringement, it grants publishers a monopoly on the news itself, preventing AI from doing what human journalists do every day.

3. The “Machine-Readable” Trap (TDM Opt-Outs)

EU law allows data mining unless the rightsholder opts out in a “machine-readable” manner. Like Company argued that their natural-language “Terms of Service” should count because the AI can “read”.

The Reality: This anthropomorphises the crawler. Crawling relies on standardised protocols (like robots.txt), not an AI analysing the legal footer of billions of websites. If natural language counts as “machine-readable,” compliance becomes impossible.

The Anatomy of the Misunderstanding

The following table gives an overview of the cognitive dissonance that defined the legal landscape this year:

Table 1: The Plaintiff’s Metaphor vs. The Engineering Reality

The Global Context: A Fractured Landscape

The CJEU’s pending decision has left Europe sandwiched between a permissive US, a pragmatic UK, and a restrictive Germany.

Table 2: Comparative AI Copyright Jurisprudence

Conclusion: Swimming in Probabilistic Waters

The Like Company lawsuit was built on a coherent, internal logic, but it was a logic of the 19th century applied to the 21st. It viewed the AI as a thief, the training set as a stolen library, and the output as a bootleg broadcast.

But as we look back on 2025, it is clear these metaphors led us astray. The “Dolphin” case asked the CJEU to expand copyright law to cover the statistical analysis of information.

We are moving from an era of Information Scarcity (where copying was difficult) to an era of Information Ubiquity (where generation is free). We cannot hold back this tide with the levee of old copyright definitions.

As we enter 2026, let’s hope we stop looking for the monkey in the machine. There is no monkey. There is no file. There are only probabilities. And if we outlaw probabilities, we might as well outlaw the weather forecast. But then again, being Belgian, my weather forecast feels like a continuous hallucination and outlawing it might make sense.


Get the article’s core concepts in a unique auditory summary—listen to the AI-generated track below!

This song was created entirely using artificial intelligence tools. If you enjoy this experiment, keep an eye out for each song illustrating our December articles.

Written byCaroline De Cock, LL.M., Head of Research