The Publisher Is Also a Data Broker. We Should Probably Talk About That.
Consider how a library works. You walk in. You go from shelf to shelf. You borrow. Nobody keeps records of what you read, once you bring the book back, which shelves you lingered at, or how long you spent in the periodicals room. That protection is not accidental: it is a founding principle of librarianship, enshrined in professional ethics and, in many jurisdictions, in law. The library is the room where curiosity is allowed to be private.
Now consider what happens when the library moves online. The catalogue is owned by a corporation that offers only a subscription model, with nearly no ownership option for the content. The journals live behind a platform. The library is a mere intermediary between its patrons and the publisher. The reference management tool, the repository, the peer review interface, the metrics dashboard: each one provided by the same multinational, each one logging every action you take. The room where curiosity was allowed to be private has become a room where every keystroke is recorded by a company whose other business is selling those records.
A psychologist named Eiko Fried found out what that looks like in practice when he submitted a data access request to Elsevier in December 2021. What came back was an email containing hundreds of thousands of data points stretching back years: his affiliations, his reviews, the requests for peer review he had declined, his IP addresses tracing back to his home, the precise moments he logged in, the websites he had visited, every article he had downloaded or viewed, every click recorded.
Researcher Bernhard Stahl, at the University of Groningen, read about this and identified the problem clearly. “They control the academic process with an infrastructure that serves their interests, not ours”, he told the university paper. “And they use the collected data to provide analytics services to whoever pays for them.”
Whoever pays for them. The rest of this post is about who that turned out to be, in the most extensively documented cases, and what it means for the policy conversations currently taking place in Brussels.
Publishing was always a business. Data brokerage is a different one.
Academic publishers have never been charities. The model of locking publicly funded research behind paywalls, charging institutions for content they already produced for free, and extracting labour from researchers who peer-review without payment is not new. The critique has been loud for decades. For most of that history, the scandal was about money. It is still about money, but the room has changed.
Profit margins in scientific publishing sit above 20 percent, with Elsevier/RELX reaching 37-40% and announcing a 10% increase in profits for 2024. The latter is higher than leading tech firms such as Google or Apple, and enables the distribution of hundreds of millions of euros to shareholders and top executives every year. Substantial parts of those profits derive, directly or indirectly, from public investments in higher education. In its 2021 annual report, RELX/Elsevier declared revenues from university library subscription fees and open access “read and publish” agreements of close to 2 billion euros.
And RELX is not alone. In the first half of 2025, Springer Nature reported operating profit of €241 million on revenue of €926 million, a margin of 26 percent. The company attributed growth partly to open access momentum and partly to what it described as a sustained focus on “developing artificial intelligence tools to transform the publication process, provide more value to communities and create new revenue streams.” Open access, in other words, is the growth vehicle. AI tools are the next ones. The subscription wall has not been replaced by a more equitable model: it has been replaced by a data extraction model that is still being built.
Scientific publishers are, to a certain extent, following the playbook of large tech companies by expanding their business models towards data analytics. As part of that strategy, they have acquired companies occupying crucial positions throughout the research cycle. RELX/Elsevier now owns Scopus (search and discovery), Mendeley (reference management), and Pure (research repositories), among others. Each platform is a data collection point. Platforms such as Pure facilitate not just the collection of download numbers but also the recording of personal information and behavioural data. The Young Academy Groningen and the Open Science Community Groningen have noted that some such platforms have at times operated through spyware.
The room where curiosity was protected has not simply moved online: it has been acquired, rebuilt, and wired for surveillance. And the corporation that did the wiring also owns a business whose clients include governments and law enforcement agencies.
RELX is not just Elsevier
Elsevier is a subsidiary of RELX Group. RELX also owns LexisNexis, whose core business has little to do with academic journals.
LexisNexis’ marketing materials describe a database of 1.4 billion unique online digital identities from 4.5 billion devices. Its ThreatMetrix for Government product describes the ability to provide governments “a holistic, singular view of your citizens.” LexisNexis also maintains a “Special Services” division specialising in counterintelligence and investigative solutions.
According to SPARC’s analysis of vendor data privacy practices, RELX’s “risk” business, which provides services to corporations, governments, and law enforcement agencies based on expansive databases of personal data, has surpassed its Elsevier division in revenue and profitability.
The same corporate group that records every click on ScienceDirect is also, through a different subsidiary, selling access to profiles of over a billion individuals to government agencies worldwide. The data may flow through different business units. It flows through the same corporate structure, is governed by the same board, and serves the same shareholder interests. The room where the researcher sits, and the room where the government analyst queries a deportation list, are connected by a corporate parent.
Thomson Reuters, parent company of Westlaw, operates a comparable corporate structure. According to research published in the NYU Review of Law and Social Change, Thomson Reuters and RELX have built and maintained surveillance tools for local, state, and federal law enforcement entities, with surveillance representing a new source of revenue as traditional publishing and legal research become less lucrative. Both companies operate globally, with clients across multiple jurisdictions.
In the US: from the library database to the deportation flight
The most extensively documented case of where this infrastructure leads is in the United States, where both LexisNexis and Thomson Reuters hold government contracts connecting their data products directly to immigration enforcement.
According to a 2025 investigation by Prism Reports, LexisNexis’s Accurint tool reportedly provides more than 11,000 ICE agents access to analytics that “automate” decisions about vetting, screening, and targeting people for deportation, the vast majority of whom are people of colour. ICE’s own records suggest that LexisNexis conducts large-scale surveillance for civil immigration arrests.
Scholar Sarah Lamdan, writing in the NYU Review of Law and Social Change, documented a specific professional contradiction that emerged from these contracts: both Thomson Reuters and RELX Group hold contracts with ICE to supply the agency with personal data and analytic tools to mine that data, meaning that lawyers who publicly condemn ICE enforcement may not know that their legal research subscriptions contribute to the corporate infrastructure that enables it. The room in which the lawyer prepares their case against a deportation order and the room in which the deportation order was generated are, in this sense, furnished by the same company.
The library profession in the US also responded. The Library Journal reported that these contracts caught the attention of academic librarians and students nationwide. Lake Washington Institute of Technology declined to renew its LexisNexis contract after librarians concluded that the issue was a practical concern for their significant population of immigrant and first-generation students.
In 2021, the #NoTechforICE campaign succeeded in pressuring Thomson Reuters to terminate its collaboration with ICE. However, some of the damage was already done. Within days of the 2024 U.S. presidential election result, ICE was already moving to expand its surveillance apparatus again. A contract terminated does not delete the datasets. It does not change the corporate structure. It does not alter the incentive that made building the surveillance apparatus commercially attractive in the first place.
In the US, SPARC has noted that library subscription dollars, directly or indirectly, now support these surveillance systems. Every institution that holds a Westlaw contract, every university paying for ScienceDirect access, contributes to the corporate infrastructure that makes this possible. That should not be read as an accusation against librarians, researchers, or universities: they often have no viable alternative.
The academic freedom dimension this debate keeps missing
Back in Europe, the risks are at the moment less acute in the specific sense described above, but the structural problem is the same. Outsourcing the research infrastructure to commercial parties carries significant risks beyond individual privacy. Vendor lock-in allows companies like RELX/Elsevier to drive up prices for services such as Pure in an uncontrolled manner and to extend their monopoly position. More broadly, commercial companies are in a position to monetise public goods and academic futures: the Young Academy Groningen’s call for action flags the potential use of collected data as input for hiring and recruitment decisions, as one example of how this could work at the expense of staff, students, and ultimately taxpayers.
The academic freedom risk also has a global dimension that does not require ICE contracts to be real. A researcher anywhere who uses ScienceDirect to study government corruption, or who searches for literature on political dissent, leaves a data trail in the same private room that turned out not to be private. Elsevier’s privacy notice describes, as SPARC documented, the disclosure of detailed user data, including geolocation data, sensitive personal information, and inference data used to create profiles on individuals, both for wide-ranging internal use and to external third parties, including affiliates and business and joint venture partners.
In an era where governments are increasingly purchasing data products from private brokers to monitor dissidents and minorities, the question of who ends up paying is not a theoretical one.
This is also why the question of who controls research information infrastructure extends well beyond the campus. As Knowledge Rights 21 (KR21) has argued, the imperative is clear: research information must be open to ensure that it does not end up locked inside proprietary infrastructures, giving control over research assessment to those who control the data. The key question is how far commercial interests can be considered legitimate when control over research information undermines the achievement of policy goals.
The policy window Brussels keeps looking past
Brussels has a busy legislative season touching the knowledge ecosystem: the Digital Omnibus, the European Research Area (ERA) Act, the Data Act, and the forthcoming Digital Fairness Act. None has been directed at this problem.
As we have written in our coverage of the #ProtectOurFutureMemory campaign, the Digital Fairness Act offers one opening. That piece focused on unfair licensing terms: clauses that allow publishers to withdraw content unilaterally, override EU text and data mining rights by contract, and impose foreign legal frameworks on European institutions. All real and urgent. But the post-open-access surveillance architecture goes further. It is not primarily a licensing problem. It is a data infrastructure problem, and it requires a data infrastructure response.
The ERA Act, currently under negotiation, could include meaningful provisions on research assessment and the independence of research infrastructure. As KR21 has outlined, the proposed update to the Open Data Directive through the EU’s Digital Omnibus provides a timely opportunity to deliver meaningful change on research information openness, including obligations on publishers to release data produced through publicly funded research and to prevent that data from being locked into proprietary systems. Without it, the ERA Act risks passing without addressing who controls the platforms through which publicly funded research flows, and on what terms that data can be extracted and sold elsewhere.
The GDPR applies in principle to all of this. In practice, the enforcement gap between what publishers collect and what institutions can challenge remains wide. As SPARC’s analysis of ScienceDirect’s privacy practices found, the significant expertise and capacity required for any institution to understand even one vendor’s privacy practices creates a profound power asymmetry between vendors and libraries. Most university procurement offices do not have the resources to audit what a publisher’s privacy notice commits to. And as experience shows, the data analytics empires were built in the gap between principle and enforcement.
The room we handed over
We began with a library: a room where curiosity was private, protected by professional ethics and institutional practice built up over decades. We then watched that room move online, and we accepted, mostly without debate, that the digital version of it would be owned and operated by corporations with interests that diverge from our own. We paid subscription fees that funded the acquisition of more platforms, more data collection points, and more products sold to more clients. In the US, some of those clients turned out to be immigration enforcement agencies. Globally, the clients are whoever can pay.
The Young Academy Groningen and the Open Science Community Groningen have put five questions to their university board:
- what is the institution’s position on commercial control over the academic publishing cycle;
- what steps are being taken to reduce vendor dependence;
- what awareness has been created among staff about surveillance practices;
- what is the institution’s position on data ethics in relation to publisher agreements; and,
- most directly, to what extent is implicitly subsidising shareholder payouts reconcilable with the open science principles the university advocates?
These are worth putting to European institutions at scale, and to the EU’s current legislative process. The ERA Act, the Digital Fairness Act, and the Digital Omnibus together represent a genuine window. Whether that window is used to address the deeper infrastructure question, or confined to licensing clause reform while the surveillance architecture goes untouched, will determine a great deal about who the knowledge system serves over the next decade.
The room where curiosity was once protected belongs to someone else now. The door was handed over gradually, in procurement decisions and contract renewals, in platform migrations and open access negotiations, in each individual choice that seemed reasonable at the time. The ERA Act, the Digital Fairness Act, and the Digital Omnibus will either draw a line here or they will not. There is no neutral outcome: legislation that fails to address this actively ratifies it.
Get the article’s core concepts in a unique auditory summary—listen to the AI-generated track below!
This song was created entirely using artificial intelligence tools. If you enjoy this experiment, keep an eye out for each song illustrating our December articles.
Written by Caroline De Cock, LL.M., Head of Research
