Walling Off the Web: How Cloudflare’s Plan to ‘Save’ Content Could Break the Internet
The economics of the open web are changing at a breathtaking pace. For two decades, the implicit bargain was simple: creators produce content, and in exchange, search engines like Google send them traffic, which can be monetized through ads. That model, for all its flaws, funded a vast and diverse ecosystem of information. But the rise of AI “answer engines” has shattered this quid pro quo. Why click a link when an AI can give you the answer directly?
Into this crisis steps Cloudflare, one of the most powerful and pivotal infrastructure companies on the internet. With its new “pay-per-crawl” initiative, CEO Matthew Prince has positioned Cloudflare as the architect of a new business model for the web—one where AI companies must pay for the data they consume.
On the surface, the goal is noble. In a recent interview with Stratechery, Prince laid out a stark choice: either journalists and creators starve, the entire media landscape is co-opted by a handful of powerful AI patrons like the Medicis of old, or we figure out a new way for creators to get paid. Cloudflare’s solution is to create scarcity by blocking AI crawlers, thereby forcing AI companies to the negotiating table.
The problem, however, is that in its zeal to build this new model, Cloudflare is setting a precedent that could irrevocably damage the very fabric of the open internet. As Mike Masnick of Techdirt has powerfully argued, we are cheering on the construction of walled gardens in the name of fighting AI, and in the process, we risk breaking everything that made the web great. This is a profound mistake, and one that deserves intense scrutiny.
The Promise: A New Deal for Creators
Matthew Prince’s diagnosis of the problem is compelling. The world is rapidly shifting from search engines that provide a “treasure map” of links to answer engines that deliver a finished product. This shift severs the connection between content creation and traffic, gutting the business model that supported the web. Data from Cloudflare’s own network shows it is becoming exponentially harder to get a click from an AI than it ever was from Google search.
Publishers are feeling this acutely, telling Cloudflare they are “dying.” In response, Cloudflare is leveraging its position as a gatekeeper for a vast portion of the internet to enforce a new rule: if you want to crawl our customers’ content to train your AI, you need to pay for it. The logic is simple market economics: introduce scarcity to create value.
For many beleaguered publishers, this sounds like a lifeline. But the implementation of this vision is where the promise begins to unravel, revealing a fundamental misunderstanding—or perhaps a willful redefinition—of how the web is supposed to work.
The Peril: Proprietary Standards & Conflating Crawlers with Users
The first critical flaw in Cloudflare’s approach, as highlighted by Luke Hogg and Tim Hwang; is its sharp break from the web’s long-standing tradition of open, cooperative standards. For decades, the internet has operated on a voluntary protocol known as robots.txt. This simple text file allowed site owners to signal their preferences, and well-behaved crawlers would respect those wishes, embodying the cooperative spirit of the web without needing a central authority.
Cloudflare’s solution abandons this decentralized model. Instead of relying on open standards, it uses its proprietary network to unilaterally shut out automated visitors with a “one-click” kill-switch.
The second critical flaw in Cloudflare’s approach, as highlighted by Mike Masnick on Techdirt, is its dangerous conflation of two very different activities: bulk scraping for AI training and user-directed queries.
It is one thing for a website owner to block a bot from systematically downloading their entire site to feed a massive AI model. That is a choice about how one’s content is used in bulk. It is another thing entirely to block a technological tool from accessing a single, publicly available URL on behalf of a specific user to answer a specific question. The latter is not a bot “crawling” the web; it is a user “browsing” it with a modern tool.
This distinction came into sharp focus when Cloudflare publicly accused AI company Perplexity of using “stealth, undeclared crawlers to evade website no-crawl directives.” Cloudflare described an experiment where they set up secret domains with robots.txt files forbidding crawling, then asked Perplexity questions about those domains. When Perplexity provided answers, Cloudflare cried foul.
But as Masnick points out, this test didn’t prove Perplexity was illicitly scraping. It proved that when a user asks a question, Perplexity’s tool goes to the public URL to find the answer—which is precisely how a browser works. Cloudflare is attempting to expand the definition of robots.txt from a directive for mass crawlers to a restriction on individual user access via intermediary tools. This is a radical and dangerous shift. It breaks the web’s fundamental promise: if content is made public, people should be able to access and use it.
The Collateral Damage of Walled Gardens
This aggressive posture is already causing significant and predictable collateral damage, creating a chilling effect across the open web.
1. The Death of the Archive: Archival services are the first victims. As Masnick reports, Reddit has begun blocking the Internet Archive’s crawler to prevent AI companies from accessing its content for free via the Wayback Machine. In the quest to monetize fresh data for AI training, we are actively destroying the historical preservation of human knowledge and online culture. Similarly, Common Crawl, a non-profit that provides crucial web archives for academic and public-interest research, is finding itself shut out from vast swathes of the internet, neutering its ability to create a public good simply because AI companies also found its archives useful.
2. Breaking Accessibility and Utility: Where does this blocking end? As Masnick correctly asks, if we normalize the idea that any technological intermediary can be blocked, we put essential accessibility tools at risk. Many visually impaired users rely on screen readers and other software to “read” the web for them. Under this new paradigm, are they next to be blocked? Furthermore, it breaks legitimate and helpful uses of AI. Masnick gives a personal example of using an AI tool to help edit his articles by having it read his source material to check for accuracy—a task that is now impossible with publishers who have walled off their content.
3. A Two-Tier Internet: The most likely outcome of this strategy is not a renaissance for small creators. Instead, it will be a “two-tier internet” where large platforms and media conglomerates strike multi-million dollar licensing deals with big AI companies, while everyone else—researchers, journalists, archivists, startups, and individual users—is locked out. We are not democratizing access; we are cementing the power of the largest incumbents on both sides of the transaction, creating a web that is less open and less accessible for all.
The Irony of Power
There is a deep irony in Cloudflare leading this charge. For years, the company has been a reluctant arbiter of content, carefully navigating complex moderation issues and often stating its desire not to be the judge of what is good or bad online. In his interview, Prince draws a line: Cloudflare wants the power to stop cyberattacks, but not the power to be an “editor.”
Yet, what is this “pay-per-crawl” system if not a form of editorial control at the most fundamental level? By building the tollbooths, policing the traffic, and deciding which bots get access under what terms, Cloudflare is stepping squarely into the role of a market-making gatekeeper—a position of immense power. While Prince frames this as merely facilitating a market, the act of creating and enforcing the rules of that market is a profound exercise of power over who gets to access information and on what terms.
The crusade to monetize AI may have started with good intentions, but it is paving a path toward a fragmented, expensive, and closed internet. The solutions to creator compensation must be found in ways that preserve the web’s core principles of openness and accessibility, not ones that sacrifice them.
As Mike Masnick concludes, “The web’s core principle wasn’t ‘open to everyone except the technologies we don’t like.’ It was ‘open, period.'” In its effort to cure the economic disease ailing online content, Cloudflare is prescribing a medicine that may kill the patient.
Written by Caroline De Cock, LL.M. , Head of Research.
