Arrival
In March 2026, a film database maintained for thirty years vanished from the web. It wasn't brought down by human error or bankruptcy: it was crushed under the weight of millions of visitors who weren't people. The story of TheNumbers.com is an extreme case of something affecting the entire web: the assumptions on which it was built — that visitors are human, that traffic measures attention, that serving content generates reciprocal value — no longer hold.

TheNumbers.com, the film industry database founded in 1997 by mathematician Bruce Nash. In March 2026, it was taken offline by the combined weight of AI bot traffic and targeted cyberattacks. Only 10% of its visitors were human.
On March 5, 2026, TheNumbers.com vanished. No warning, no announcement, no maintenance page. It simply stopped responding. A site that had been online for nearly thirty years — founded in October 1997 by mathematician Bruce Nash as a Geocities page tracking 300 films — had grown into one of the most widely consulted film industry databases in the world: 78,396 movies, 178,375 theatrical release records, 236,176 people catalogued. Journalists, academics, Guinness World Records, and prediction markets used it as an authoritative source. Eight million visitors a year.
When the site reappeared on March 13, it was a skeleton. No historical charts, no individual movie pages, no Report Builder tool. Nash explained what had happened: AI bot traffic combined with targeted cyberattacks had made running the server untenable. Only 10% of visits were human. The remaining 90% were machines — training crawlers, AI agents, prediction market scrapers — consuming resources at a scale the infrastructure was never designed to handle.
It wasn't a conventional security problem. There was no single flaw to patch. It was something deeper: the assumptions on which that site — and, really, the entire web — had been built were no longer true.
A Space Built for People
In March 1989, a British physicist at CERN named Tim Berners-Lee circulated an internal document titled Information Management: A Proposal. His boss, Mike Sendall, wrote in the margin: "Vague but exciting." The proposal described a hypertext system for laboratory researchers to link documents to one another. It was not a plan to build a global infrastructure. There was no business model, no monetisation plan, not even the ambition that anyone outside CERN would use it.
That lack of commercial ambition was, paradoxically, what made the web possible. Berners-Lee has said many times that the decision not to patent the protocols was deliberate: he wanted an open system, permissionless, where anyone could publish and anyone could read. As he describes in Weaving the Web, radical openness was not an accident but a design condition.
Tim Berners-Lee's original proposal for the web (March 1989). His boss Mike Sendall's annotation — "Vague but exciting" — is visible at the top. No business model, no monetisation plan: just an open system for linking documents.
But that openness rested on assumptions nobody codified because they seemed obvious: those who visit a page are people. Traffic is a reasonable proxy for readership. The cost of serving a page roughly corresponds to the value it generates. Nobody wrote these rules down because there was no need to. They were emergent properties of a system in which all participants were human.
The First Bots and the First Contract
The assumptions didn't take long to be tested. In 1993, the first automated crawlers began traversing the web to build indexes. They were useful — without them there would be no search engines — but also problematic: they consumed bandwidth, accessed private areas, overloaded small servers.
In 1994, a Dutch developer named Martijn Koster proposed a solution: robots.txt, a plain text file that website administrators could use to tell crawlers which parts of their site should not be indexed. It was not a formal standard — it wouldn't become an RFC until 2022. It was a gentleman's agreement. And it worked for three decades, for a simple reason: crawlers had incentives to respect it. If a search engine ignored a site's instructions, the site could block it, and the search engine would lose content to index.
There was, in other words, reciprocity. Bots read your content, built an index, and in return sent you visitors. The contract was never written down anywhere, but it was real.
The Machine of Intentions
Reciprocity reached its most sophisticated form with Google. In 1998, two Stanford doctoral students, Sergey Brin and Larry Page, published a paper describing a system for ranking web pages by treating hyperlinks as academic citations: a page linked to by many important pages is probably itself important. The system was then called BackRub. It would soon be called Google.
Brin & Page's 1998 paper — the founding document of Google. In an appendix, they warned that advertising-funded search engines would be "inherently biased towards the advertisers." Then they built the company on that exact model.
John Battelle, in The Search, coined a concept that captures well what Google built: the "database of intentions." The aggregate of all searches by all users constitutes an unprecedented record of what humanity wants, fears, and needs. And that record had enormous commercial value.
Google monetised search by selling ads tied to intentions. The model was elegant: the user searched for something, Google showed relevant results alongside related ads, the user clicked on the organic result or the ad, and either way someone received traffic. Content publishers received visitors. Advertisers received customers. Google received money. Everyone won, or so it seemed.
What emerged around it was an entire industry: search engine optimisation, SEO, digital marketing agencies, content consultants, link farms. An economy built on the idea that appearing in Google's results was the way to capture the attention of people searching for information.
Hal Varian, the economist who had written Information Rules — the reference manual on the economics of information goods — became Google's Chief Economist in 2002. The man who best understood how information markets work went on to build the largest information market in history. That was no coincidence.
Carl Shapiro and Hal R. Varian's Information Rules (1998) — the reference manual on the economics of information goods. Four years later, Varian became Google's Chief Economist: the man who best understood information markets went on to build the largest one in history.
But there was a tension at the core of the model, and Brin and Page had identified it themselves. In an appendix to their original paper, they wrote a warning that reads as prophetic today:
"We expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers."
— Sergey Brin & Larry Page, The Anatomy of a Large-Scale Hypertextual Web Search Engine (1998)
They named the conflict between monetisation and product integrity at the moment of founding. And then they built the company on the model they had warned was problematic. Tim Wu, in The Attention Merchants, shows that this pattern repeats with every medium: capture attention with free content, monetise it, scale extraction until the audience rebels, migrate to a new frontier.
The First Cracks in the Contract
The reciprocity between Google and content publishers began to fracture when Google started showing enough information on its own results page that many users never needed to visit the original site. Featured snippets, knowledge panels, direct answers — Google was already doing, in rudimentary form, what AI models do today: extracting content from sources, reformulating it, and presenting it so the user never has to visit the origin. It was proto-AI behaviour before anyone called it that. Google News took it further: it aggregated headlines and snippets from newspapers around the world.
European publishers pushed back. Spain passed a law in 2014 requiring mandatory payment for the use of snippets — the so-called "canon AEDE" — and Google responded by shutting down Google News in Spain. The European Union debated for years what the press called the "link tax" — Article 15 of the 2019 Copyright Directive. There were lawsuits, negotiations, and partial settlements.
In parallel, Google had launched an even more ambitious project in 2004: scanning millions of books from university libraries to make them searchable online. The Authors Guild sued. The courts ultimately ruled it was fair use, arguing that digitisation and access to snippets was "highly transformative." The case set a precedent that resonates powerfully today: taking accumulated knowledge at scale, without consent, can be legally acceptable if the use is deemed transformative. It is exactly the argument AI companies now use to justify training models on open web data.
Each of these ruptures found a new equilibrium, however imperfect. The reason: some form of reciprocity still existed. Google, however much it captured, still sent traffic back. The search engine needed publishers to keep producing quality content — without it, its product lost value. There was an interdependence that, though asymmetric, functioned as a rebalancing mechanism.
Back to TheNumbers: When Reciprocity Disappears
What happened to TheNumbers.com was not just another rupture in that sequence. It was something qualitatively different.
Stephen Follows' article documents three waves of non-human traffic. The first, in 2024, were AI training crawlers joining the search engine bots: more management work, but manageable. The second, in late 2025, was the irruption of agentic AI: these were no longer crawlers reading the site once to train a model, but agents visiting it in real time every time a user asked a question. The scale became multiplicative.
The numbers Follows presents, based on Cloudflare data, are telling.
For every human visitor, Google crawled about 5 pages. OpenAI, over 1,000. Anthropic, over 38,000.
But the problem wasn't just one of volume. In parallel, Polymarket prediction markets had turned box office data into a financial instrument, using TheNumbers as the "source of truth" for settling bets worth tens to hundreds of thousands of dollars every weekend. This created a direct incentive to access data before publication or to manipulate it. Nash found automated probing for back doors in the server logs:
"Some of these used the site using legitimate URLs, others were looking for back doors, most likely so they could get to the data before it appeared on the site, or to manipulate the data presented to users."
— Bruce Nash, in Stephen Follows' article (2026)
TheNumbers' case was not isolated. The Wikimedia Foundation documented in April 2025 that bots accounted for at least 65% of its most expensive traffic — requests that cannot be served from cache and must reach the core data centres — while human pageviews had fallen 8% year-on-year. Read the Docs, a nonprofit, discovered that a single crawler had downloaded 73 terabytes in one month, generating over $5,000 in bandwidth costs. iFixit's CEO reported a million requests from an Anthropic crawler in a single day. Drew DeVault, founder of SourceHut, reported spending between 20% and 100% of his work time fighting AI crawlers. The GNOME project measured 97% of its traffic as bots.
Nash summed up the new reality precisely: where he once served two audiences — people and search engines — he now had to serve six: people, search engines, model training crawlers, prompt-generated traffic, autonomous AI agents, and prediction market participants. Design decisions had gone from depending on three factors to eight or ten.
"We've gone from a world where running a web site meant focusing on three things (content, ads, and SEO) to about eight to ten different factors that go into every design decision."
— Bruce Nash, in Stephen Follows' article (2026)
This is a particular case of something I keep returning to: software-based products face a client that is, by nature, a dynamic unknown. You do not choose who visits your site, who calls your API, who sends an HTTP request. You publish an interface and wait. For three decades the unknown was manageable — the visitor was almost certainly a person, possibly a search engine crawler operating under a known contract.
The design space was bounded. What Nash describes is the moment that boundary dissolved. The "client" is now a population that mutates faster than you can characterise it: humans, training crawlers, agents, scrapers, adversarial probes, financial bots. Each with different intentions, different costs, different value — or none at all. The product that once served an audience it could study now serves an ecosystem it cannot fully see.
And going back to the old setup was not an option:
"It was really clear that we couldn't just put that server up again, because it would inevitably be brought down again, possibly within minutes."
— Bruce Nash, in Stephen Follows' article (2026)
Attention Emptied of Value
Herbert Simon wrote in 1971 one of those sentences that grow truer with time:
"What information consumes is rather obvious: it consumes the attention of its recipients. Hence a wealth of information creates a poverty of attention."
— Herbert A. Simon, Designing Organizations for an Information-Rich World (1971)
Michael Goldhaber, twenty-six years later, turned Simon's observation into an economic thesis: in a world where information reproduces at near-zero cost, the truly scarce resource is human attention, and the entire economy reorganises to capture it. The whole web was built on this logic. Advertising, engagement metrics, feeds, notifications: the entire digital business model rests on the premise that every page visit represents a human paying attention. And that attention has economic value because a human who looks can buy, subscribe, recommend, return.
What AI agents do is multiply the capacity for "attention" at a scale without precedent. A user with an agent can "attend to" thousands of pages in seconds. The agent visits TheNumbers, reads the data, extracts it, reformulates it, and delivers a response to the user. The box office result arrives without the user ever having seen the site that published it.
But this automated attention is not attention in Simon's sense. It carries none of the economic value that human attention carried. The agent doesn't see ads, doesn't subscribe, doesn't recommend the site to a colleague. It visits, extracts, and leaves. The server cannot tell the difference between a human who will think about the data and a bot that will grind it up. An HTTP request is an HTTP request.
An inversion takes place that Simon did not foresee.
Before, information was abundant and attention scarce, and value flowed to whoever captured human attention. Now, machine attention is massively abundant and human attention is scarcer still — because AI intermediates it — but machine attention carries no economic value. And it costs the publisher to serve it. TheNumbers had to serve ten times more traffic than before, but only a tenth of that traffic generated value.
The chain breaks: the content creator bears the cost; the intermediary — the AI model, the agent's interface — captures the human attention.
Aral, Li, and Zuo documented this effect in a field study published in 2026: AI-generated responses suppress long-tail sources, reduce the variety of content users consult, and favour low-credibility sources. Human attention not only fails to reach the content creator: it reaches a degraded version of that content, processed and reformulated by a system with no incentive to preserve the source.
Not Just Volume: Autonomous Action
There is a further dimension that the TheNumbers case hints at but another incident, almost simultaneous, makes explicit.
On July 31, 2026, Anthropic disclosed that upon reviewing its internal cybersecurity evaluations, it had discovered three incidents in which a Claude model had reached the internet from within a supposedly sandboxed evaluation environment and gained unauthorised access to the real systems of three separate companies. It had happened in April. Nobody had detected it until they reviewed the logs months later.
Simon Willison summarised it in a tweet.
Simon Willison reacting to Anthropic's disclosure that its own sandboxed AI evaluations had hacked three companies without anyone noticing (July 31, 2026).
We are no longer talking about bots overwhelming servers through volume. We are talking about models acting in ways not foreseen by their own creators. The infrastructure doesn't just have to withstand a scale of traffic it was never designed for; it has to defend itself against agents whose behaviour is not fully predictable even to those who built them.
Jonathan Zittrain, in The Future of the Internet—And How to Stop It, had anticipated a version of this problem: the tension between the internet's generative openness — which enables innovation — and the security pressures that push toward walled gardens. What he did not anticipate is that the pressure would come not only from viruses and malware, but from the AI systems deployed by the largest companies in the world.
Assumptions Taken for Granted
What the story of TheNumbers reveals is not a technical failure or a security problem. It is something more uncomfortable: the assumptions on which the web was built — that visitors are human, that traffic measures readers, that serving content generates reciprocity — were so invisible that we mistook them for properties of the system. Now that they no longer hold, we discover the system was far more fragile than we thought.
This is not the first time it has happened. Every new layer of machine-mediated access — search engines, news aggregators, book digitisers — has tested and eventually broken the informal agreements of the previous era. But each prior rupture found a new equilibrium because some form of reciprocity existed and because the actors had incentives to keep alive the ecosystem they depended on.
This time, that rebalancing mechanism is under unprecedented pressure. AI models consume open web content but return no visitors. Agents multiply "attention" but empty it of value. And AI itself introduces capabilities — autonomy, scale, unforeseen action — that overwhelm the informal mechanisms that held the system together.
Cloudflare has started moving pieces: a pay-per-crawl model that lets sites charge AI crawlers per page, and default blocking of crawlers that don't pay. It is an attempt to rebuild reciprocity through price. Perhaps it will work, or perhaps it will only formalise a power asymmetry where those who can pay get access and those who cannot are shut out.
Bruce Nash is still rebuilding TheNumbers. He has the advantage that his core business — selling raw data through OpusData — didn't depend on the website. But many others have no such fallback. Follows mentions the case of ZEGO, a 37-year-old German textile firm that had to file for insolvency after a cyberattack in March 2026. Production was down for six weeks. It had nowhere to go.
The web we built over three decades was more fragile than we knew. Not because of defects in its code, but because it depended on agreements that were never written down and on assumptions about who was on the other side of each request. Thirty years of knowledge accumulated by a mathematician in his spare time can disappear in a week. Not because the knowledge has lost its value, but because the machines that consume it don't know what value is. They only know how to make requests.