Simo Virokannas

Writings and ramblings

The End of the Infinite Internet

TL;DR: The AI industry’s scaling strategy has two hidden assumptions: that useful human information is effectively infinite, and that producing information faster is equivalent to producing more value. Neither assumption is true.

The AI industry has spent the last few years behaving as though intelligence was a simple matter of scale. Manufacture more GPUs, build data centers, use more parameters, get more training data, make better models, invent more money.

There is a fundamental flaw in that equation: there’s less and less humans in it. Well, not literally, at least not yet. But we are running out of the useful things humans have written.

Epoch AI estimates that the effective stock of high-quality, repetition-adjusted human-generated public text is around 300 trillion tokens. Their models suggest that, if current scaling trends continue, this stock could be fully utilized sometime between 2026 and 2032, with the median projection being around 2028. In other words, the data wall isn’t necessarily here yet, but it is possible it is or at least close enough to be visible.

That estimate was already uncomfortable when it was published in 2025. And it is considerably more interesting now that we’ve potentially reached that milestone.

The problem isn’t only that humans aren’t producing enough new text. The quality of the text being produced is changing too.

We have spent the last decade putting smartphones in the hands of everyone, including children, replacing boredom with infinite entertainment, replacing memorization with search, replacing writing with autocomplete, replacing conversations with doomscrolling. The evidence does not establish that smartphones are making entire generations less intelligent. It does establish that excessive digital device use is associated with worse educational outcomes, particularly when the devices become a source of distraction.

The OECD’s PISA 2022 report found that around 30% of students reported being distracted by digital devices in most or every mathematics lesson, and those students scored about 15 points lower in mathematics even after accounting for socioeconomic and school characteristics. Students using digital devices for leisure for more than an hour a day also tended to perform worse.

The important word is excessive. Moderate use of technology can be beneficial. The OECD actually found that students using digital devices for learning for up to an hour a day performed better than students who used none. It is overuse and distraction that led to poorer results.

So here’s an interesting situation:

The machines are getting better at generating information at precisely the moment when the humans generating the underlying information are increasingly assisted, mediated and distracted by machines.

It actually works

This article uses the term AI, or Artificial Intelligence, repeatedly. I want to make it perfectly clear that I personally believe in only the first of the two letters to be true. The use of the abbreviation is with the sole purpose of referring to something people call something with that name. Any semantic personification of an artificial process or a machine is purely anecdotal.

AI can be extremely useful. There have been many studies, albeit often biased, but still based on data:

  • Stanford/MIT study of 5,179 customer-support agents found that generative AI increased productivity by 14% on average, with a 34% improvement among novice and lower-skilled workers, already in 2023.
  • Around the same time, GitHub’s controlled experiment with professional developers found that developers using Copilot completed a coding task 55% faster than the control group, while also having a slightly higher completion rate (and I’m really hoping this doesn’t get interpreted as ‘Copilot made developers 55% faster’ – in a controlled experiment, with a controlled input and expected outcome, with no room for imagination or creativity, the process of completing the task itself was 55% faster).

These are real, measured productivity gains.

Writing a support response 14% faster is useful. Writing code 55% faster is useful. Producing a marketing draft in thirty seconds instead of thirty minutes is useful. But a task completed faster doesn’t necessarily mean a more productive company.

Why?

A developer can produce twice as much code, a salesperson can produce ten times as many emails, marketing can kick out a thousand pieces of content, and a company can deploy an AI chatbot across every department.

None of this guarantees that anyone wanted that much more code, more emails, auto-generated marketing content or even a chatbot. Sometimes AI means faster production of things nobody needed.

The research articles referenced in this article were dug up using an AI prompt. Did you need this article? Probably not.

The productivity paradox

This creates an awkward distinction: AI is very good at reducing the cost of producing information. It is extraordinarily good at summarizing information. It is not necessarily any good at all at increasing the value of information.

Many corporate AI initiatives feel strangely circular. A company spends millions teaching employees how to use AI, employees use AI to generate more material, managers use AI to summarize the material, other employees use AI to rewrite the summaries, and eventually somebody produces a presentation explaining how much productivity was created.

The output is enormous, but the value is harder to find.

McKinsey’s 2025 State of AI survey found that 88% of organizations were regularly using AI, yet only 39% reported an enterprise-level EBIT (Earnings Before Interest and Taxes) impact. Nearly two-thirds had not yet begun scaling AI across the enterprise.

That is the part of the AI “revolution” that doesn’t get quite as much attention. We have become very good at demonstrating that AI can make somebody do something faster. We are much worse at demonstrating that it makes a company richer.

Let’s walk through some problems plaguing these experiments, starting with scope.

The scope problem

The easiest things to automate are usually tasks, not businesses.

AI can summarize a meeting. It can’t make a meeting worthwhile.

AI can write code. It can’t decide what software should exist or what code should exist.

An AI chatbot can answer customer questions. It can’t determine whether the customer needed to call in the first place.

AI can generate ten thousand marketing messages. It can’t make ten thousand people, or even one, want the product.

The difference is between doing something faster and doing something valuable.

When novelty wears off and companies have to justify the billions they are spending, that distinction is going to become increasingly important.

The data problem

Large language models are neural networks trained to compress enormous statistical regularities in their training data into model parameters.

The obvious strategy for making them better (training the network) has been to give them more. More books, more websites, more code, more conversations, more images, more audio, more documentation.

But once the useful human material has been consumed, there isn’t an infinite second internet waiting underneath it.

What’s left is synthetic data.

The models can generate training examples for other models, which is useful in some domains, particularly where answers can be verified mathematically or procedurally. But synthetic data has an obvious limitation.

If I photocopy a book a million times, I haven’t created a million books. I have created a million copies of one book.

This isn’t only theoretical. A 2024 Nature study demonstrated what researchers call model collapse, where repeatedly training generative models on model-generated data causes them to progressively lose information about the original data distribution.

The researchers found that indiscriminately training on generated content causes the tails of the original distribution to disappear. In other words, unusual, rare and less predictable information gets lost first.

The Internet is becoming contaminated.

The internet used to be a gigantic record of human activity. People wrote reviews because they had used products, (some) programmers wrote documentation because they had solved problems, scientists published research because they had discovered something, people wrote forum posts because something had happened to them. This data was based on realities (or imagination, both things AI is missing), and all of this provided value to the dataset. The problem isn’t that synthetic information is worthless. The problem is that it increasingly becomes difficult to distinguish information about reality from information generated by a machine predicting what information about reality should look like.

Machines are beginning to poison the well.

The internet continues growing, but its usefulness as a training dataset declines.

Eventually, perhaps, more of the information produced will either be intentionally synthetic or written by a machine pretending to be a human who had something to say.

Grasping for reality

Once the public internet becomes saturated, the valuable data isn’t necessarily going to be another billion web pages; instead, it is going to be any data that machines cannot manufacture themselves.

Industrial processes, scientific experiments, medical research, financial transactions, proprietary software, customer interactions, sensor networks, images and video of the physical world, human decisions and expert demonstrations.

Reality becomes the desired dataset.

This is one reason the next phase of AI may look less like web scraping and more like ownership.

Who has the best proprietary data? Who has millions of customers generating interactions? Who has the largest software repositories? Who has access to scientific research? Who has billions of videos showing how humans actually behave?

The few companies with direct access to human activity have an increasingly valuable resource.

Google has its search, YouTube, Android and Maps. Meta has billions of users generating content and interactions. Microsoft has enterprise software and enormous amounts of business data.

The pure AI companies have something else: they have the models, but the models may not be the most valuable part.

The IPOcalypse

This becomes particularly interesting when these companies start reaching for the public markets.

Private companies can justify almost any amount of spending by promising that the next model will change everything. Public companies eventually have to explain revenue, margins, capital expenditure and free cash flow.

Anthropic has now confidentially filed for an IPO and is reportedly preparing for a possible public debut as soon as this fall. OpenAI, meanwhile, appears to be preparing for a possible IPO further down the road. The valuations being discussed assume extraordinary levels of future growth.

What happens if the investors decide the growth isn’t worth the price?

What happens if inference becomes cheap and customers discover that they don’t actually care which company’s model produced the answer?

The risk for them isn’t that AI stops working. The risk is that AI starts working too similarly across competing companies. If every major model is good enough, model intelligence starts to look more like a commodity. Commodities have terrible margins. Terrible margins aren’t attractive to investors. IPOcalypse.

After the crash

If any of the major AI companies eventually lose 70% or 80% of their market value after going public, I don’t think they disappear. I think they will get boring. And a lot of this happens to companies going public even if they’re successful:

Research projects get cancelled, hiring freezes, compute budgets get scrutinized, experimentation disappears, enterprise contracts become more important, margins become more important, cash flow becomes more important.

The companies stop asking, “How do we build the best possible model?” – instead, “How do we make more money with the model we already have?”

And this is probably healthy. This would also be a great time to rename it to something less anthropomorphizable to remove misconceptions about what a machine can or can’t do for us as humanity (please, if you know a better word for anthropomorphizable, send it in a comment, WordPress is yelling at me in red ink).

The survivors will move up the stack, selling workflows rather than tokens, acquiring proprietary data rather than endlessly scraping the public internet, turning AI into a component of products rather than selling AI itself as the product.

OpenAI would have to become much more than a model company, Anthropic would have to prove that enterprise demand can support its enormous infrastructure costs, Google and Microsoft could simply absorb AI into businesses they already have, Meta could continue using increasingly cheap intelligence to make its existing platforms more valuable.

The companies that survive a valuation collapse won’t necessarily be the companies with the best models, but they’re more likely to be the companies with something else worth owning: customers, infrastructure, data and enough capital.

The end of the infinite Internet

For years, the AI industry has been built around an assumption that felt reasonable because it was true for a surprisingly long time: There will always be more data.

But there is no infinite Internet. There is only an Internet built by a finite number of humans, and we are now asking those humans to either use the machines to produce their output, or to compete with machines at producing the very material the machines need to learn from.

We spent the first phase of the AI “revolution” teaching machines everything humans had already said.

The AI industry thought intelligence could be scaled by manufacturing more compute, but it eventually runs into the content production cap of a finite civilization. The next frontier isn’t an infinite Internet. It’s reality.


Comments

One response to “The End of the Infinite Internet”

  1. Reima Karppinen

    Nyt oli kiinnostava! Ite pohtinut noita samoja. Mietin anthropomorphizable kohdal, että onpa nerokas ilmaisu, joten ei ole parempaa vastinetta tarjolla.

    Mukava lukea IHMISEN kirjoittamaa tekstiä. Kiitos (taas)!

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.