The AI race did not end because the machines hit a wall. It ended because we lost the ruler.
That distinction is everything, and from here in 2036 it is easy to miss. The plateau looks like a ceiling — as if intelligence simply ran out of road around 2031. It didn't. What ran out was our ability to tell two systems apart. When you cannot measure which model is better, "better" stops being a thing companies can sell, and an industry that lived on the promise of progress quietly loses the reason to make any.
The collapse was visible years before anyone named it. By early 2025, the field's most-cited benchmarks were already saturating: leading models clustered above 90% accuracy on MMLU, the standard knowledge test, leaving no room to separate the frontier.1 On the public Chatbot Arena leaderboard, the rating gap between the top model and the tenth-ranked model narrowed from 11.9% to 5.4% in a single year; the gap between the top two shrank to 0.7%.2 The race was photo-finishing. Everyone crossed the line at once.
The conventional wisdom
The expectation, in 2025, was that this was a tooling problem with a tooling fix. Benchmarks saturate; you build harder benchmarks. And the field did exactly that — Humanity's Last Exam, released that January with 2,500 expert-written questions explicitly designed to resist saturation, was the flagship attempt to rebuild the ruler.3 The assumption underneath it was that measurement could always be made to outrun capability, and that as long as some test still hurt, the contest would continue.
That assumption held for the science and broke for the market. New benchmarks kept differentiating models in the lab. But buyers do not run evaluation suites. By the time capability gaps shrank to fractions of a percent on the tests anyone trusted, the difference had already dropped below the threshold a customer could feel — or pay for. The ruler still worked. It just no longer described anything a purchasing decision could turn on.

Force one: convergence enhances taste and obsolesces the score
When products become objectively identical, competition does not stop. It relocates — to whatever remains unmeasurable.
This is the first force, and it is the one nobody priced in. Convergence enhanced the value of the subjective. Warmth, brevity, the texture of a model's prose, the particular way it said no — qualities that were rounding error when capability gaps were large became the entire basis of choice once those gaps closed. A profession of taste-makers grew to arbitrate them: people paid to tell a German engineering firm which functionally identical model "felt" more precise. (How many such advisers there are, and what they charge, is a 2036 detail this magazine can only estimate; the dynamic is real, the headcount is not audited.)
What this obsolesced was the benchmark score as a commercial object. For roughly forty years, computing sold itself on numbers a buyer could compare — clock speed, capacity, throughput. That era ended quietly around 2031. "Faster" and "smarter" stopped moving product because they stopped meaning anything a customer could verify. Marketing departments grew. Evaluation labs shrank.
Force two: identical products retrieve the gentlemen's agreement
The second force is older than computing, and convergence retrieved it: the tacit understanding among rivals not to compete on price.
The mechanism needs no conspiracy. When products are indistinguishable, price competition is mutual ruin — so firms, observing each other, simply settle near the same number without ever speaking. Antitrust scholars call this tacit collusion, and they have warned for years that it slips through the law's fingers: the Sherman Act requires proof of an agreement, and there is none to find.4 In 2024, the FTC and the Department of Justice, alongside UK and EU regulators, issued a joint statement warning that algorithms could let competitors "collude on terms or business strategies" without ever coordinating.5 That was written about pricing software. It described, in advance, what an entire industry of equivalent products would drift into on its own.

The reversal
Here is the turn, and it is the one that should have been obvious.
A test is not a scoreboard. It is an incentive. For forty years the benchmark did invisible work: it told a company that spending a billion dollars on research would show up as a number it could charge for. Measurement was the thing that made innovation pay. Pull the ruler out, and the spending becomes irrational — not because the science failed, but because nothing rewards it. The tool we built to track progress turned out to be the tool that funded it, and we only learned that by removing it.
This is the cruelest inversion in the whole story. The harder-benchmark project was meant to keep the race alive. By proving the old measures dead before any new one took hold commercially, it helped usher in the interregnum it was designed to prevent — the years with no working ruler at all. Note the speculative line here: that 2036 sits inside a measurement vacuum is a projection, not a reported fact. What is documented is the saturation, the convergence, and the warning. The vacuum is where those lines point if nothing intervenes.
And the strangest part is what the present-day data refuses to confirm. The original fear was a research collapse — papers drying up, talent fleeing. The opposite is on the record. AI publications more than doubled from 2013 to 2023, from roughly 102,000 to over 242,000, still climbing at the moment the benchmarks died.6 The output never stopped. What this piece imagines stopping, by 2036, is not the work but the willingness to fund the unmeasurable kind — the slow, speculative research whose payoff no longer shows up on any ruler a buyer reads. That is the speculation. The thriving research pipeline of 2025 is the fact it has to argue against.
The lesson
Mature markets do not announce themselves. They arrive as the quiet that follows a number nobody can move.
The durable principle is this: what you can measure, you will improve; what you stop measuring, you stop improving — and you will mistake the second for a law of nature. The AI plateau of 2036, if it holds, will not be a verdict on what machines can do. It will be a verdict on what we chose to count. The benchmark was never just a test. It was the price tag on progress, and we took it off to see what was underneath. The answer was nothing — not because nothing was there, but because, without the tag, no one could be paid to look.
This article is speculative journalism written from a fixed 2036 vantage. The benchmark saturation, model convergence, the harder-benchmark response, the antitrust warnings about tacit collusion, and the growth of AI research through 2023 are real and cited below. The 2036 end-state — the measurement vacuum, the taste economy, the collapse of funding for unmeasurable research — is one plausible trajectory, not a forecast, and is flagged as speculation throughout. No real company is alleged to have broken any law.
