Research

AI Experts Underestimated the Field's Speed, Study Finds

A Forecasting Research Institute study finds AI experts consistently lowballed progress: IMO gold came five years early and Anthropic's revenue runs five times forecasts.

Top AI experts badly underestimated how fast the field is moving, study finds
Top AI experts badly underestimated how fast the field is moving, study findsAI-generated
By Elena Vasquez5 min read

Updated

Why it matters

  • AI reached gold-medal level at the International Mathematical Olympiad five years ahead of the median expert forecast, per the Forecasting Research Institute.
  • Anthropic's annualized revenue is about five times what experts predicted.
  • Forecasts for real-world uses such as self-driving cars show a more mixed picture, indicating underestimation was not uniform.

AI reached gold-medal level at the International Mathematical Olympiad five years ahead of the median expert forecast. That is the headline finding from the Forecasting Research Institute, which examined how leading AI experts' predictions compare with what has actually happened — and found the experts consistently lowballed the pace of progress.

The misses are not marginal. They are large, repeated, and pointed in one direction: AI capabilities and commercial momentum have outrun professional forecasters' expectations.

The study, released by the Forecasting Research Institute, looked at predictions made by leading AI experts and measured them against observed outcomes. Two data points anchor the analysis.

The first is mathematical reasoning. Experts were asked when AI systems would perform at gold-medal level at the International Mathematical Olympiad, one of the most demanding competitions for pre-university mathematicians. The median forecast set a date years in the future. Reality arrived five years earlier than that median projection.

The second is commercial. Anthropic's annualized revenue runs at roughly five times what experts predicted it would at this point, according to the study. A fivefold revenue gap between forecast and outcome is not a rounding error. It signals that demand for frontier AI products has scaled far faster than even informed insiders anticipated.

Why the misses matter

Forecasting is not an academic exercise in this field. Timelines for AI capability milestones feed directly into policy debates, corporate capital allocation, and safety research prioritization. When median expert forecasts are systematically too slow, institutions that calibrate to those forecasts — regulators writing compliance windows, companies planning infrastructure, labs scheduling safety work — inherit the error.

The International Mathematical Olympiad benchmark carries particular weight because competition mathematics has long served as a proxy for general reasoning ability. A system that solves IMO problems at gold-medal level demonstrates multistep, verifiable reasoning, not pattern matching on scraped text. Experts treated that milestone as a distant marker of general capability. It arrived half a decade early.

The Anthropic revenue figure matters for a different reason. Revenue is the market's aggregate judgment on whether frontier AI is useful enough to pay for. Experts forecasting a figure that reality exceeded by roughly a factor of five underestimated not just technical progress but adoption — how quickly businesses and consumers would commit real money to AI products.

A more complicated picture outside the lab

The study's findings are not uniform across all domains, and the Forecasting Research Institute is explicit about that. Forecasts for real-world, physically embedded applications such as self-driving cars paint a more mixed picture.

That asymmetry is itself informative. Tasks with clean, digital inputs and outputs — proofs, code, text — have advanced faster than experts expected. Tasks that require acting in the messy physical world, where sensors, edge cases, liability, and regulation interact, have not reliably beaten the forecasters' clock. The gap between the two suggests the bottleneck for real-world AI deployment is at least as much about integration and environment as about raw model capability.

It also complicates a simple reading of the study. The experts did not uniformly underestimate everything. They underestimated the pace at the frontier of cognitive tasks and in the commercial market for AI, while their predictions for embodied, real-world deployment held up better. Any conclusion that experts are simply always too pessimistic would overread the data.

What systematic underestimation does to planning

A five-year error on a capability milestone is significant by any planning standard. Government AI strategies, compute procurement cycles, and safety research programs are typically built on multi-year horizons. If the reference forecasts those plans rely on run five years slow, the plans arrive late by construction.

The commercial error compounds this. If revenue at frontier labs runs five times above expectations, then compute demand, energy requirements, and competitive pressure among AI developers all scale beyond what planners budgeted. Each of those downstream quantities was presumably forecast with reference to the same community of experts whose median predictions the study found wanting.

The Forecasting Research Institute's work therefore carries a practical implication: consumers of AI forecasts — policymakers, investors, safety researchers — should treat median expert timelines as a reference point with a demonstrated bias, not as a neutral baseline. The direction of the historical error, at least on the cognitive and commercial benchmarks examined, has been toward underestimating speed.

The limits of the finding

Three caveats bound the conclusion. First, the study examined specific benchmark predictions — the IMO milestone, Anthropic's revenue, self-driving deployment among them — not every claim experts have made. A systematic bias on these measures does not prove bias on all measures.

Second, forecasting error cuts both ways in principle. The mixed results on self-driving cars show that experts have not been uniformly slow. Any model of "experts always lag reality" fails against the study's own more nuanced findings.

Third, past forecasting error does not by itself determine the future pace. That AI beat the IMO forecast by five years does not guarantee the next milestone will arrive early. Forecasters who overcorrect into aggressive timelines could err in the opposite direction on the next generation of predictions.

The so-what

The study's core finding stands regardless of those caveats: on the measures examined, the people paid closest attention to AI — the leading experts — consistently expected less, and slower, than the field delivered. A gold-medal IMO performance five years ahead of the median forecast and an Anthropic revenue run rate roughly five times predictions are the two hardest numbers in the finding, and both point the same way.

For anyone whose plans depend on when AI systems will clear the next capability bar, the study suggests stress-testing those plans against timelines considerably faster than the expert median — while remembering that domains like self-driving, where forecasts have held up better, may continue to move at their own pace.

Original: forecastingresearch.substack.com

Share this article:

More from Elena Vasquez

Elena Vasquez

Show full bio

Market editor covering media and advertising at AI In Context.

122 articles

Related articles

  1. OpenAI Predicts AI-Made Discoveries by 2026 as Intelligence Costs Plunge
  2. The Cost of Reaching a Fixed AI Performance Level Is Dropping Fast
  3. OpenAI Recruits Elite Mathematicians After Research Release Stumbles

« Previous articleNext article »