Research

AlphaGo Veteran Warns Today's LLMs Can't Actually Reason

A core AlphaGo team member who just left Google DeepMind argues LLMs still run on intuition alone and calls for AlphaGo-style reasoning architectures.

Don’t be fooled—LLMs don’t reason
Don’t be fooled—LLMs don’t reasonAI-generated
By Rebecca Stone6 min read

Updated

Why it matters

  • Thore Graepel, a core member of the AlphaGo team, says he left Google DeepMind to pursue a fresh approach to machine reasoning modeled on AlphaGo's search architecture.
  • AlphaGo's policy network gave move 37 roughly a one in 10,000 chance of being played by an expert human; the move was selected by explicit game-tree search with thousands of branches.
  • Graepel argues chain-of-thought output does not constitute reasoning because models lack persistent epistemic state, separate knowledge from manipulation, and often concoct reasoning traces after the fact.

A core member of the AlphaGo team has left Google DeepMind over the conviction that today's large language models cannot reason—and that fixing this requires rebuilding AI systems around an explicit, inspectable reasoning process modeled on AlphaGo's search.

Thore Graepel, now chair of machine learning at University College London, makes the argument in a new essay. His reference point is move 37 in game two of the March 2016 match between AlphaGo and Lee Sedol in Seoul, where AlphaGo triumphed 4-1 over one of the greatest professional Go players of all time.

The move—a stone placed on the fifth line—looked so absurd that some commentators thought it was a programming glitch. Lee Sedol thought so too, at first. "I thought AlphaGo was based on probability calculation and that it was merely a machine," Lee said afterwards. "But when I saw this move, I changed my mind. Surely, AlphaGo is creative."

Graepel, who helped build the program, says the popular reading of move 37 as a flash of machine intuition is wrong. The intuition module—AlphaGo's policy network, trained to guess what a strong human would play—regarded the move as nothing special, giving it roughly a one in 10,000 chance of being made by an expert human player. What actually selected it was AlphaGo's search machinery, which explicitly constructed and searched a game tree with thousands of branches, weighing the future consequences of proposed moves.

The contrast with Deep Blue, which defeated reigning world chess champion Garry Kasparov in 1997 by looking six to eight moves ahead per player and evaluating 200 million positions per second using human hard-coded rules, frames the stakes. Go is vastly more complex: computing even a fraction of the possible outcomes would take a supercomputer billions of years. To win, AlphaGo had to sense who was ahead at a glance and invent moves no human had thought to play.

Graepel maps this architecture onto the two-system model of human cognition popularized by Daniel Kahneman. System 1 is fast, gut-level, effortless; System 2 is slow, step-by-step, deliberative. AlphaGo's networks supplied the hunches; its search supplied the deliberation, testing those hunches against moves and countermoves. Neither half worked alone. "Intuition alone would never have opted for move 37, and brute-force search would have struggled to sieve through all the many possible moves," he writes.

Chain of thought is not reasoning

Today's LLMs, in Graepel's telling, run almost entirely on System 1. A large language model picks the next token, over and over—fast, associative pattern completion across almost every subject people write about.

The field's response after ChatGPT debuted—chain of thought, where models generate intermediate steps that decompose a problem before committing to an answer—has produced real gains, above all in mathematics and coding. But Graepel argues it does not introduce a genuinely separate reasoning mechanism. The intermediate reasoning is still produced by the same next-token prediction process, just iterated for longer.

He identifies three shortcomings that disqualify chatbot behavior as reasoning in the sense a scientist would recognize. First, these models maintain no explicit, persistent, inspectable epistemic state—no open ledger of the hypotheses under consideration, confidence in various explanations, evidence being weighed, and unresolved questions held open for revision as new information arrives. Second, there is no clean separation between what a system knows and how it manipulates that knowledge; knowledge and reasoning remain interwoven in the neural network's weights, with no independently represented set of beliefs. Third, research has demonstrated that models often concoct chains of thought after the fact, reaching an answer by one route while reporting another.

The stakes extend beyond benchmark debates. In medicine, engineering, and scientific research, Graepel argues, it matters not only what a system concludes but how it arrives at the conclusion. When a medical diagnosis goes wrong, users need to pinpoint the failure: faulty reasoning, invalid evidence, or incorrect assumptions. A post-hoc narrative cannot support that audit.

The AlphaGo prescription

Graepel's proposed fix draws directly on AlphaGo's design. AlphaGo maintained a game tree—a data structure containing all the variations and possible futures it had considered, each move and position annotated with judgments from its neural networks. It updated the tree as reasoning progressed, then synthesized the information to choose a move.

A general reasoning system, he argues, should maintain an analogous epistemic state: what is settled, what is doubted, what has been ruled out, which questions remain open. Reasoning then becomes a sequence of moves that change that state to advance knowledge and reduce uncertainty—deducing consequences, decomposing problems, and deciding what question to ask, calculation to perform, or experiment to run next.

He concedes the difficulty. Open-world reasoning is harder than Go or chess: the current state of affairs is only partially known, the set of available actions is large and variable, and consequences of actions are stochastic or unknown.

But he argues recent LLM advances now supply the components. LLMs can suggest ways of tackling a problem based on what is known and what resources are available. They can interact with tools via APIs or code and help assess whether a claim is supported by available evidence. Crucially, an independent part of the system must evaluate each move by how much it actually resolves uncertainty, updating beliefs only when the change is backed by evidence. Once those rules are enforced, the model can accumulate certified knowledge and improve its reasoning policy by learning from past reasoning experiences. Graepel describes such a system as "the scientific method on steroids, with the purpose of producing knowledge that can withstand scrutiny."

Scale sharpens intuition, not deliberation

Graepel closes with a direct rejection of the scaling consensus. "I do not think we reach trustworthy machine intelligence by making system 1 bigger," he writes. "Scale sharpens intuition, but it does not make intuition more deliberative."

Move 37 mattered, he says, because a machine held a position, weighed possible futures, and chose the move its artificial instincts would likely have rejected. Society needs comparable creative moves in drug discovery, materials, climate, and diagnosis—fields where the board looks nothing like a Go board and nobody hands over the rules. Such insights will come only from systems whose conclusions arise from an auditable sequence of evidence, inference, and belief revision, he concludes, "rather than from a convincing story told after the fact."

The essay arrives as labs pour resources into exactly the approach Graepel questions—ever-larger models with longer chain-of-thought traces—making his departure from DeepMind and his architectural critique a data point in one of the field's central debates: whether reasoning emerges from scale or must be engineered.

Original: vialogue.wordpress.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

178 articles

Related articles

  1. Ten Years After Move 37, DeepMind Says AlphaGo Set the AGI Agenda
  2. AI Beats the Best Stratego Player Ever—On 16 GPUs and Pocket Change
  3. Ataraxos Beats Stratego's Greatest Player, Ending a Human Stronghold
  4. Google DeepMind's SIMA 2 Turns AI Into a Gaming Companion
  5. Google DeepMind researcher quits, calls push for superintelligence irresponsible

« Previous article