Former AlphaGo researcher Thore Graepel argues that LLMs need auditable belief states and search, not simply longer chains of thought.

Large language models have become better at producing step-by-step answers, but that apparent deliberation does not necessarily amount to reasoning, according to Thore Graepel, a former Google DeepMind researcher and core member of the AlphaGo team.
In an analysis published by MIT Technology Review, Graepel argues that current LLMs remain fundamentally next-token prediction systems. He says trustworthy machine intelligence will require a separate reasoning architecture that records beliefs, tests possible explanations, updates its conclusions with evidence, and makes its decision process inspectable.
The argument matters as AI builders and enterprise buyers increasingly deploy models for coding, research, diagnosis, planning, and other tasks where a plausible answer is not enough. Graepel’s proposal is not a product announcement or a new benchmark result. It is a technical critique from a prominent researcher who recently left Google DeepMind and is calling for a different direction in AI development.
Graepel’s central distinction is between fluent pattern completion and deliberative reasoning. An LLM generates one token after another based on statistical relationships learned from data. That process can produce useful explanations, mathematical solutions, and code, but the underlying mechanism remains the same even when the model generates intermediate steps.
Those intermediate steps are commonly called chain of thought. Graepel acknowledges that chain of thought can improve performance, particularly in mathematics and coding. His objection is that generating more text does not create an independent reasoning system. The model is still extending a sequence through next-token prediction rather than manipulating an explicit set of propositions or hypotheses.
He identifies three weaknesses. Current models generally lack a persistent and inspectable record of what they believe, how confident they are, what evidence supports each claim, and which questions remain unresolved. They also combine stored knowledge with the operations used to manipulate it inside the neural network’s weights, rather than maintaining a clean separation between knowledge and reasoning. Finally, their written explanations may not faithfully represent how an answer was produced.
That last problem is especially consequential for high-stakes systems. If an AI makes a mistake in medicine, engineering, or scientific research, users need to determine whether the failure came from bad evidence, an invalid assumption, or a flawed inference. A convincing explanation generated after the answer may not provide that diagnostic information.
Graepel uses AlphaGo to illustrate the kind of architecture he believes modern AI needs. During its 2016 match against South Korean champion Lee Sedol, AlphaGo played an unusual move known as move 37. The move initially looked like an error, but AlphaGo eventually won the game and the match 4-1.
The important point, Graepel argues, is that the move did not come from intuition alone. AlphaGo combined a neural policy network, which estimated promising moves, with a search system that explored possible continuations. The search built a game tree containing alternative positions and future moves, then evaluated those branches before selecting an action.
In this design, neural networks supplied fast judgments while search provided structured deliberation. Graepel compares the arrangement with the behavioral distinction between rapid, intuitive thinking and slower, analytical thinking. The two components supported one another: intuition narrowed the possibilities, while search tested those possibilities against future consequences.
The comparison has limits. Go supplies a defined board, a fixed set of legal actions, and rules that determine the outcome of each move. Open-ended problems in science or business do not provide those conditions. The state of the world is incomplete, available actions vary, and consequences may be uncertain.
Even so, Graepel says the architecture offers a useful template. A general-purpose reasoning system could maintain an “epistemic state”—a structured record of settled claims, doubts, rejected explanations, open questions, and supporting evidence. Each reasoning step would update that state rather than merely append more generated text.
The strongest claims in this story come from Graepel’s MIT Technology Review analysis, not from a newly reported experiment or independent product evaluation. The article provides historical and architectural detail about AlphaGo, but it does not present a new benchmark comparing current LLMs with a proposed epistemic-state system.
Graepel also refers to research showing that language models can produce explanations that do not reliably reflect the processes that led to their answers. The supplied article does not identify specific studies, datasets, or model results in enough detail to assess the size or generality of that effect here.
That distinction is important. The criticism does not mean current models are useless or incapable of solving multi-step problems. It means that performance on a task should not automatically be treated as proof that a model reasons in the same sense as a system that explicitly tracks alternatives, evidence, and belief revisions.
The article also presents Graepel’s own design direction. He proposes using LLMs to suggest strategies, interact with tools through APIs or code, and assess whether claims are supported by evidence. An independent evaluator would then judge whether each step actually reduces uncertainty before allowing the system to update its knowledge.
For AI builders, the practical signal will be whether new systems move beyond longer prompts and longer generated reasoning traces. Developers should watch for architectures that maintain inspectable intermediate state, separate factual memory from inference procedures, and log how evidence changes a conclusion.
Tool use will be another key area. A system that can run code, query databases, retrieve documents, or design an experiment may become more reliable if those actions are connected to explicit tests rather than treated as decorations around a language-model response. The relevant question is not simply whether an agent calls a tool, but whether the result changes a structured belief state in a traceable way.
Enterprise buyers should also ask vendors how they audit model decisions, reproduce failures, and distinguish retrieved evidence from generated assertions. Existing chain-of-thought interfaces may make an answer look more reasoned without making the underlying process more verifiable.
Researchers, meanwhile, will need to show whether architectures inspired by AlphaGo can handle open-world tasks without becoming too expensive or too slow. Maintaining and evaluating many hypotheses may improve reliability, but it could also increase compute, latency, engineering complexity, and the burden of validating each update.
Graepel’s intervention is best understood as a challenge to how the AI industry labels progress. Better results on reasoning benchmarks and more elaborate explanations are valuable, but they do not by themselves establish that a model has an auditable reasoning process.
The harder test is operational: can a system preserve uncertainty, expose the evidence behind a conclusion, revise its beliefs when new information arrives, and show where a failure occurred? If AI is to support scientific discovery, medicine, and other high-consequence workflows, those capabilities may matter as much as raw model scale. The next generation of AI systems will be judged not only by the answers they produce, but by whether users can inspect and trust the path taken to reach them.