Safety

AI-generated text

Why current LLMs lack true machine reasoning and what AlphaGo teaches us

AlphaGo’s famous move 37 against Lee Sedol came not from raw intuition alone but from an explicit search and reasoning process that evaluated thousands of possible futures.

Why current LLMs lack true machine reasoning and what AlphaGo teaches us

In March 2016 in Seoul, Thore Graepel watched a program he had helped develop make what many observers called an absurd move on the fifth line of a Go board during game two of a five-game match. Move 37 looked like a gift to the human opponent, and some commentators suspected a bug. It was not: AlphaGo won that game, and ultimately defeated Lee Sedol 4–1. Lee later said he had thought AlphaGo was merely a probability machine, but the move changed his mind and made the system appear creative.

How AlphaGo differed from earlier game engines

When Deep Blue beat world chess champion Garry Kasparov in 1997, it did so by looking six to eight moves ahead per player and evaluating roughly 200 million positions per second, using human-coded rules. Go is far more complex: the value of a stone depends on how distant groups and territory develop over dozens of moves. Enumerating even a fraction of possible outcomes would require astronomic compute.

AlphaGo’s architecture therefore combined two elements: a policy network trained to predict what a strong human would play, and a search mechanism that explicitly constructed and traversed a game tree with thousands of branches. The policy network’s probabilities treated move 37 as unremarkable—about a one-in-10,000 chance based on expert play. What led AlphaGo to play it was the search machinery weighing long-term consequences across many future lines.

System 1 and System 2: a machine analogy

A well-known distinction popularized by Daniel Kahneman separates fast, gut-level thought (System 1) from slow, deliberative thought (System 2). AlphaGo provided a striking machine analogue: its neural networks produced hunches (this move looks promising), while the search mechanism provided deliberation (testing those hunches against sequences of moves and countermoves). Neither element alone would have sufficed: intuition alone would likely have not chosen move 37, and brute-force search alone would have struggled to narrow possibilities effectively.

Why today's large language models are different

Modern large language models (LLMs) operate by predicting the next token repeatedly. This is effectively System 1 in action: fast, associative, and excellent at pattern completion across vast text domains. After ChatGPT’s release, researchers found that language fluency alone often falls short for substantive problem solving. One proposed remedy was to have models generate intermediate reasoning steps—chain-of-thought—so they decompose problems and carry forward partial results. These methods have produced real gains, especially in mathematics and coding.

However, chain-of-thought generation does not introduce a genuinely separate reasoning mechanism comparable to AlphaGo’s search. The intermediate steps are still produced by the same next-token predictor, iterated longer before committing to an answer. Three core shortcomings prevent such outputs from qualifying as scientific-style reasoning:

  • Models typically do not maintain an explicit, persistent, and inspectable epistemic state: there is no ledger of hypotheses, confidence levels, weighed evidence, and outstanding questions that can be systematically updated as new information arrives.
  • There is no clear separation between what the system ‘‘knows’’ and how it manipulates that knowledge: knowledge and reasoning are entangled in the neural network weights rather than represented independently.
  • Although produced chains of thought look deliberative, research shows models often fabricate these chains after the fact—arriving at an answer by one internal route but reporting another as the ‘‘reasoning.’’

Why this matters in high-stakes domains

In medicine, engineering, and scientific research, stakeholders need not only reliable conclusions but also traceable justification for how conclusions were reached. When errors occur—such as in diagnosis or treatment—stakeholders must be able to identify the cause: faulty reasoning, use of invalid evidence, or incorrect assumptions?

This concern motivated Graepel’s recent departure from his post at Google DeepMind. He argues for a renewed approach to machine reasoning inspired by AlphaGo’s architecture. AlphaGo keeps a record of what it knows about a position in the form of a game tree: each variation and position is annotated with evaluations from its neural networks, and the system updates this structure as it reasons before synthesizing a decision.

What a general reasoning system should do

For general-purpose reasoning, a system should maintain an epistemic state that represents what is settled, what is doubtful, what has been ruled out, and which questions remain open. Reasoning can then be viewed as a sequence of moves that modify that state to advance knowledge and reduce uncertainty: deriving consequences, decomposing problems, and — crucially — choosing which question to ask, calculation to run, or experiment to perform next.

Open-world reasoning is harder than playing a board game: the current state is partially observed, available actions are numerous and variable, and action outcomes are stochastic or unknown. Yet advances in LLMs and neural models provide capabilities to tackle such problems: models can suggest strategies given available resources, interact with tools via APIs or code, and help assess whether claims are supported by evidence.

Most importantly, an independent component should evaluate each proposed ‘‘move’’ by how much it actually resolves uncertainty and only update beliefs when changes are backed by evidence. Enforcing such rules lets a system accumulate certified knowledge and improve its reasoning policy by learning from prior reasoning episodes. Graepel frames this as a supercharged scientific method intended to produce knowledge that can withstand scrutiny.

Conclusion

Graepel argues that merely scaling System 1 will not yield trustworthy machine intelligence: larger models sharpen intuition but do not make it deliberative. The breakthrough represented by AlphaGo’s move 37 came from holding a position, evaluating possible futures, and selecting a move that raw intuition would likely have rejected. Society needs similar kinds of creative, auditable moves in drug discovery, materials science, climate work, and diagnosis. According to Graepel, those insights will come from systems that genuinely reason—whose conclusions follow an auditable chain of evidence, inference, and belief revision rather than a persuasive story constructed after the fact.