MIT Technology Review Essay Says LLMs Do Not Truly Reason Like AlphaGo
MIT Technology Review essay argues LLMs do not truly reason, contrasting next-token prediction with AlphaGo's search and intuition.
The author recalls watching AlphaGo play move 37 in the second game of the five-game match in Seoul in March 2016. The move seemed absurd enough that some commentators thought it was a programming glitch. AlphaGo won the game and defeated Lee, one of the greatest professional Go players, 4-1. “I thought AlphaGo was based on probability calculation and that it was merely a machine,” Lee said afterward, according to the essay. “But when I saw this move, I changed my mind. Surely, AlphaGo is creative.”
The essay says that view misses what happened. In 1997, Deep Blue beat Garry Kasparov by looking six to eight moves ahead and evaluating 200 million chess positions per second with hard-coded rules. Go is much more complex; calculating even a fraction of possible outcomes would take a supercomputer billions of years. AlphaGo had to sense who was ahead and invent moves humans had not played. The author argues move 37 came not from pure intuition but from reasoning.
AlphaGo, according to the essay, combined two systems. Its policy network was trained to guess what a strong human would play. That intuitive part saw move 37 as unremarkable, giving it roughly a one in 10,000 chance of being played by an expert human. Its search machinery chose the move by constructing and searching a game tree with thousands of branches, each a possible future. The essay compares this to Daniel Kahneman's System 1 and System 2: the networks supplied hunches, while search supplied deliberation. Neither half worked alone. Intuition would not have chosen move 37, and brute-force search would have struggled to sort through the possibilities.
Today's AI models work differently, the essay says. A large language model picks the next token repeatedly, which amounts to System 1 in action: fast, associative, and good at pattern completion across many subjects. After ChatGPT's debut, the field recognized that language fluency alone was not enough. Models were made to deliberate by generating intermediate steps that decompose a problem, carry partial results, and influence later reasoning—a process known as chain of thought. The gains are real, especially in mathematics and coding, the essay says. But it argues this is not a genuinely separate reasoning mechanism. The intermediate reasoning is still produced by the same next-token prediction process, iterated longer before the model answers.
The essay lists three shortcomings. First, these models typically maintain no explicit, persistent, and inspectable epistemic state: no open ledger of hypotheses, confidence in explanations, evidence being weighed, and unresolved questions that would be revised as new information arrives. Second, they lack a clean separation between what the system knows and how it manipulates that knowledge; knowledge and reasoning are interwoven in neural network weights, with no independent, explicitly represented set of beliefs. Third, research has shown that chatbots often produce chains of thought after the fact, reaching an answer by one route but reporting another. The essay says this matters in high-stakes fields such as medicine, engineering, and scientific research, where how a system reaches a conclusion matters as much as the conclusion itself.
The author concludes that if future AI systems are to produce trustworthy results and novel insights in science and medicine, they need genuine reasoning capabilities of the kind AlphaGo displayed. That requires mechanisms beyond longer chains of next-token prediction.