Why AlphaGo’s Reasoning Still Outperforms Today’s LLMs—and What’s Missing
In 2016, AlphaGo’s victory over Go champion Lee Sedol showcased a breakthrough in AI reasoning, combining neural networks with deliberative search. Today, large language models (LLMs) dominate the AI landscape, but they still lack the genuine reasoning capabilities that made AlphaGo revolutionary. Former Google DeepMind researcher Thore Graepel argues that the future of AI depends on revisiting AlphaGo’s hybrid approach—and fixing what LLMs get wrong.
Editor, Lazyfounder

In 2016, AlphaGo’s victory over Go champion Lee Sedol showcased a breakthrough in AI reasoning, combining neural networks with deliberative search. Today, large language models (LLMs) dominate the AI landscape, but they still lack the genuine reasoning capabilities that made AlphaGo revolutionary. Former Google DeepMind researcher Thore Graepel argues that the future of AI depends on revisiting AlphaGo’s hybrid approach—and fixing what LLMs get wrong.
30 SEC SUMMARY
- AlphaGo’s 2016 victory over Go champion Lee Sedol highlighted its ability to combine intuitive pattern recognition with deliberative reasoning, a feature missing in today’s large language models (LLMs).
- Current LLMs rely on fast, associative pattern completion (System 1) but lack genuine reasoning capabilities like AlphaGo’s deliberative search (System 2).
- Thore Graepel, former Google DeepMind researcher, argues for a new approach to machine reasoning inspired by AlphaGo’s architecture, emphasizing auditable and evidence-based AI.
- Chain-of-thought reasoning in LLMs often fabricates intermediate steps after reaching conclusions, rather than performing true deliberation.
- Trustworthy AI in high-stakes fields requires explicit epistemic states, separable knowledge and reasoning, and auditable inference processes.
TABLE OF CONTENTS
- AlphaGo’s Hybrid Reasoning
- Limitations of Current LLMs
- A New Approach to Machine Reasoning
- Background: The Evolution of AI Reasoning
- What this means
- Key takeaways
- FAQ
- Sources
KEY HIGHLIGHTS
- AlphaGo defeated Go champion Lee Sedol 4-1 in a five-game match in 2016, with move 37 in game two standing out as a creative and strategic breakthrough.
- AlphaGo combined intuitive pattern recognition (System 1) with deliberative search (System 2) to evaluate thousands of possible moves and their consequences.
- Current LLMs rely on fast, associative pattern completion but lack a separate reasoning mechanism like AlphaGo’s search.
- Chain-of-thought reasoning in LLMs often fabricates intermediate steps after the fact, rather than performing true deliberation.
- Thore Graepel left Google DeepMind to advocate for a new approach to machine reasoning, inspired by AlphaGo’s architecture.
- Trustworthy AI requires systems with auditable reasoning, evidence-based belief revision, and the ability to resolve uncertainty.
AlphaGo’s Hybrid Reasoning
In March 2016, AlphaGo, developed by Google DeepMind, defeated Lee Sedol, one of the greatest professional Go players of all time, in a five-game match held in Seoul. The victory was a milestone in artificial intelligence, particularly because of move 37 in game two. Initially perceived as an absurd or mistaken move, it was later recognized as a creative and strategic breakthrough.
Unlike Deep Blue, which defeated chess champion Garry Kasparov in 1997 by evaluating 200 million positions per second using human-coded rules, AlphaGo employed a hybrid approach. It combined a neural network trained to mimic expert human moves with a deliberative search mechanism that explored thousands of possible future moves and their consequences.
According to MIT Technology Review, AlphaGo’s policy network estimated that move 37 had a roughly one-in-10,000 chance of being played by a human expert. However, its search machinery looked beyond immediate plausibility, weighing the long-term consequences of the move by constructing and evaluating a game tree with thousands of branches.
Limitations of Current LLMs
Large language models (LLMs) like ChatGPT rely on fast, associative pattern completion, akin to what psychologist Daniel Kahneman describes as "System 1" thinking—automatic and intuitive. However, they lack a genuinely separate reasoning mechanism, or "System 2," which involves deliberate, effortful problem-solving.
Chain-of-thought reasoning, often touted as a solution to improve LLM outputs, does not introduce a separate reasoning process. MIT Technology Review reports that LLMs often concoct intermediate reasoning steps after reaching a conclusion, rather than using deliberation to arrive at it. This raises questions about the reliability of their outputs, especially in high-stakes applications like medicine, engineering, or scientific research.
Unlike AlphaGo, which maintained a transparent game tree to track considered moves, positions, and neural network judgments, LLMs lack explicit epistemic states—representations of settled knowledge, doubts, or ruled-out possibilities. This opacity makes it difficult to audit their reasoning or verify their conclusions.
A New Approach to Machine Reasoning
Thore Graepel, Chair of Machine Learning at University College London and a former Google DeepMind researcher, has left the company to advocate for a fresh approach to machine reasoning. His argument centers on the need for AI systems that combine the intuitive strengths of neural networks with deliberative reasoning mechanisms inspired by AlphaGo’s architecture.
Graepel suggests that trustworthy AI requires systems with auditable inference processes, evidence-based belief revision, and the ability to resolve uncertainty. For example, an independent component of a reasoning system should evaluate proposed moves or hypotheses based on how much they reduce uncertainty, updating beliefs only when supported by evidence.
However, applying this approach to open-world problems—such as medical diagnosis or scientific research—poses unique challenges. Unlike board games, where the state of play is fully known and rules are fixed, real-world scenarios involve partial observability, variable actions, and stochastic or unknown consequences.
Background: The Evolution of AI Reasoning
AlphaGo’s victory over Lee Sedol marked a turning point in AI, demonstrating that machines could master complex, intuitive games like Go by combining neural networks with search-based reasoning. This hybrid model contrasted with earlier AI systems, such as Deep Blue, which relied on brute-force computation and human-coded rules.
The limitations of current LLMs in reasoning have sparked debates about the future of AI development. While LLMs excel at pattern recognition and associative tasks, their lack of genuine deliberative reasoning raises concerns about their suitability for applications requiring transparency, accountability, and accuracy.
Recent advances in AI, such as Google DeepMind’s SynthID Bio for watermarking AI-designed proteins, highlight the growing role of AI in fields like biosecurity and scientific research. However, these applications also underscore the need for reliable, auditable reasoning in AI systems.
What this means
Lazyfounder analysis — our interpretation, not reported fact.
Graepel’s critique highlights a fundamental tension in AI development: the gap between what LLMs appear to do and what they actually achieve. For founders and operators building AI-driven products, this isn’t just an academic concern—it’s a practical one. If your product relies on LLMs for tasks like decision support, creative problem-solving, or data analysis, the lack of genuine reasoning could introduce risks, from subtle errors to catastrophic failures in high-stakes domains.
The lesson from AlphaGo’s success is that hybrid systems—those combining neural networks with structured, auditable reasoning—may be the key to building AI that is both powerful and trustworthy. For startups, this suggests an opportunity: rather than treating LLMs as black-box solutions, integrating them with deliberative reasoning mechanisms could unlock new use cases in fields like healthcare, law, or scientific research. The challenge, of course, is that building such systems is non-trivial. It requires not just technical innovation but also a shift in how we evaluate AI performance—prioritizing transparency and reliability over fluency and speed.
For now, founders should be cautious about overestimating what LLMs can do. While they are invaluable tools for pattern recognition and content generation, they are not a substitute for genuine reasoning. If your product relies on AI for critical decisions, consider supplementing LLMs with rule-based systems, human oversight, or other mechanisms to ensure accuracy and accountability.
Key takeaways
- AlphaGo’s 2016 victory over Lee Sedol demonstrated a hybrid approach to reasoning, combining neural networks with deliberative search.
- LLMs lack genuine reasoning capabilities, relying instead on pattern completion and post-hoc rationalization.
- Experts like Thore Graepel advocate for AI systems with auditable, evidence-based reasoning to improve reliability in critical applications.
- Current AI models miss key features like explicit epistemic states and separable knowledge and reasoning mechanisms.
- Open-world reasoning poses unique challenges, including partial observability and stochastic consequences, which board-game AI like AlphaGo did not address.
FAQ
Why was AlphaGo’s move 37 in game two significant?
Move 37 was initially perceived as absurd or mistaken, but it later proved to be a creative and strategic breakthrough. It demonstrated AlphaGo’s ability to innovate beyond human intuition, as its policy network estimated the move had a one-in-10,000 chance of being played by a human expert. The move highlighted the power of combining neural networks with deliberative search to explore unconventional strategies.
How do current LLMs differ from AlphaGo in terms of reasoning?
AlphaGo combined intuitive pattern recognition (System 1) with a deliberative search mechanism (System 2) that evaluated thousands of possible moves and their consequences. In contrast, current LLMs rely primarily on fast, associative pattern completion (System 1) and lack a separate, auditable reasoning process. Chain-of-thought reasoning in LLMs often fabricates intermediate steps after reaching a conclusion, rather than performing genuine deliberation.
What are the limitations of chain-of-thought reasoning in LLMs?
Chain-of-thought reasoning in LLMs does not introduce a separate reasoning mechanism. Instead, it often generates intermediate steps after arriving at a conclusion, which can lead to fabricated or unreliable rationalizations. This undermines trust in the model’s outputs, particularly in high-stakes applications where transparency and accuracy are critical.
What does Thore Graepel propose as a solution to improve AI reasoning?
Graepel advocates for a new approach to machine reasoning inspired by AlphaGo’s architecture. He argues that trustworthy AI requires systems with auditable inference processes, explicit epistemic states (representing knowledge, doubts, and possibilities), and evidence-based belief revision. This approach would enable AI to resolve uncertainty and update beliefs only when supported by evidence, making it more reliable for real-world applications.
Why is open-world reasoning more challenging than board-game AI like AlphaGo?
Open-world reasoning involves scenarios where the current state of affairs is partially known, actions are variable, and consequences are stochastic or unknown. In contrast, board games like Go or chess have fully observable states, fixed rules, and clear outcomes. This makes open-world reasoning significantly more complex, as AI systems must account for uncertainty, partial information, and dynamic environments.
Related on Lazyfounder
Sources
- MIT Technology Review · 2026-10-02
Don’t be fooled—LLMs don’t reason
This story is an original summary drafted with AI by Lazyfounder from the reporting listed above and checked by automated validation. Facts are attributed to their original publishers; sections marked as analysis are Lazyfounder's. Where a source is in another language, facts were machine-translated and quotations are reported, not reproduced. Read the original coverage via the links, and see our AI policy and corrections policy.
About the author
Editor, Lazyfounder
Tarun Mottlia edits LazyFounders, covering Indian startups, funding rounds, AI and product launches. Every story on the site is AI-assisted and checked against its cited sources before publication.
More stories by Tarun MottliaGet the LazyFounder Brief
Startup, funding and AI news in a five-minute read. Join the early-access list.


