The Illusion of Thinking

Summary: Apple researchers (Shojaee et al. 2025) show that Large Reasoning Models exhibit three distinct performance regimes — and a counter-intuitive scaling limit — when tested on controllable puzzles, suggesting current “reasoning” is sophisticated pattern matching rather than generalizable planning.

Sources: Academia/the-illusion-of-thinking.pdf

Last updated: 2026-05-06


Setup

Shojaee, Mirzadeh et al. (Apple, 2025) tested five frontier LRMs — o3-mini (medium and high), DeepSeek-R1, DeepSeek-R1-Qwen-32B, and Claude 3.7 Sonnet Thinking — on four controllable puzzle environments: Tower of Hanoi, Checker Jumping, River Crossing, and Blocks World. These puzzles allow fine-grained complexity control while maintaining consistent logical structure, and crucially allow verification of not just final answers but intermediate reasoning traces.

Prior evaluations on maths benchmarks suffered from data contamination and measured only final accuracy. This study examined how models think, not just whether they get the right answer.

The key models compared are thinking/non-thinking pairs: Claude 3.7 Sonnet with and without thinking, DeepSeek-R1 vs DeepSeek-V3.

Three regimes of complexity

The central finding — three regimes — holds consistently across all four puzzle types and all models tested:

Regime 1 — Low complexity: Standard LLMs (without extended thinking) outperform reasoning models at equivalent compute. The thinking mechanism adds overhead without adding value.

Regime 2 — Medium complexity: Reasoning models gain clear advantage. Extended chain-of-thought with self-reflection helps.

Regime 3 — High complexity: Both model types collapse to zero accuracy. Thinking models delay the collapse but cannot prevent it. Beyond the critical threshold, performance is zero regardless of how many tokens are available.

The counter-intuitive scaling limit

As complexity increases, LRMs initially allocate more thinking tokens. But as problems approach the critical collapse threshold, models reduce reasoning effort despite having ample token budget remaining. They give up before running out of compute.

This is not a context-length issue. It is a fundamental limit in how LRM thinking scales with compositional depth. The pattern is most pronounced in o3-mini variants; Claude 3.7 Thinking shows a less severe but similar curve.

Overthinking at low complexity

On simple problems, reasoning models often find the correct answer early in the thinking trace but then continue generating incorrect alternatives — wasting the token budget. This “overthinking” phenomenon represents real compute cost: the model doesn’t know when to stop.

Algorithms do not help

When the researchers provided the explicit recursive algorithm for Tower of Hanoi in the prompt (removing the need to find a solution strategy), performance did not improve. Collapse occurred at roughly the same point as in the standard condition.

This is the paper’s most striking result: the failure is not in problem-solving strategy but in step execution. Models cannot reliably follow prescribed logical steps across many sequential moves, even when the steps are given to them.

Puzzle-type inconsistency

Models handle Tower of Hanoi (requiring 100+ sequential moves at N=10) better than River Crossing (requiring only ~11 moves at N=3). This cross-puzzle inconsistency almost certainly reflects training-data exposure: Tower of Hanoi examples are abundant online; River Crossing at N>2 is rare. Performance tracks training exposure, not planning ability.

Implications

LRM “reasoning” is best understood as very sophisticated pattern matching. The three-regime finding aligns with the theoretical argument in jaeger-relevance-realization: genuine agency involves open-ended problem-solving that cannot be captured algorithmically — and these results show that something algorithmic is missing in LRMs at high complexity.

These limits are fundamental, not scaling issues. Throwing more tokens or more compute at high-complexity problems does not help. This is direct empirical evidence bearing on the questions raised in llm-and-mind and in borg-llm-meaning’s philosophical analysis.