Site navigation

New Apple Paper Pours Cold Water on AI Reasoning

Graham Turner

,

AI reasoning limits
A new research paper from Apple casts doubt on the much-touted reasoning abilities of today’s top AI models, suggesting they often rely on pattern-matching rather than real logic.

As tech giants race to develop artificial intelligence that can “think” like humans, a new study from Apple researchers delivers a sobering reality check: The much-hyped reasoning abilities of today’s most advanced AI models may be nothing more than an elaborate illusion.

The paper, titled The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models, examines how Large Reasoning Models (LRMs) – such as OpenAI’s o1 and o3, Google’s Gemini Flash Thinking, and Anthropic’s Claude 3.7 Sonnet – perform when faced with increasingly complex logic puzzles.

The findings reveal that these models don’t truly reason (in their current form, anyway). Instead, they rely on pattern recognition and statistical guesswork – until they hit a wall and fail completely.

The Breaking Point of AI “Thinking”

To test the limits of AI reasoning, Apple’s researchers used classic logic puzzles like the Tower of Hanoi, River Crossing, and block-stacking challenges – tasks that require multi-step planning, recursion, and strict adherence to rules. These puzzles are simple for humans to learn, but become progressively harder as complexity increases (such as adding more discs to the Tower of Hanoi).

The results showed clear differences across complexity levels. For the simpler stuff, standard AI models (LLMs) outperformed specialised “reasoning” models (LRMs), suggesting that the additional “thinking steps” introduced by LRMs may sometimes reduce performance.

For medium-complexity problems, LRMs demonstrated a relative advantage, solving multi-step tasks more effectively – though only to a point.

At high complexity, both LLMs and LRMs went full 4:30pm-on-a-Friday: total accuracy collapse, failing entirely despite having more than enough computing power.

Even more telling was how the models behaved under pressure. As puzzles grew harder, the AI initially expended more effort (using more “thinking tokens”) – but then, in an arguably very human manner, it gave up early, reducing effort just as the problems became most challenging – though the process that leads to this abandonment of the task is quite different to how we get there.

Guessing. 

One of the paper’s most interesting (and potentially damning) findings was that even when researchers gave the models the correct algorithm – essentially handing them a step-by-step solution – they still failed. This suggests that the issue isn’t just a lack of knowledge, but a fundamental inability to follow structured logic.

“Current evaluations focus on final answer accuracy in math and coding, but they don’t reveal how models arrive at answers,” the researchers wrote. “Our study shows that what looks like reasoning is often just pattern matching – and when patterns break down, so does the AI.”

Why This Matters for the Future of AI

The study arrives at a pivotal moment in AI development. Companies like OpenAI, Google, and Anthropic have heavily marketed their models as capable of “human-like reasoning,” fuelling speculation about near-term artificial general intelligence (AGI) – even though no one can agree what that actually means.

Irregardless, Apple’s research suggests that today’s AI, no matter how fluent, still lacks true understanding and is long way away from anyone’s definition of AGI.


Recommended reading


The paper also raises questions about benchmarking. Most AI evaluations rely on math and coding tests where answers can be memorised or statistically predicted. But in controlled puzzles requiring genuine reasoning, the models falter – exposing a gap between performance and true intelligence.

Is it a coincidence that Apple’s come out with a paper that rains on the AI parade as the company noticeably lags behind in the AI race? Probably not.

That doesn’t change the fact that while AI can simulate reasoning – it doesn’t truly reason. That’s okay, it doesn’t mean the technology is junk – it just means if you read something about AI that has the word “reasoning” in it, give it a second glace to make sure you know what you’re being sold.

Graham Turner

Sub Editor

Latest News

AI

Nvidia Launches Open Secure AI Alliance for AI Safety and Security

AI Business Recruitment

Nearly a Quarter of Orgs Reducing Entry-level Hiring Due to AI Automation

Business

Scottish Businesses Turn to Self-funding as Growth Confidence Dips in H2

Data Finance

Payment Leaders are Struggling to Get Real-time Data