News
DeepMind Researcher: AI Structurally Cannot Make Einstein's Leap

Tom Zahavy of Google DeepMind argues in a paper presented at ICML 2026 that language models have mastered theorem proving but structurally lack the capacity for the kind of creative leap Einstein made in conceiving general relativity.
Contents
Google DeepMind researcher Tom Zahavy has put forward a thesis that tempers the enthusiasm around AI's mathematical achievements: today's language models can prove complex theorems, but structurally cannot make the kind of creative leap that led Albert Einstein to general relativity in 1907.
Zahavy doesn't dispute that language models are making progress in mathematics. What he questions is something narrower and more fundamental: whether they can conceive of an entirely new way of looking at a problem before the data describing that problem even exists. In his view, this distinction determines whether AI will ever achieve a breakthrough on the scale of theoretical physics a century ago, or remain merely an excellent executor of frameworks devised by others.
Three types of reasoning
Zahavy's argument builds on philosopher Charles Sanders Peirce's classic distinction. Induction means finding patterns in data, deduction means logically deriving conclusions from accepted premises, and abduction means proposing an entirely new set of premises before it's even known whether they will hold up. Modern language models, the researcher argues, have mastered the first two modes impressively. The third, in his view, they cannot perform at all.
The elevator thought experiment
The paper's central example is the so-called Einstein elevator. In 1907, the physicist imagined an observer sealed inside a freely falling elevator who feels no force of gravity at all. Einstein called this his "happiest thought" - it became the seed of the equivalence principle, the foundation of general relativity. Zahavy poses the question directly: could today's AI model imagine such an experiment and independently derive a new physical principle from it. The paper's answer is unequivocal: no.
There is no logical path to these laws, only intuition resting on sympathetic understanding of experience can reach them - Albert Einstein
This line, which Einstein said on a different occasion, captures the essence of the problem Zahavy describes. Language models learn from text describing other people's experiences, but never live through those experiences themselves. What they lack, as the paper puts it, is the starting point from which physical intuition arises, the kind that leads to an entirely new way of seeing a familiar phenomenon.
Mathematical records nonetheless
The paradox is that the same generation of models Zahavy denies the capacity for abduction has, in recent months, been breaking records in deductive mathematics. OpenAI's GPT-5.6 Sol Ultra solved the Circle Double Covering Conjecture, and Anthropic's Fable 5 found a counterexample disproving the 87-year-old Jacobi conjecture. These are proof that the models can move deftly within existing mathematical and logical frameworks.
Zahavy doesn't deny the importance of these achievements, but he refuses to call them a creative breakthrough in the sense that Einstein's leap was. Solving a difficult conjecture within an existing mathematical apparatus is, in his view, still deduction at scale, not the proposal of a new set of axioms that didn't exist before.
A test for abstract reasoning
The paper cites results from the ARC-AGI-3 benchmark, designed specifically to test the ability to reason beyond learned patterns. AI systems solve fewer than 1 percent of its tasks, while humans score close to 100 percent. For Zahavy, this is a stark illustration of his thesis: scaling models improves performance on familiar task types, but doesn't close the gap in the ability to generate entirely new conceptual frameworks.
The publication feeds into a broader debate unfolding in parallel at this year's International Congress of Mathematicians, where Terence Tao spoke of mathematics possibly shifting from an era of "proof scarcity" to an era of "proof abundance" generated by machines. In such a scenario, the bottleneck would no longer be finding a proof but understanding and verifying it, a task that still requires a different kind of thinking than proving itself.
Implications for AI-for-science
The biggest consequences of Zahavy's thesis concern startups building models for autonomous scientific discovery, which assume that simply scaling up language models will eventually be enough to reach breakthroughs comparable to 20th-century theoretical physics. If Zahavy is right, the path to such breakthroughs doesn't run through ever-larger text models, but through embodied, multimodal world models capable of simulating physical intervention rather than merely describing it in words.
For the AI industry, this is an argument against treating each new record in solving open mathematical problems as proof of approaching general artificial intelligence. Zahavy's paper doesn't claim that models will stop getting more capable, only that proficiency in deduction and proficiency in inventing new theories are two different skills, and that the current architecture of language models mainly develops the former.


