The Cycle of Discovery

How do we fundamentally discover new things? In a letter to Maurice Solovine, Albert Einstein described discovery as a cycle: sensory experience leads, via an intuitive “jump,” to a system of axioms; logical deduction then generates theorems that can be tested against experience. Generative AI has become highly effective at two parts of this cycle. It performs induction—statistical pattern matching across large volumes of data – and it is rapidly advancing in deduction, the formal derivation of consequences from given premises. Systems such as AlphaProof have reached Olympiad-level performance in mathematics, and later models have attained gold-medal standards on similar problems.

Three Modes of Inference

Zahavy’s position paper argues that a third mode of inference—abduction—is missing. Drawing on Charles Sanders Peirce’s framework, he distinguishes the three forms by the structural roles of rule, case, and result:

  • Deduction applies a known rule to a case to produce a result. It preserves truth.
  • Induction extracts a rule from many cases and results. It generalizes from frequency.
  • Abduction invents a case (or a new rule) that would explain a surprising result. It is the creative generation of a hypothesis where none existed in the prior symbolic repertoire.

The Case of General Relativity

Einstein’s formulation of general relativity serves as the central case study. At the time, Newtonian gravity faced no empirical crisis that demanded a new theory. The equivalence of inertial and gravitational mass had been verified to high precision; the sole notable anomaly, the advance of Mercury’s perihelion, was widely attributed to an undiscovered planet rather than a failure of the laws themselves. There was therefore no substantial error signal or supervised dataset that an inductive system could compress into a better model. The axioms of the new theory could not be deduced from existing premises either, because they were the premises. The decisive step was abductive: the “happiest thought” in which Einstein imagined an observer in free fall and recognized that the local effects of gravity and uniform acceleration are indistinguishable. That insight, grounded in embodied mental simulation rather than in linguistic or statistical manipulation of existing texts, supplied the equivalence principle and the subsequent geometric axioms.

Where Current Models Succeed—and Stop

Once those axioms are supplied, the subsequent mathematical labor—identifying the correct curvature tensors, correcting intermediate errors about the Newtonian limit, and deriving the field equations—falls within the reach of modern formal-reasoning systems. Zahavy concedes that a contemporary large language model, given Einstein’s physical assumptions as input, could plausibly complete the deductive phase. The bottleneck lies upstream: the translation of simulated sensory experience into formal axioms that have no prior symbolic precedent.

The Limits of Symbol Manipulation

Current language models operate as high-dimensional symbol manipulators. They recombine and optimize within an established linguistic space. Techniques such as prompt engineering or iterative refinement improve performance inside that space but cannot supply the sensory grounding required for manipulative abduction—the active interaction with a physical (or physically consistent) model that generates new conceptual primitives. Zahavy likens this limitation to the Chinese Room: the system can rearrange the language of physics without access to the referents that give the language meaning.

Toward Physically Consistent World Models

The paper therefore identifies physically consistent, action-controllable world models as a necessary substrate. Passive video predictors that merely continue statistical patterns are insufficient. Models that permit counterfactual intervention—allowing an agent to “cut the cable” of an elevator or otherwise manipulate the simulated environment—could function as synthetic laboratories. Sensory and interactive feedback from such environments would provide the grounding needed to propose axioms where no linguistic template yet exists. The proposal is framed specifically for the physical sciences, where the object of study is external material reality; in more abstract domains the form of the required simulation would differ, though the necessity of an abductive jump remains.

A Position, Not a Verdict

Zahavy presents the argument as a position paper, not as a claim that large language models are incapable of useful scientific contribution or that scaling alone cannot produce further advances. The analysis focuses on the structural conditions under which a system could replicate the particular kind of paradigm-forming leap illustrated by general relativity. The critical missing step is the jump from experience to axiom; without a mechanism for that step, the cycle of scientific invention remains incomplete.

The paper is available as a preprint on the PhilSci-Archive (Zahavy, 2026).

Leave a comment

Your email address will not be published. Required fields are marked *