LLMs Can't Jump Hits Hacker News, Challenging AI's Claim to Scientific Discovery
- Sophie Larsen

- Aug 6
- 14 min read
Google DeepMind researcher Tom Zahavy has challenged a central AI promise, despite rapid gains in formal reasoning. His paper, “LLMs Can’t Jump,” argues that current models cannot make the conceptual leaps behind major scientific discoveries. The claim reached Hacker News and reopened a difficult question about what machine reasoning actually accomplishes.
The paper does not deny that AI can prove theorems, search programs, or improve known algorithms. It separates those abilities from inventing the premises that make an entirely new theory possible. That distinction puts today’s AI scientist projects under more pressure than another disappointing benchmark would.
Zahavy uses Albert Einstein’s path to general relativity as the main test. A model might manipulate Einstein’s equations after receiving the right principles. The harder task is generating those principles before experimental evidence makes them obvious.
The Paper Draws a Line Between Proof and Discovery
The central claim is narrow but consequential: solving from supplied premises is not the same as inventing the premises.
Zahavy’s position paper divides reasoning into induction, deduction, and abduction. Each process serves a different role in scientific work.
Induction extracts patterns or general rules from examples. Modern language models perform a computational version of that process during training. They learn statistical relationships among tokens across enormous collections of human-produced material.
Deduction begins with accepted premises and derives their consequences. Formal mathematics offers a clean example. Once axioms, definitions, and a target conjecture are specified, a system can search for a valid proof.
Abduction moves in the opposite direction. It proposes a possible explanation for an observation, especially when the correct premises remain unknown. The explanation is not guaranteed to be true, but it creates something that deduction can later test.
Zahavy identifies abduction with the “jump” in his title. The term comes from a diagram Einstein sent to his friend Maurice Solovine. Einstein described discovery as movement from sensory experience toward axioms, followed by deductions that return to observable reality.
That first movement is not a routine inference from the available data. It introduces a conceptual framework that tells scientists which facts matter and how those facts belong together.
The paper argues that LLMs remain strongest after that framework exists. They can search combinations, expose contradictions, formalize arguments, and follow consequences. Those are valuable scientific functions, but they operate downstream from the decisive conceptual choice.
The distinction explains why the paper attracted attention beyond academic philosophy. Claims about automated science often combine several different capabilities under one label. Literature review, hypothesis generation, experimental planning, theorem proving, and paradigm creation all become “discovery.”
Zahavy rejects that compression. A system that generates thousands of variations within an established representation has not necessarily invented a new representation. A system that solves a stated problem has not necessarily decided which unstated problem deserves attention.
The argument also shifts evaluation away from polished scientific language. An LLM can describe tensors, gravity, and causal inference fluently. Fluency does not establish that its symbols connect to a manipulable model of physical reality.
The Hacker News discussion matters because developers regularly encounter the same distinction at a smaller scale. A coding model can resolve a defined issue while missing that the specification encodes the wrong product assumption.
In research, the cost of that limitation is larger. An automated system can optimize an elegant answer to a question that no longer deserves to organize the field.
Why General Relativity Is the Paper’s Hard Case
Einstein’s achievement matters here because the decisive theory arrived before a rich dataset made its structure easy to infer.
General relativity did not emerge from a large collection of failed Newtonian predictions. Newtonian mechanics described most observable gravitational behavior extremely well. Its success gave researchers little statistical pressure to replace its foundations.
Mercury’s unusual orbital precession was a prominent exception. Astronomers had already used missing objects to explain orbital anomalies, including the prediction of Neptune. Some researchers therefore proposed another planet, Vulcan, near the Sun.
That response preserved the existing framework. It treated the residual as a missing object rather than evidence that gravity required a new account.
An inductive optimizer would face a similar temptation. If a familiar theory produces low error across most observations, adding a local correction is cheaper than reconstructing space, time, and gravity.
Einstein followed a different path. His reasoning connected gravity with acceleration through imagined physical situations. His famous falling-observer and elevator scenarios helped motivate the equivalence principle.
The equivalence principle says that a local observer cannot always distinguish a gravitational field from acceleration. That proposition reorganized the problem before Einstein possessed the final mathematical machinery.
The paper treats this process as manipulative abduction. The thinker constructs a scenario, changes its conditions, and examines the consequences. The result is a proposed premise, not merely another prediction from established rules.
Einstein still needed years of mathematical work. He collaborated with Marcel Grossmann, explored tensor calculus, made wrong turns, and temporarily adopted the flawed Entwurf theory.
Those struggles do not weaken Zahavy’s distinction. They show how deduction, search, and verification can dominate after a new physical intuition defines the search space.
A capable AI assistant might have accelerated that later work. It might have checked assumptions, explored candidate tensors, or found inconsistencies sooner. None of those achievements explains how it would originate the equivalence principle.
This creates a demanding test for AI discovery claims. Remove general relativity from the training record, restrict knowledge to the relevant historical period, and ask a system to reconstruct the theory.
Success would require more than reproducing a known derivation. The system would need to identify the right conceptual conflict, invent useful thought experiments, and formulate new premises.
It would also need to avoid hidden leakage. Modern scientific language contains concepts shaped by relativity, even when a prompt excludes Einstein’s published work. Training data can quietly preserve the answer through later terminology and assumptions.
The paper therefore offers an argument, not a completed controlled experiment. General relativity is a case study used to expose a capability boundary. It does not prove that every possible computational system will fail.
That qualification matters. “Current LLMs lack an identified mechanism for this jump” is defensible. “No machine can ever perform abduction” would require a much stronger philosophical and technical case.
Zahavy’s target is primarily the language model as a disembodied statistical system. The paper leaves room for architectures that connect language, action, simulation, and causal intervention.
Hacker News Meets the AI Scientist Debate
The dispute pressures systems marketed as AI scientists because most operate within objectives and representations chosen by humans.
The Hacker News thread placed Zahavy’s argument before an audience that frequently tests frontier models in practical settings. That audience is familiar with both surprising model successes and brittle failures.
The disagreement is not simply whether LLMs are intelligent. The more useful question asks which parts of discovery remain supplied by researchers, benchmark designers, and software scaffolding.
Consider an automated research agent asked to improve an optimization algorithm. Humans usually define the benchmark, evaluation metric, programming language, available tools, and acceptable output format.
The agent can still find a genuinely new implementation. Its result can outperform existing human-written code and offer real economic or scientific value.
However, the system searches inside a world already made legible. The objective function tells it what counts as progress. The representation determines which candidate ideas can be expressed.
Google DeepMind’s AlphaEvolve illustrates both the strength and boundary of this approach. The system combines language models with automated evaluators and evolutionary search to improve algorithms.
According to Google’s AlphaEvolve announcement, the system found improvements across mathematics and computing tasks. These results deserve attention even if they do not settle the abduction question.
AlphaEvolve can generate candidate programs, evaluate them, retain productive changes, and repeat the process. This loop can discover solutions that were not explicitly written in its training material.
Yet the evaluator remains central. It supplies a machine-readable signal that ranks candidate programs. The system knows which direction constitutes improvement because researchers encoded that direction.
Zahavy’s challenge begins where that signal becomes uncertain. A scientific revolution can change the variables, standards, or ontology that define success. There may be no stable evaluator for the idea before the idea exists.
The same issue applies to automated paper-generation systems. An AI scientist can retrieve literature, propose hypotheses, run experiments, and draft a report. These steps compress substantial research labor.
Most such systems still inherit a domain, dataset, experimental protocol, and scoring rule. Their hypotheses often recombine concepts found in the supplied literature.
Recombination should not be dismissed as worthless. Human scientists also reuse established concepts, techniques, and instruments. Most important research extends a paradigm rather than replacing it.
The problem comes from inflating incremental automation into evidence of autonomous scientific invention. A system can produce novel results without displaying the specific kind of novelty under debate.
This is why “new” needs more precise definitions. A generated artifact can be absent from the training set, statistically unusual, useful, and patentable. It can satisfy all four conditions without creating a new explanatory framework.
Conversely, a scientific premise can sound simple after humans accept it. Its importance lies in reorganizing observations, not in producing an unusually complicated sentence.
Teams evaluating research agents should therefore trace where framing enters the workflow. Who selected the problem? Who chose the measurable target? Who supplied the representational language? Who decided that an anomalous output deserved investigation?
Those questions reveal whether the model is navigating a map or redrawing it.
Formal Reasoning Success Makes the Challenge Sharper
AI’s rapid progress in deduction strengthens Zahavy’s contrast because it shows how far systems can advance after humans formalize the problem.
Google DeepMind’s AlphaProof offers the clearest comparison. The system combines language-model capabilities with reinforcement learning in a formal proof environment.
A Nature paper reported olympiad-level formal mathematical reasoning from AlphaProof and AlphaGeometry 2. Their performance showed that machine reasoning can handle difficult, carefully specified mathematical problems.
Formal environments provide unusually reliable feedback. A proof assistant can verify whether each logical step follows from accepted rules. Incorrect proofs fail mechanically instead of receiving an ambiguous human judgment.
That verification loop supports large-scale search and learning. The model can explore candidates while the formal system prevents persuasive but invalid arguments from passing as solutions.
This is a major advance over ordinary chatbot reasoning. A chatbot can produce a plausible proof containing a subtle error. A formally checked system must construct an object that satisfies explicit logical constraints.
Zahavy does not deny this progress. His argument depends on taking it seriously.
Give a machine the right axioms, formal language, and conjecture, and its search abilities can become formidable. It may even find proofs that surprise expert mathematicians.
The open question is who creates the axioms and conjectures. Mathematics does not progress only by proving items from a prepared list. Researchers also invent definitions, choose abstractions, and decide which statements expose deeper structure.
These upstream choices resist simple evaluation. A new definition may initially appear unnecessary. Its value can emerge only after it connects several problems or makes a previously invisible regularity expressible.
Scientific premises face an even harder constraint. They must connect formal symbols to observations in the world. Their success depends on prediction, intervention, instrumentation, and causal interpretation.
An LLM trained on text receives the finished record of those connections. Papers describe which variables mattered after researchers selected them. Textbooks present mature theories without preserving every abandoned conceptual path.
Training on that record can make discovery look cleaner than it was. The model sees the language produced after a jump, not necessarily the experiences that made the jump possible.
This creates an evaluation trap. If researchers ask a model to rediscover a famous result, the model may reconstruct patterns embedded throughout later writing. If they ask for an unknown theory, nobody possesses an immediate answer key.
Automated verification works best when success is formal and local. Scientific invention often demands delayed, expensive, or uncertain validation. Experiments can take months, instruments can introduce errors, and competing theories can fit the same observations.
The difference does not make AI irrelevant. It changes where confidence belongs.
Models can search deductions, compare literatures, detect contradictions, and generate candidate experiments. Those capabilities can widen the range of hypotheses that human researchers inspect.
They can also help maintain the external memory that long investigations require. A searchable technical knowledge base can preserve rejected assumptions, experimental context, and links among documents.
That support becomes especially useful when discovery spans people and years. However, organizing evidence still does not decide which unexplained detail deserves a new theory.
The strongest near-term model is therefore collaborative. Humans and machines can divide work according to their verified strengths without pretending the division has already disappeared.
World Models Are the Proposed Bridge, Not a Proven Answer
Zahavy’s proposed remedy replaces passive text prediction with systems that can act inside consistent simulations and test counterfactual situations.
A world model represents how an environment changes over time. An action-controllable world model also lets an agent intervene and observe the resulting state.
That distinction matters for abduction. Watching many objects fall can reveal a pattern. Choosing to remove a support, alter gravity, or change an observer’s motion can reveal causal structure.
Zahavy argues that scientific invention requires this second relationship. A system needs more than descriptions of experiments. It needs a synthetic laboratory where proposed actions produce coherent consequences.
Google DeepMind’s Genie research points toward that direction. The original Genie paper described generative interactive environments learned from video.
Such systems can transform visual data into spaces that respond to actions. They offer a path beyond a language model that only discusses physical situations in tokens.
However, visual realism is not enough. A generated apple that falls convincingly may reflect frequent image sequences rather than an internal law of gravity.
A scientific world model must remain consistent under unusual interventions. It should preserve causal relationships when an agent creates conditions rarely represented in training data.
That requirement is demanding. Generative video models often prioritize plausible frames. Scientific simulation requires stable quantities, repeatable dynamics, and reliable counterfactuals.
The model also needs a way to translate simulation into concepts. Experiencing a regularity does not automatically produce an axiom that explains it.
This translation is the paper’s real bottleneck. An agent must choose which properties of a simulation deserve symbolic representation. It must ignore incidental details while retaining causal structure.
That selection process resembles the framing problem the architecture was supposed to solve. Developers can give a system predefined objects, actions, and state variables, but those choices embed human assumptions.
A richer latent environment reduces some constraints while introducing others. If the latent space reflects correlations from training videos, the model can still reproduce familiar appearances without discovering the governing mechanism.
Embodiment also remains ambiguous. Einstein did not need to jump from an actual roof or ride inside a falling elevator. His physical experience supported a mental simulation.
A machine might likewise acquire grounding through sufficiently consistent virtual interaction. Nothing in the argument establishes that biological sensation is uniquely capable of supporting abduction.
The important requirement is functional. The system must form interventions, predict their consequences, compare explanations, and create representations not handed to it in advance.
Even that list might be incomplete. Scientific judgment includes deciding which surprises matter, which simplifications remain acceptable, and when an unlikely premise deserves prolonged investigation.
Human priors shape those decisions. Einstein valued principles such as covariance and conceptual unity. Kepler’s commitments influenced his search for astronomical order.
A machine will also need priors. Designers must decide whether those preferences are learned, programmed, evolved, or derived through interaction.
Strong priors can drive productive exploration, but they can also create systematic blindness. An abductive machine could generate confident theories from misleading simulations.
This makes verification more important, not less. A system that proposes new premises needs aggressive tests that distinguish conceptual novelty from uncontrolled hallucination.
Current LLM safeguards focus heavily on factual accuracy within known knowledge. Scientific invention requires proposing statements that existing knowledge cannot already certify.
The desired system must therefore be imaginative and distrustful of its own imagination. Building both properties into one research loop remains unresolved.
“Can’t” Is the Paper’s Most Vulnerable Word
The paper identifies a real gap, but its strongest architectural conclusion exceeds the evidence supplied by one historical case study.
First, novelty is difficult to classify. A proposed hypothesis can combine existing concepts and still reorganize a field. Human invention rarely emerges without inherited language, instruments, and prior theories.
Calling all recombination induction risks defining abduction as whatever machines have not yet accomplished. That definition would move whenever a system produces a surprising result.
Second, language models do not operate alone in many modern systems. Tool use, memory, reinforcement learning, formal verification, search, and environmental interaction can change the overall capability.
A pure next-token predictor may lack a reliable abductive mechanism. An agent built around that predictor can still implement processes absent from the base model.
The relevant unit of analysis may therefore be the entire system. Saying “an LLM cannot jump” becomes less informative when the deployed researcher includes simulators, evaluators, instruments, and persistent memory.
Third, the Einstein example establishes a very high bar. Reconstructing general relativity from historical knowledge is closer to reproducing a scientific revolution than measuring ordinary scientific invention.
A system could contribute meaningful new hypotheses in chemistry, biology, or computer science without independently rebuilding the conceptual foundations of physics.
Demanding an Einstein-scale result before recognizing machine abduction would undercount smaller but authentic changes in representation.
Fourth, the paper’s diagnosis needs falsifiable evaluations. Researchers need tests where the relevant framework is absent from training, the environment supports intervention, and success cannot come from memorized terminology.
Those conditions are difficult to create. Synthetic scientific worlds could provide hidden laws, but success there might reflect the biases of their designers.
Real-world tests avoid some synthetic assumptions but make novelty and causality harder to score. They also create long feedback cycles and potentially serious safety concerns.
The best interpretation is therefore provisional. Current evidence does not show that language models can independently originate paradigm-level scientific premises. It also does not establish a permanent impossibility.
The paper is most persuasive when it asks researchers to disaggregate capabilities. It is less persuasive when “structurally incapable” sounds like a theorem about every future system containing an LLM.
That tension should shape product claims. Vendors should identify which stages their research agents automate and which stages still depend on human framing.
Researchers should publish full traces showing how objectives, prompts, tools, and evaluators influenced a reported discovery. A polished final paper cannot reveal those dependencies.
They should also compare model-generated hypotheses against strong search baselines. Evolutionary search, retrieval, and template recombination can produce surprising outputs without requiring a new reasoning category.
Finally, independent experts must judge whether a result changes the representation or merely improves performance inside it. That assessment will often remain debatable.
The Hacker News reaction is useful precisely because practitioners resist clean philosophical boundaries. They see models behave differently across contexts, prompts, tools, and levels of scaffolding.
Anecdotes from either side remain insufficient. Impressive outputs do not prove grounded understanding, while familiar failures do not prove an architectural impossibility.
What the Next AI Discovery Claims Must Show
The decisive evidence will come from systems that originate useful premises under controlled knowledge limits, not from better scientific prose.
The first signal to watch is a serious “historical cutoff” experiment. Researchers should train or constrain a system so later theories and terminology cannot leak into its evidence.
The system should then receive observations, instruments, and opportunities for intervention available before a major discovery. It must formulate a premise that was not represented in its starting materials.
If a system repeatedly succeeds under those controls, Zahavy’s claim weakens. If it only reconstructs theories when later language remains available, the paper’s diagnosis gains support.
The second signal is progress in action-controllable world models. These systems must remain physically consistent during unusual interventions, not simply generate convincing visual continuations.
Researchers should test whether agents can identify hidden variables, design discriminating experiments, and create compact causal representations. Success must transfer beyond the exact environments used for training.
Reliable transfer would support the idea that simulation can supply the grounding Zahavy finds missing. Failure outside familiar environments would suggest that visual interaction still masks induction.
The third signal is the provenance of claimed AI discoveries. Labs should disclose who defined the objective, selected the representation, built the evaluator, and recognized the result’s importance.
A system deserves stronger credit when it changes one of those upstream elements. Improving a metric within a human-authored benchmark remains valuable, but it answers a different question.
For developers and enterprise buyers, this distinction affects workflow design now. AI can accelerate search, formalization, comparison, and verification without becoming the final authority on framing.
Teams should keep human review closest to decisions about goals, assumptions, causal explanations, and acceptable evidence. Those are the points where a fluent answer can quietly narrow the problem.
Knowledge workers should also preserve the context behind conclusions. Source material, rejected hypotheses, and decision histories help people notice when an assistant is optimizing yesterday’s frame.
The “LLMs Can’t Jump” debate will not end with a single benchmark or Hacker News thread. Its value lies in demanding a cleaner standard for discovery.
When the next AI scientist claim arrives, ask what the system received before it began. Then ask whether it found a better answer or created a better question.
That distinction determines whether AI crossed Zahavy’s gap, or simply moved faster after humans built the bridge.


