top of page

Jihao Liu Says AI Disproved the YTD Conjecture, but the Real Result Is More Precise

Aug 23
11 min read

Jihao Liu says generative AI helped construct the first counterexample in his 79-page YTD preprint, despite decades of work supporting versions of the conjecture. The paper presents a smooth projective fivefold that is K-polystable but allegedly lacks a constant scalar curvature Kähler metric. If the proof survives expert review, that object breaks the classic general formulation of the Yau-Tian-Donaldson conjecture.

Liu submitted the paper to arXiv on August 19, 2026, at 17:19 UTC. The abstract credits GPT-5.6-sol, Fable 5, and the Danus system with obtaining the main result. An appendix written with Bin Dong and Guoxiong Gao discusses how generative AI was used.

That disclosure explains why the story spread quickly. It does not settle the mathematics by itself. The paper is a new preprint, not a peer-reviewed publication, and the claimed counterexample requires specialist checking across several technical layers.

The deeper story is also narrower than the viral summary that “AI disproved YTD.” The result targets the equivalence between ordinary K-polystability and the existence of constant scalar curvature Kähler metrics for general polarized varieties. Important YTD theorems for Fano varieties remain intact, while stronger stability formulations were already replacing the classic statement.

What the AI-Assisted YTD Preprint Actually Claims

The central claim is a specific mathematical mismatch, not a general declaration that decades of Kähler geometry were wrong.

The Yau-Tian-Donaldson conjecture connects two different descriptions of a geometric object. One side asks whether a polarized variety admits a canonical metric. The other asks whether that variety satisfies an algebraic stability condition.

A polarized variety is a projective variety paired with an ample line bundle, which supplies the embedding data used in algebraic geometry. The metric in question has constant scalar curvature, meaning its scalar curvature stays uniform across the underlying Kähler manifold.

The conjecture’s appeal comes from this bridge. A difficult nonlinear analytic existence problem would become equivalent to an algebraic criterion called K-polystability. Mathematicians could study canonical geometry through degenerations and numerical invariants instead of solving a differential equation directly.

Liu’s YTD preprint claims to separate those two sides. It constructs a polarized smooth projective fivefold, a complex projective variety of complex dimension five, that passes the K-polystability test but does not admit the expected constant scalar curvature Kähler metric.

A single valid example is enough. An equivalence fails when one object satisfies the proposed condition but lacks the promised consequence.

The distinction between “K-polystable” and “has no constant scalar curvature metric” therefore carries the whole result. If either half fails, the paper does not disprove the classic conjecture. Reviewers must check the construction, the stability calculation, the analytic obstruction, and the logical connection between them.

The preprint’s date matters because the Bilibili trend appeared only days after the arXiv submission. As of August 23, the paper has not had time to pass conventional journal review. Its presence on arXiv establishes that the claim was publicly posted, not that the argument has been certified.

The author is not an anonymous social media account. Liu is an algebraic geometer at Peking University and a member of the Beijing International Center for Mathematical Research. His research profile lists birational geometry, singularities, foliations, and explicit algebraic geometry among his main interests.

That background raises the claim above an unsupported chatbot output. It still leaves a verification gap. Professional authors can make mistakes, AI-generated arguments can hide subtle errors, and arXiv performs moderation rather than full peer review.

The correct headline is therefore conditional. Liu and his paper claim a disproof, while independent mathematical consensus is still forming.

Why YTD Was Already More Complicated Than One Conjecture

The phrase “the YTD conjecture” now covers several statements with different scopes, stability notions, and established results.

The broad idea developed from work associated with Shing-Tung Yau, Gang Tian, and Simon Donaldson. It predicts that the existence of canonical Kähler metrics should correspond to algebraic stability.

That principle has produced major theorems. It is not a single untouched proposition that survived unchanged until one AI system found a counterexample.

For Fano varieties, important YTD formulations connecting K-polystability with Kähler-Einstein metrics have been proved. A Fano variety has an ample anticanonical bundle, placing it in a special geometric setting with additional structure. Liu’s claimed example does not erase those theorems.

The new preprint concerns the broader constant scalar curvature problem for general polarized varieties. This setting is harder because the relevant analytic and algebraic spaces are more complicated. The appropriate definition of stability has also remained central to the debate.

Ordinary K-stability tests a polarized variety through algebraic degenerations called test configurations. Each configuration receives a numerical invariant. Roughly, negative behavior reveals a destabilizing degeneration, while polystability allows only specified trivial cases at the boundary.

The problem is whether this collection of algebraic tests captures every analytic way that a metric can fail to exist. Liu’s proposed example says it does not. The variety passes ordinary K-polystability, yet an analytic obstruction remains.

That possibility did not arrive without warning. Researchers have developed stronger notions, including uniform K-stability and non-Archimedean formulations. These conditions quantify stability rather than merely requiring a nonnegative invariant.

Sébastien Boucksom and Mattias Jonsson described a version based on uniform modified K-stability in their 2026 non-Archimedean approach. Their framework works with a completed space of non-Archimedean metrics, extending beyond classical test configurations.

A separate May 2026 preprint presented a uniform formulation connecting constant scalar curvature metrics with an automorphism-adjusted form of uniform K-stability. That work illustrates why experts may view Liu’s result as a correction to the stability condition, not the collapse of the entire YTD program.

This is the key interpretive divide. The viral version says AI defeated a famous conjecture. The technical version says ordinary K-polystability appears too weak for the most general constant scalar curvature problem.

Those statements differ in scope. The first attracts attention, while the second tells mathematicians what must change.

A valid counterexample would show that the original algebraic test omitted some destabilizing behavior. It would strengthen the case for completed, uniform, or filtration-based stability notions that can see more than standard test configurations.

The result could therefore support the larger YTD philosophy. Algebraic stability may still characterize canonical metrics, but the correct algebraic definition must be stronger.

The Real Opponent Is Ordinary K-Polystability

This paper pits ordinary K-polystability against stronger stability frameworks, not AI against every human proof of YTD.

The classic formulation promises that K-polystability is sufficient for a constant scalar curvature Kähler metric. Liu’s fivefold allegedly exposes a case where that promise fails.

The stronger frameworks respond by enlarging what counts as a destabilizing direction. Instead of checking only conventional algebraic test configurations, they examine filtrations, non-Archimedean metrics, or quantitative lower bounds.

That difference can sound abstract, but its function is straightforward. A stability test is useful only when it detects every obstruction relevant to metric existence. A test that misses one class of degenerations can label an object stable when the analytic geometry says otherwise.

Liu’s construction reportedly exploits exactly this gap. The polarized fivefold is K-polystable under the ordinary definition. However, the paper argues that it fails the stronger uniform condition needed to guarantee the metric.

If that argument is correct, the counterexample does not leave the field without a replacement. It identifies the boundary between a weaker condition and a condition designed to capture the missing analytic behavior.

This is why the dimensions and category matter. A fivefold is far removed from the simplest surfaces or threefolds that often guide geometric intuition. Higher-dimensional constructions can combine bundles, symmetry, and intersection data in ways that lower-dimensional examples cannot.

The paper’s example also does not imply that ordinary K-polystability is useless. It remains necessary in broad settings and sufficient in important special cases. The alleged failure concerns sufficiency at full generality.

Think of the distinction as a diagnostic test. A negative result can reliably identify a problem, while a positive result still misses one rare condition. The test retains value, but its positive prediction needs a stronger threshold.

That framing also explains why earlier positive results are not automatically contradicted. A theorem proved under a Fano assumption, a discreteness condition, or uniform stability remains a theorem. A general counterexample only defeats statements whose hypotheses include the new example.

Mathematical news often loses these quantifiers. “For every polarized smooth projective variety” becomes “YTD,” while “K-polystable” becomes “stable.” Once those details disappear, a correction to one formulation sounds like a rejection of the whole research program.

The most consequential outcome would be a revised consensus: ordinary K-polystability is necessary but insufficient, while an enhanced non-Archimedean or uniform condition gives the correct equivalence.

That would still be a major result. It would redirect proofs, examples, and moduli questions toward the stronger invariant. It would also clarify why earlier attempts to prove the ordinary formulation encountered persistent technical barriers.

The counterexample could become a benchmark object. Researchers would ask which stability definitions reject it, which invariants reveal the obstruction, and whether related examples exist in lower dimensions.

Those questions matter more than whether the event counts as a victory for one AI vendor. The lasting mathematical value lies in diagnosing the missing condition.

How GPT-5.6-sol, Fable 5, and Danus Entered the Work

The paper presents AI as part of a research system, while a named mathematician remains responsible for the published proof and its claims.

The abstract specifically identifies GPT-5.6-sol, Fable 5, and Danus. That level of disclosure is unusually direct compared with papers that acknowledge only generic language-model assistance.

GPT-5.6-sol and Fable 5 are described as generative models used in obtaining the main result. Danus is a software framework for coordinating AI-driven mathematical exploration. Its public Danus repository provides a concrete artifact that researchers can inspect and adapt.

A multi-model workflow can divide mathematical labor into candidate generation, literature retrieval, symbolic manipulation, criticism, and revision. One model proposes a construction, another attacks its assumptions, and an orchestrator maintains the evolving state.

That process differs from asking a consumer chatbot one question. Long mathematical arguments require stable notation, persistent goals, and repeated checking. An orchestration layer can preserve intermediate claims and route them through multiple attempts.

However, orchestration does not turn model output into a theorem. Language models can reproduce plausible definitions while reversing a quantifier, applying a result outside its hypotheses, or overlooking a singular case.

The fivefold construction must therefore stand independently of its discovery process. Readers should be able to verify every lemma without trusting the models that suggested it.

This principle also prevents two opposite mistakes. Skeptics should not reject a proof merely because AI helped discover it. Enthusiasts should not accept the proof merely because several advanced models participated.

Mathematics ultimately judges the written argument. A correct proof remains correct regardless of whether its first idea came from a person, a computer search, or a language model.

Attribution is harder. The arXiv entry names Liu as the paper’s author, while Dong and Gao join the appendix about AI use. The models are credited as tools rather than authors because they cannot accept responsibility, resolve editorial questions, or answer objections.

The workflow still signals a change in research practice. A specialist can now ask models to survey adjacent constructions, combine techniques, and search for exceptions at a scale that would consume substantial human time.

Counterexamples are especially compatible with this style. The system needs one object satisfying a finite collection of properties. Once proposed, that object can sometimes be checked through explicit calculations and established theorems.

A universal proof presents a different burden. It must cover every object in a class and control every exceptional case. A fluent but incomplete argument can survive many review passes before its hidden gap becomes visible.

Liu’s claim sits between these categories. The existence of one fivefold would refute the equivalence, but proving that its analytic metric does not exist requires more than displaying a simple numerical witness.

The work therefore should not be described as an automated search producing an instantly checkable answer. It is a long geometric argument with AI-assisted discovery, synthesis, and drafting, according to the paper.

This distinction matters for enterprises evaluating AI research agents. The relevant capability is not unattended truth production. It is accelerated hypothesis generation paired with expert verification and auditable records.

Teams adopting similar workflows need to preserve prompts, model versions, intermediate outputs, failed branches, and human corrections. A searchable AI knowledge base can help maintain that provenance, although it cannot validate the mathematics.

The strongest lesson is operational. Models can widen the search space, but accountability stays with the researchers who select, check, and publish the result.

What the YTD Claim Still Does Not Prove

A new arXiv submission can establish priority for a claim, but it cannot substitute for independent expert verification.

As of August 23, only four days separate the submission date from the present analysis. That is too little time for the usual cycle of seminar discussion, line-by-line checking, revisions, and journal review.

The most obvious uncertainty is whether the fivefold is genuinely K-polystable under the exact definition used by the classic conjecture. Stability calculations can depend on all possible test configurations, not merely an obvious family.

A second uncertainty concerns the claimed absence of a constant scalar curvature Kähler metric. The obstruction must apply to the specified polarization and survive every relevant symmetry or automorphism issue.

A third concerns the bridge between ordinary and uniform stability. If the proof silently imports a theorem with a stronger hypothesis, the construction may expose a definitional mismatch rather than a counterexample.

These are not accusations of error. They are the standard questions raised by an ambitious preprint, with additional attention because generative systems helped produce the argument.

The AI disclosure also creates a reproducibility problem. Naming models tells readers which systems participated, but not whether other researchers can reconstruct the discovery path. Hosted models change, internal versions may be unavailable, and stochastic outputs vary.

Even a complete prompt transcript would not make the proof true. It would show how the authors arrived at it and help identify which steps received meaningful human scrutiny.

Formal verification could address part of this concern, but the preprint is not presented as a machine-checked proof in a system such as Lean. Translating advanced Kähler and algebraic geometry into a proof assistant would itself require substantial foundational work.

Peer review is also not infallible. The better standard is converging verification from specialists who try to reproduce the stability analysis, challenge the analytic obstruction, and simplify the construction.

Public discussion has already shown both excitement and caution. Some mathematicians treat AI-assisted counterexamples as evidence that research automation is accelerating. Others emphasize that specialist guidance and proof checking still constitute most of the intellectual burden.

Both views can be true. AI may have generated the decisive construction, while human expertise remains necessary to recognize its relevance and establish every claim around it.

Readers should also resist drawing labor-market conclusions from one paper. A model-assisted result in a highly structured area does not show that autonomous systems can select valuable problems, sustain a research program, or train the next generation of mathematicians.

The event does show that dismissal is no longer adequate. When an established researcher publicly credits multiple models with a major conjecture-level result, the research community must evaluate the proof and the workflow seriously.

The safest current verdict is precise: a credible expert has posted a detailed AI-assisted disproof claim, but the claim has not yet accumulated independent validation proportional to its importance.

Three Signals That Will Decide How This Result Ages

The next stage is not another viral headline. It is expert replication, a revised stability map, and a durable account of the AI workflow.

The first signal is independent mathematical verification. Specialists need to confirm the fivefold’s construction, its ordinary K-polystability, and the nonexistence of the claimed metric. A concise independent proof, a seminar consensus, or a public correction would carry more weight than engagement numbers.

If experts reproduce the argument, the interpretation strengthens immediately. If they identify a repairable gap, attention shifts to whether the construction can be salvaged. If the central stability or obstruction claim fails, the disproof does not stand.

The second signal is how researchers position the example against stronger YTD formulations. A useful counterexample should reveal exactly which enhanced condition detects its instability.

That comparison will determine whether the paper closes a program or sharpens it. The most likely durable reading is that ordinary K-polystability was too weak, while uniform or completed non-Archimedean stability captures the missing direction.

Watch for follow-up papers that calculate those stronger invariants on Liu’s example. Also watch for attempts to reduce its dimension or produce a broader family of counterexamples. Either development would show that the object reveals a structural boundary rather than an isolated anomaly.

The third signal is disclosure quality. The appendix creates a precedent by naming the models and the orchestration system. Future revisions can make that record more useful by separating machine proposals, human interventions, discarded arguments, and final verification.

That evidence will shape how journals and research institutions evaluate AI-assisted mathematics. Transparent workflows can support credit assignment and error analysis. Vague claims that “AI solved it” cannot.

For developers and knowledge workers, the immediate lesson is to treat provenance as part of the output. Preserve source material, model context, decisions, and corrections through a disciplined research workflow. The same practice applies when AI analyzes code, contracts, experiments, or market data.

The YTD story is therefore neither a clean machine triumph nor a reason to ignore AI-generated research. It is a live test of a hybrid process: models search and synthesize, experts choose and verify, and the wider community tries to break the result.

The question to ask over the coming months is concrete. Can independent geometers validate the fivefold, identify the stronger stability condition it violates, and reconstruct enough of the AI process to trust how the result was produced?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page