top of page

Shanmu Jin Used AI on the Crouzeix Conjecture, but the Proof Still Needs Human Judgment

Shanmu Jin, a neurosurgery resident and postdoctoral researcher in Beijing, claims a proof of the Crouzeix conjecture after a 16-hour autonomous AI run. The result targets a problem that has resisted specialists since Michel Crouzeix formulated it in 2004.

The headline is remarkable, but it compresses several different claims. An AI system reportedly helped produce a key theorem. Jin then assembled and published the mathematical argument under his own name. Experts have examined the manuscript, but it remains a non-peer-reviewed preprint.

There is another complication. Emiel Lorist and Felix Schwenninger posted an independent solution days later, using a different proof strategy. That second manuscript turns the story from a single viral AI achievement into a test of how mathematical credit and verification should work.

What Shanmu Jin Actually Published

The confirmed event is the publication of a complete claimed proof, not the formal acceptance of a solved conjecture.

Jin submitted a manuscript titled “The Numerical Range Is a 2-Spectral Set” on July 24, 2026. The hosting platform posted the preprint on July 27.

The paper states that it proves the optimal constant in the Crouzeix conjecture. Its central result concerns the numerical range of a complex matrix, a compact region containing more operational information than the matrix’s eigenvalues alone.

For a matrix A, its numerical range contains values produced by applying A to every unit vector and measuring the resulting quadratic form. This set provides a geometric way to study matrix behavior.

The conjecture asks whether the following inequality always holds:

\[

\lVert p(A)\rVert \leq 2\max_{z\in W(A)}\lvert p(z)\rvert

\]

Here, \(p\) is a polynomial, \(W(A)\) is the numerical range, and the left side measures the size of the transformed matrix.

In practical terms, the statement connects a difficult matrix computation with the maximum value of a scalar function over a geometric region. The number 2 is the proposed universal constant.

Jin’s preprint says the constant is both valid and optimal. It introduces what the paper calls a positive-real completion theorem, involving a constraint on a normalized matrix-valued Carathéodory function.

That terminology requires care. A Carathéodory function is an analytic function whose real part remains positive. Jin’s argument uses such functions to control a matrix norm through an operator identity.

According to a SIAM News manuscript by numerical analysts Alex Townsend and Anne Greenbaum, Jin described a substantial role for OpenAI’s GPT-5.6 Sol. They wrote that a key theorem emerged during an autonomous run lasting about 16 hours.

That does not mean a chatbot answered the conjecture in one conversational response. The reported process involved a long-running research system exploring arguments, checking paths, and developing a candidate result.

Jin still had to recognize the result’s relevance, convert it into a coherent proof, and take responsibility for the manuscript. Those steps separate a mathematical contribution from an unfiltered model transcript.

The public evidence also supports Jin’s institutional affiliation. His manuscript lists Peking Union Medical College Hospital, and previous medical research identifies him with its neurosurgery and brain-tumor programs.

The available sources do not establish that the hospital directed this mathematical work. They also do not show that the institution reviewed or endorsed the proof.

Calling it “a hospital’s AI solving mathematics” would therefore be misleading. The documented subject is one physician-researcher, using a general AI system during work that crossed medicine and operator theory.

Jin reportedly reached the topic while investigating mathematical tools connected with transcranial ultrasound research. That origin explains why someone outside a traditional mathematics department encountered the conjecture.

It does not establish that the final theorem has a direct clinical application. The mathematical and medical threads should remain separate until Jin publishes a more detailed account of that connection.

Why the Crouzeix Conjecture Matters

The Crouzeix conjecture matters because it seeks a universal safety margin for evaluating functions of matrices, not because it is famous outside mathematics.

Functions of matrices appear throughout scientific computing. They help express solutions to differential equations, model changing systems, estimate stability, and analyze iterative algorithms.

Computing such functions can be difficult when a matrix is non-normal. A non-normal matrix does not commute with its conjugate transpose, so its behavior can be far less predictable than its eigenvalues suggest.

Two matrices may have similar eigenvalues while reacting very differently to perturbations. Their powers and polynomial transforms can also grow in unexpected ways.

The numerical range offers a broader geometric picture. Crouzeix’s proposal says that this picture controls every polynomial function of the matrix within a factor of 2.

Michel Crouzeix established an initial universal bound of 11.08. Later work by Crouzeix and César Palencia reduced the known general constant to \(1+\sqrt{2}\), approximately 2.414.

Their 2017 result came strikingly close to 2. Yet closing that final gap required more than numerical evidence or proofs for special matrix families.

Researchers subsequently verified the conjecture for restricted classes. Those results included certain small matrices, nilpotent matrices, and matrices with specially structured numerical ranges.

For example, a 2024 study proved versions of the statement for several classes of matrices associated with finite shift compressions. Its published analysis also documented how difficult extensions become as matrix dimension and structure change.

A universal proof must cover every square complex matrix. It must also address every eligible polynomial, or more generally suitable analytic functions, without relying on a favorable special form.

That breadth is why a short-looking inequality survived for more than two decades. Its notation hides interactions among complex analysis, operator theory, and numerical linear algebra.

The conjecture also serves as a bridge between pure and applied mathematics. Operator theorists care about spectral sets and functional calculus. Numerical analysts care about stable bounds for matrix algorithms.

A proof would not instantly replace existing numerical software. It would instead settle the best universal constant behind a large family of estimates.

That distinction matters when evaluating the AI story. The value lies in a reusable theorem with a correct proof, not in a benchmark score or a dramatic chat transcript.

The result also pressures AI companies to define “solving” more carefully. Finding a plausible lemma is different from producing a complete argument. Producing an argument is different from surviving expert review.

Mathematics offers unusually strong verification standards, but natural-language proofs are not automatically machine-checkable. Specialists must still examine definitions, quantifiers, limiting arguments, and hidden regularity assumptions.

That makes the Crouzeix case more informative than an AI system solving a contest question. The answer was not known in advance, and ordinary answer matching cannot verify it.

The Real Contest Is AI Discovery Versus Human Verification

The main tension is not physicians versus mathematicians, but rapid AI-assisted discovery versus the slower process that turns a manuscript into accepted mathematics.

Townsend and Greenbaum reportedly approached Jin’s manuscript with skepticism. Both have relevant backgrounds in numerical linear algebra, and Greenbaum has published directly on the conjecture.

Their account says they investigated the proof and discussed it with other experts. It also reports that Michel Crouzeix examined the manuscript and believed it was correct.

That is significant evidence. It is not equivalent to peer review, journal acceptance, or a machine-checked formalization.

Expert confidence can change when a proof circulates more widely. Small gaps may require repairs, while larger issues may invalidate a central step.

Preprint history contains every possible outcome. Some surprising arguments survive almost unchanged. Others need revisions, and some fail after specialists identify an overlooked assumption.

The AI component makes provenance another verification target. Readers need to know which theorem came from the autonomous run, what prompts and tools were used, and how Jin checked the output.

Without those records, outsiders can evaluate the paper but cannot independently study the claimed research process. The mathematical conclusion and the AI attribution therefore have different evidence levels.

The conclusion appears in a public manuscript. The 16-hour account comes from Jin’s description to other researchers. No complete execution trace has been publicly linked with the preprint.

That gap does not discredit the proof. It limits claims about autonomy, reproducibility, and the exact division of labor.

“AI-assisted” is currently the safest description. It recognizes the reported role of the model without pretending that a software agent independently selected the problem, certified the result, and published it.

The case also challenges traditional assumptions about expertise. Jin is not presented as a career operator theorist, yet he appears to have navigated highly specialized literature and produced a manuscript experts took seriously.

AI systems can reduce the cost of entering an unfamiliar technical field. They can retrieve definitions, compare arguments, generate examples, and sustain exploratory searches longer than a busy researcher might manage manually.

However, reduced entry costs can increase the verification burden. A field may receive more plausible manuscripts from authors who lack established relationships with its specialists.

Editors and reviewers then face a larger stream of work with uncertain provenance. Good results can arrive through unconventional routes, while polished errors can consume extensive expert time.

This is where mathematical judgment becomes more valuable, not less. A model can accelerate conjecture formation and proof search, but communities still decide which arguments deserve trust.

Developers building AI research systems should notice the operational lesson. The useful product is not merely an answer generator. It is a traceable environment for hypotheses, citations, failed attempts, and validation.

Knowledge workers face a similar problem at smaller scale. An AI-generated conclusion becomes easier to audit when its source material and intermediate reasoning remain searchable.

A searchable knowledge base cannot validate a theorem. It can preserve the evidence trail that reviewers need when evaluating AI-assisted work.

A Second Proof Changes the Crouzeix Story

An independent proof makes the conjecture look more likely to be settled, while making exclusive credit for the solution more complicated.

Lorist and Schwenninger submitted “A solution to Crouzeix’s conjecture” to arXiv on August 4, 2026. That was eight days after Jin’s preprint became public.

Their arXiv manuscript presents a different route. It combines earlier tools for weaker estimates with a perturbation lemma for 2-dilations.

A dilation represents an operator inside a larger space where its structure becomes easier to control. The authors apply their lemma to iterates in a double-layer potential representation.

That mechanism differs from Jin’s positive-real completion theorem. The two manuscripts therefore appear to offer independent approaches, rather than minor variations of the same AI-generated idea.

Independent proofs can strengthen confidence because an unnoticed flaw is less likely to infect two genuinely different arguments. They can also reveal which concepts are essential and which belong only to one technique.

Yet timing matters. Jin’s manuscript was publicly posted first. Lorist and Schwenninger submitted their paper later, although their research may have begun before Jin’s release.

Publication priority does not answer every question about discovery. Mathematicians will examine submission records, acknowledgments, revisions, and any communication among the authors.

The second proof also weakens the simplest viral narrative. The Crouzeix conjecture is not currently represented by one AI-assisted manuscript standing alone against decades of failure.

Instead, the field now has two claimed general solutions posted within a short period. One reportedly emerged through human-AI collaboration, while the other uses a conventional authorial presentation.

This convergence can support several interpretations. The field’s existing methods may have reached a point where the final gap was ready to close. New AI systems may have accelerated one route. Coincidence may also have compressed the timeline.

No public evidence establishes that the second team used Jin’s manuscript or the same AI technique. It would be irresponsible to frame the papers as a direct race without more documentation.

The two-proof situation also creates a useful natural experiment. Reviewers can compare length, dependencies, generality, and repairability.

A shorter proof is not automatically better. A longer proof is not automatically safer. The key question is whether every step follows from stated assumptions.

If both survive, the mathematical community may credit multiple solutions while still recognizing Jin’s earlier public posting. It may also treat the AI-assisted derivation as historically important, regardless of which proof becomes standard.

If only one survives, the narrative changes substantially. A flaw in Jin’s proof would not disprove the conjecture if the independent manuscript remains valid.

Conversely, a flaw in the later paper would leave Jin’s argument carrying more weight. The existence of two submissions reduces dependence on either single manuscript.

This is the central reversal in the story. AI may have helped produce an answer, but acceptance no longer depends on believing an AI success claim.

The manuscripts themselves must compete under ordinary mathematical standards. Their provenance attracts attention, while their arguments determine the outcome.

What the Claims Still Do Not Establish

The strongest responsible conclusion is that Jin produced a serious AI-assisted proof claim that experts find promising, not that AI has conclusively solved the problem.

Neither public manuscript had completed conventional peer review by August 15, 2026. Both should therefore be described as claimed solutions.

Peer review is imperfect, but it creates structured scrutiny. Referees verify novelty, inspect dependencies, test central lemmas, and request clarification where an argument moves too quickly.

For a result of this scope, informal review will also matter. Specialists are likely to present the arguments in seminars, reconstruct the proofs, and test them against known difficult cases.

Jin’s central theorem deserves particular attention. A completion result can fail if a positivity condition is weaker than the proof assumes or if an operator identity does not extend to a boundary case.

That observation does not identify a flaw. It describes the type of stress testing any new universal argument requires.

Researchers must also check the passage from normalized analytic functions back to arbitrary matrices and polynomials. Universal statements often hide delicate approximation or continuity steps.

The model’s role creates another uncertainty. Autonomous systems can generate internally consistent chains containing a subtle false lemma. Repeated self-checking does not guarantee independence because the same model may reproduce the same mistake.

Human review reduces that risk only when reviewers reconstruct the logic. Reading a polished proof and finding it plausible is not enough.

Formal verification would provide a stronger signal. Translating the proof into Lean, Isabelle, or another proof assistant would force every dependency and inference into an explicit form.

That process may take months. Operator theory uses analytic machinery that can be expensive to formalize, particularly when library coverage is incomplete.

A failed formalization would not automatically invalidate the proof. It might expose missing library components rather than a mathematical error.

A successful formalization would be unusually valuable. It would separate confidence in the theorem from confidence in the language model that helped generate it.

The story also offers no evidence that AI can routinely solve comparable open problems. One success, if confirmed, does not establish a general rate of autonomous mathematical discovery.

Selection effects are substantial. People publicize successful runs, while failed searches and abandoned proof attempts receive less attention.

The reported 16-hour run also says little about total effort. Jin may have spent far longer selecting the problem, preparing context, evaluating outputs, and repairing the manuscript.

Readers should resist comparing 16 compute hours with 22 years of human work. The AI system relied on mathematical literature and methods created during those years.

Likewise, Jin’s medical background should not become a novelty device that minimizes his contribution. Physicians can possess substantial research training, and postdoctoral work often crosses disciplinary boundaries.

The accurate surprise is narrower. A researcher outside the conjecture’s established specialist community used AI to enter the problem and produce a proof that leading experts considered credible.

That is already important. Inflating it into fully autonomous machine discovery makes the documented achievement harder to understand.

Three Signals Will Decide Whether This Result Holds

The next phase belongs to verification, comparison, and disclosure rather than another round of viral claims.

The first signal is a detailed expert response to Jin’s core theorem. A journal report, public technical note, or specialist seminar identifying no fatal gap would strengthen the case considerably.

The most useful review would explain why the positive-real completion step works. General praise from a prominent mathematician carries less information than a reconstruction of the argument.

A published correction would not necessarily defeat the result. Many valid proofs require local repairs. The decisive issue is whether the central mechanism survives.

The second signal is how the two proofs compare under scrutiny. If independent reviewers validate both Jin’s method and the Lorist-Schwenninger perturbation argument, confidence in the conjecture’s resolution will rise sharply.

Agreement between distinct mechanisms would also reduce the chance that both papers share one hidden dependency. Researchers could then choose the cleaner route for future applications.

If one proof fails, attention will shift to whether the other avoids the same obstacle. That outcome would weaken broad claims about immediate settlement without returning the field to its earlier position.

The third signal is whether Jin or OpenAI releases a reproducible account of the AI-assisted process. Useful disclosure would include the initial prompt, model settings, tool access, intermediate artifacts, and human edits.

A complete transcript may contain irrelevant or sensitive material. Researchers could still publish a curated technical record that preserves the decisive steps.

Such a record would let AI researchers distinguish retrieval from original derivation. It would also show whether the system proposed the key theorem, supplied its proof, detected errors, or mainly organized Jin’s ideas.

That distinction matters beyond one conjecture. Research institutions need standards for assigning credit, disclosing model assistance, and retaining evidence when autonomous systems contribute to publications.

The Crouzeix conjecture now offers a unusually concrete test. There is a public theorem, a named human author, a reported autonomous run, and an independent competing proof.

Developers should watch whether the result becomes peer-reviewed or formally verified. Researchers should examine whether the AI process can be reproduced on related problems without access to the solution.

Everyone else should treat the episode as evidence of changing research practice, not as proof that mathematical verification has become obsolete.

The most consequential outcome would not be an AI system replacing mathematicians. It would be a workflow where more people can explore specialized problems while experts and formal tools preserve the standard of proof.

Will the Crouzeix argument survive line-by-line review, and will its AI research trail become inspectable? Those answers will determine whether this is a durable scientific milestone or an instructive preprint controversy.

Get started for free

A local first AI Assistant w/ Personal Knowledge Management

For better AI experience,

remio only supports Windows 10+ (x64) and M-Chip Macs currently.

​Add Search Bar in Your Brain

Just Ask remio

Remember Everything

Organize Nothing

bottom of page