OpenAI Math Controversy: 25 Fields Medalists Challenge the Race for AI Proofs
OpenAI turned an extraordinary mathematical claim into an OpenAI math controversy within days. On September 8, the company said an internal AI system had solved the Navier-Stokes Millennium Prize Problem. Three days later, 25 Fields Medalists warned that AI companies and mathematicians were pursuing dangerously different goals.
The dispute is larger than whether OpenAI's proof survives review. It concerns how AI laboratories choose problems, handle confidential research, assign credit, and announce results before specialists can absorb them. The mathematicians are challenging a system that rewards the fastest answer, even when understanding and attribution arrive later.
OpenAI says its researchers and agents never saw the private work of mathematicians Tristan Buckmaster and Levent Alpöge. Buckmaster says he does not know whether their data influenced the system. That unresolved gap now sits beside a potentially historic proof, making trust part of the result itself.
How the OpenAI Math Controversy Escalated in Four Days
The decisive change was not one proof, but the speed at which a research claim became a dispute about scientific conduct.
OpenAI launched its effort on September 1 after hearing rumors that two Millennium Prize Problems had been resolved. The company later connected those rumors to Buckmaster, an NYU professor, and Alpöge, an Anthropic employee working independently.
According to OpenAI's technical account, it initially assigned agents to every unresolved Millennium Prize Problem. It then concentrated resources on Navier-Stokes after learning more about the related work.
Navier-Stokes equations describe fluid motion. The Millennium Problem asks whether smooth three-dimensional solutions always remain well behaved, or whether singularities can form in finite time.
OpenAI says its agents found a resolution on September 5, about 88 hours after the first agents started. Formalization and verification in Lean took another 17 hours.
Lean is a proof assistant that checks whether formal logical steps follow from stated definitions and axioms. Passing that check can establish internal consistency within the formal system. It does not replace scrutiny of how a theorem represents the original problem.
The company reported that the Navier-Stokes project generated approximately 2.7 million agent messages and 130 billion output tokens. Across all attempted problems, its agents exchanged 4.9 million messages and produced roughly 300 billion tokens.
Those figures show the scale of the search. They do not establish mathematical acceptance, originality, or proper attribution. Each of those questions requires a different kind of review.
OpenAI published its claim on September 8. Nature described it that day as the first claimed solution by a computer to a truly major open problem, while carefully attributing that conclusion to OpenAI.
The initial assessment captured the extraordinary stakes. A valid proof would represent a major advance in both mathematics and automated reasoning. Yet the announcement arrived before the normal process of publication, expert discussion, correction, and broad acceptance had run its course.
A separate conflict was already developing. Buckmaster and Alpöge had spent about a year studying related fluid equations with assistance from several AI systems. Their work addressed forced Euler and other connected problems, not the exact theorem OpenAI claimed to resolve.
Their path still mattered. Buckmaster said their approach emerged from earlier work by Diego Córdoba and Luis Martínez-Zoroa. He argued that it was not an obvious direction an AI would simply discover from the original problem statement within several days.
OpenAI offered a different account. It said its agents did not see Buckmaster and Alpöge's work before public release. The company also said its proof and the outside researchers' results differed significantly.
On September 10, OpenAI updated its public account after investigating whether user inputs could have influenced its system. It said Buckmaster's Codex prompts from the preceding two months could not have affected the model, including through training.
The company also recognized Alpöge and Buckmaster's priority on their forced Euler result. It maintained that no specific user data was accessed to solve Navier-Stokes.
On September 11, the conflict widened beyond the individuals involved. Twenty-five Fields Medal recipients published a signed declaration describing a severe misalignment between AI companies and the mathematical community.
Their intervention changed the story. The central question was no longer only whether one model had produced one correct proof. It became whether AI benchmark culture was damaging the institutions needed to turn answers into durable knowledge.
Why 25 Fields Medalists Rejected the Benchmark Mindset
The signatories did not reject AI-assisted mathematics. They rejected treating famous problems as scoreboards detached from understanding, mentorship, and intellectual credit.
The group included Terence Tao, Maryna Viazovska, Manjul Bhargava, Peter Scholze, June Huh, Cédric Villani, and 19 other Fields Medal recipients. Their collective standing made the statement difficult to dismiss as ordinary resistance to automation.
Their argument starts with a distinction between solving and understanding. A famous problem gives mathematicians a direction for exploration. Its value comes partly from the ideas, methods, and questions developed during that journey.
A final true-or-false answer is only one product of that process. Researchers must explain the proof, locate its central ideas, connect it with previous work, and make those ideas usable by other people.
The declaration warned that mass-producing solutions at increasing speed can disrupt this process. A torrent of technically valid claims can overwhelm the limited number of specialists capable of evaluating, explaining, and integrating them.
Verification therefore has at least three layers. A formal checker can test whether encoded steps are valid. Expert mathematicians must determine whether the formal statement matches the intended theorem. The wider community must then decide whether the result introduces reliable and useful knowledge.
The Fields Medalists focused heavily on that third layer. They argued that mathematical ideas become valuable through talks, discussions, simplifications, teaching, and careful writing. Those activities require time and willing human participation.
This makes the OpenAI math controversy a labor-allocation problem as well as a scientific dispute. AI systems can increase the supply of proposed results faster than universities can increase the supply of qualified reviewers.
The burden does not fall evenly. Specialists in a narrow field may feel obligated to inspect high-profile claims because public confidence depends on their judgment. Yet reviewing a long machine-generated proof can displace their own research, teaching, and student supervision.
AI laboratories can fund enormous inference runs. Academic departments usually cannot summon an equally large review workforce. That imbalance lets a company control the announcement tempo while outsiders inherit much of the verification cost.
The signatories also challenged the assumption that famous problems make ideal AI benchmarks. Benchmarks require recognizable targets and clear success signals. Millennium Problems appear attractive because almost everyone understands that solving one would matter.
Mathematical research has broader goals. Researchers seek explanatory structures, reusable techniques, productive conjectures, and new communities of inquiry. A proof that closes a problem without revealing accessible ideas can score well while contributing less to those goals.
This does not make the proof worthless. It changes the standard for calling the project successful.
OpenAI itself acknowledged part of that distinction. The company described its result as substantial progress and said it did not intend to claim the Millennium Prize. It also called the work a snapshot of AI development, rather than a final culmination.
That framing places the proof inside an AI capability narrative. For OpenAI, the run tested whether a coordinated system could attack a difficult frontier problem. For mathematicians, the relevant test includes what the community can understand and build afterward.
Both evaluations can be legitimate. Trouble begins when the laboratory's benchmark result is presented publicly as equivalent to a completed scientific process.
The declaration therefore asks humans controlling AI development to make different choices. It supports AI that enhances mathematical study and understanding. It opposes incentives that prioritize rapid announcements over integration into the mathematical record.
The distinction matters outside mathematics. Software teams already face similar pressure when generated code increases output faster than people can review architecture, security, and maintenance consequences.
Knowledge workers experience another version when AI produces summaries faster than anyone can validate their sources. A searchable AI knowledge base can preserve context, but it cannot decide which unsupported claim deserves trust.
The mathematicians' warning is therefore not nostalgia for slower work. It is a demand that automation include the human systems needed to interpret and maintain what machines produce.
The Main Conflict Is Speed Versus Scientific Stewardship
OpenAI optimized for finding a result quickly, while the mathematicians insist that discovery carries obligations that cannot be compressed into an inference run.
The company's system used coordinated agents to explore several approaches in parallel. Codex consolidated promising intermediate results, allowing ideas from different groups to influence later searches.
This resembles a large research organization operating at machine speed. Teams pursue separate paths, compare findings, and redirect resources toward promising leads. The difference lies in volume and tempo.
Thousands of concurrent agents can generate more mathematical material than any human group could read during the same period. That scale helps search a difficult problem space, but it also creates a filtering problem.
A correct proof might be buried among failed attempts. A useful idea might appear without a clear intellectual lineage. Several agents might independently reconstruct techniques present in training data or public literature.
The final output can therefore be formally valid while its conceptual history remains hard to explain. Formal verification answers whether the encoded proof works. It does not automatically reveal which prior ideas made the search successful.
That limitation became sensitive because OpenAI began after hearing rumors about external progress. The company says the rumor inspired its evaluation effort and that it only saw the outside work after publication.
Buckmaster's account raised different concerns. He said he and Alpöge had placed drafts and research material into Codex during their project. During conversations with OpenAI, he asked whether those sessions were available to the internal system.
He said he received an assurance that the model did not look up user data, but no immediate answer about training. OpenAI later said the relevant prompts could not have influenced the model through any route, including training.
Those statements narrow the dispute, but they do not erase the broader governance question. Researchers need to know whether confidential queries can influence model development, evaluations, or internal research agendas.
The issue extends beyond direct copying. Metadata, aggregate usage patterns, de-identified examples, and employee knowledge can all affect what a laboratory chooses to investigate. Different pathways require different safeguards.
OpenAI's updated statement addresses the specific model and period involved. It does not yet establish a universal protocol for conflicts between product users and the laboratory's internal researchers.
The dispute also includes contested conversations about publication and authorship. Buckmaster said OpenAI proposed coordinated release options that would have treated the parties differently because Alpöge worked for Anthropic.
OpenAI researcher Sébastien Bubeck rejected the allegation that he sought to remove Alpöge from authorship of Alpöge's own work. He said the discussion concerned authorship of a rewritten OpenAI proof.
Bubeck also apologized for a remark about Buckmaster risking his career. He described it as a poor choice of words and said he had intended to warn against making unfounded accusations.
Buckmaster explicitly stopped short of accusing OpenAI of data misuse. He said he did not know what its model did, how it worked, or whether the outside team's data was used.
The dispute timeline therefore contains claims, denials, clarifications, and an unresolved trust problem. It does not support a simple conclusion that the proof was stolen.
That distinction is essential. The strongest verified criticism concerns process, incentives, and transparency. A plagiarism claim would require evidence that has not been publicly established.
The 25 Fields Medalists widened their argument accordingly. Their declaration discusses hurried announcements, incomplete attribution, and the loss of mathematical understanding. It does not depend on proving that OpenAI accessed private drafts.
This is why speed versus scientific stewardship is the clearest conflict. OpenAI can be correct about data access and still face legitimate criticism about its announcement model.
Likewise, the mathematicians can oppose benchmark incentives without denying that the AI system produced important work. The dispute is over what institutions should reward once machines can generate credible frontier research at extreme scale.
A Lean-Checked Proof Is Not Yet a Millennium Prize Resolution
The proof's formal status is significant, but public descriptions have moved faster than the institutions that determine mathematical acceptance.
OpenAI says its proposed resolution shows that smooth solutions to three-dimensional Navier-Stokes equations can develop a finite-time singularity. If accepted, that would settle the problem by establishing the breakdown side of the stated alternatives.
The company also released a Lean formalization. That step offers more than a conventional manuscript alone because specialists can inspect machine-checkable definitions and dependencies.
Still, formalization does not remove every source of error. The formal theorem might differ subtly from the Clay problem. Assumptions can be encoded incorrectly. Definitions can capture a nearby result without resolving the intended question.
Specialists must compare the formal statement with the original problem and interpret each construction. They also need to understand whether the proof relies on acceptable formulations, regularity conditions, and solution concepts.
The public should therefore distinguish four claims.
First, OpenAI says an internal system generated a proposed solution. That event is documented by the company's release and published materials.
Second, Lean accepted a formalized proof. This supports logical consistency within the chosen formal framework.
Third, outside experts must determine whether the formal result exactly resolves the Navier-Stokes problem. That process was still underway when the controversy erupted.
Fourth, the Clay Mathematics Institute has a separate recognition process. Its prize rules require publication in a qualifying outlet, a waiting period of at least two years, and general acceptance by the global mathematics community.
No launch announcement can bypass those conditions. OpenAI also says it does not intend to seek the prize, but the rules remain useful as a model of deliberate verification.
The waiting period reflects a practical reality. Major proofs often change as readers find gaps, simplify arguments, or question definitions. Community acceptance cannot be scheduled around a product announcement.
The only previously resolved Millennium Problem also shows how different formal discovery can be from institutional closure. Grigori Perelman posted work on the Poincaré conjecture in the early 2000s. Years of scrutiny followed before the result became broadly accepted.
AI can shorten the search for a proof. It cannot automatically compress the time required for people to understand its consequences.
That gap creates a communications risk. Headlines saying that OpenAI "solved" Navier-Stokes can sound definitive. More accurate language says the company proposed and formalized a solution that awaits independent evaluation.
The skeptical case should not overreach in the opposite direction. A pending review does not mean the proof is probably wrong. Formal verification and early expert interest make the claim more serious than an unsupported chatbot answer.
The real uncertainty concerns equivalence, originality, explanation, and acceptance. Those issues matter even if no logical defect appears.
OpenAI's scale claims also need careful interpretation. The large token count shows an expensive and extensive computation. It does not measure the novelty or elegance of the resulting mathematics.
A short human proof can reveal a reusable idea. A long machine-generated proof can settle a theorem through intricate constructions that few people can extend. Either result can be correct, but their scientific value may unfold differently.
The distinction also affects safety. Researchers often discover hidden assumptions while explaining a result to colleagues. If AI systems produce proofs faster than humans can reconstruct them, that social error-correction layer weakens.
A useful standard would require more than downloadable proof files. Laboratories could publish readable conceptual maps, provenance records, failed approaches, model details, and formal definitions alongside major claims.
Independent teams could then reproduce both the formal check and the correspondence with the original theorem. Such practices would make AI-assisted discovery easier to trust without requiring companies to reveal every model parameter.
Until those norms exist, the phrase "AI solved a Millennium Problem" compresses several unfinished judgments into one sentence. That compression helped turn a remarkable research result into the OpenAI math controversy.
Attribution Becomes Harder When the Research Tool Is Also a Competitor
The controversy exposes a conflict that ordinary privacy policies were not designed to manage: an AI provider can serve researchers while pursuing the same discoveries internally.
Buckmaster and Alpöge reportedly used several language models during their work, including OpenAI systems. Their drafts and prompts therefore passed through tools operated by companies with their own scientific ambitions.
Researchers routinely share unpublished reasoning with software. Cloud storage holds manuscripts, email carries conjectures, and collaborative platforms record discussions. AI assistants go further because users submit detailed questions, failed approaches, and partial insights.
That interaction can expose the shape of an unfinished discovery. Even without copying text, a provider might learn that users are concentrating on a specific method or nearing a result.
OpenAI says this did not happen in the Navier-Stokes effort. Its investigation concluded that Buckmaster's relevant Codex prompts could not have affected the internal system.
The company deserves to have that finding reported clearly. The public evidence currently does not establish that OpenAI trained on or retrieved the mathematicians' private material.
However, trust requires more than a company investigating itself after a dispute. Researchers need policies established before they disclose valuable ideas.
Those policies should answer practical questions. Can product data inform internal benchmark selection? Can de-identified prompts influence research priorities? What controls separate commercial product teams from frontier research groups?
Users also need to know how conflicts are reviewed. An independent audit offers stronger assurance than a retrospective statement from the organization facing the allegation.
The competitor dimension adds another complication. Alpöge worked for Anthropic, although reports describe this project as independent work. OpenAI therefore encountered research connected to both an external academic and an employee of a rival laboratory.
That affiliation made coordination sensitive. It also illustrates why conventional authorship norms can collide with corporate boundaries.
Scientific credit usually follows intellectual contribution. Corporate publication rules can instead follow employment, confidentiality, and competitive strategy. An AI-generated proof produced inside a laboratory sits at the intersection of both systems.
The Fields Medalists' attribution warning addresses this collision. Mathematics develops through explicit links to previous ideas. A proof must be situated within that lineage, even when a machine assembled its steps.
Models complicate attribution because they absorb broad bodies of literature during training. Their outputs rarely identify which sources shaped a particular inference. Agent systems add another layer by combining intermediate results across many generated conversations.
A laboratory may therefore struggle to reconstruct provenance after success. That is not proof of misconduct. It is a design weakness for systems marketed as scientific collaborators.
Better provenance could include timestamped agent traces, cited retrieval sources, model checkpoints, and records of human interventions. Privacy-preserving audits could test whether specific user material entered the research pipeline.
Clear embargo procedures would help too. When a laboratory learns that outside researchers are approaching a result, it could appoint an independent conflict reviewer before launching a competing internal run.
Publication protocols could separate capability demonstrations from scientific claims. A laboratory might disclose that a model generated a candidate result while delaying definitive language until outside reviewers complete an initial assessment.
These measures would slow the marketing cycle. They would also protect the credibility of future results.
The OpenAI math controversy shows what happens when such rules are improvised during a race. Even a correct proof can acquire an avoidable cloud of suspicion.
For enterprise buyers, the lesson is immediate. Sensitive prompts can contain patentable methods, unreleased product plans, legal reasoning, or proprietary research. Contract terms and technical controls must address internal research use, not only external data disclosure.
For developers, provenance should become a product requirement. An answer that cannot explain its source path creates risk when it enters code, research, or regulated decisions.
For knowledge workers, retaining original documents and conversations remains essential. Generated summaries should supplement the evidence trail, not replace it.
Mathematics provides an unusually clear test because formal correctness can sometimes be checked. Other fields will face the same ownership problem with weaker verification tools and greater room for ambiguity.
Three Signals Will Decide What This Episode Changes
The next phase depends on mathematical review, enforceable data safeguards, and whether AI laboratories change how they announce scientific results.
The first signal is independent technical validation. Specialists must compare OpenAI's formal statement with the Clay formulation, inspect the construction, and test whether its assumptions close every relevant gap.
A clear consensus that the proof resolves the original problem would strengthen OpenAI's capability claim. It would not settle the attribution debate, but it would establish the mathematical importance of the result.
A significant mismatch or unrepairable gap would weaken the company's announcement. It would also reinforce the Fields Medalists' warning about promoting results before adequate review.
Readers should watch for peer-reviewed publication, detailed seminars, independent Lean reproductions, and specialist explanations. General reactions from AI commentators cannot substitute for those signals.
The second signal is a verifiable policy for researcher data. OpenAI's case-specific investigation says Buckmaster's prompts did not influence the system. The broader question is whether frontier laboratories will create auditable protections for future conflicts.
A strong policy would define barriers between user data, model development, benchmark selection, and internal scientific projects. It would also identify who independently investigates disputes.
If OpenAI or its competitors publish enforceable safeguards, researchers will have a clearer basis for trusting AI tools with unfinished work. If policies remain vague, users may avoid sharing their most valuable reasoning.
That withdrawal would harm both sides. Mathematicians would lose useful assistance, while model developers would lose expert feedback from demanding research settings.
The third signal is a change in scientific announcement practices. The 25 Fields Medalists asked AI companies to prioritize understanding, attribution, and community integration over benchmark victories.
A meaningful response would appear in the next major claim. Does the laboratory involve outside experts before announcing success? Does it publish conceptual explanations and provenance alongside formal artifacts?
Does it distinguish candidate proofs from accepted solutions in headlines and executive statements? Does it give prior researchers enough time to examine citations and intellectual dependencies?
If those practices improve, this episode may become a productive correction. AI systems could accelerate discovery while stronger institutions preserve credit and understanding.
If the same pattern repeats, the conflict will deepen. Researchers may treat frontier laboratories as competitors first and collaborators second.
That outcome would limit access to the human expertise AI systems still need. It could also fragment scientific work across corporate boundaries, private tools, and guarded academic groups.
The OpenAI math controversy is therefore not a verdict on whether AI belongs in mathematics. AI is already part of the field, from informal exploration to formal proof checking.
The unsettled issue is governance. A machine can search millions of paths, but institutions still decide what deserves belief, who receives credit, and how knowledge reaches the next generation.
Watch the proof reviews first, because correctness remains fundamental. Then watch the data rules and the next announcement, because those choices will reveal whether the industry understood the warning.
The most useful question is not whether AI can produce another famous answer. It is whether the people deploying it can preserve the conditions that make an answer scientifically valuable.



