top of page

Fields Medalists AI Warning Exposes a Fight Over What Mathematics Is For

Sep 13
12 min read

Twenty-five Fields Medal winners issued a Fields Medalists AI warning after a series of machine-generated results intensified conflict between AI companies and research mathematicians.

The September 11 declaration accepts that large language models can now contribute to important mathematical problems. Its challenge runs deeper than whether those answers are correct. The signatories argue that AI companies are optimizing for visible solutions, benchmark wins, and fast announcements. Mathematics, they say, depends on conceptual understanding, careful attribution, student development, and the slow integration of ideas.

That distinction places AI labs, particularly OpenAI, under pressure. OpenAI has promoted models as scientific collaborators and invested heavily in solving prominent problems. Yet recent disputes over proof announcements, credit, private research, and a planned Caltech event show how quickly technical progress can collide with academic norms.

The disagreement is not a simple contest between people and machines. It concerns who chooses the problems, who verifies the results, who receives credit, and who carries the cost of turning generated output into durable knowledge.

What the 25 Fields Medalists Actually Warned About

The declaration argues that solving a problem is not the same as advancing mathematics.

The signatories include Terence Tao, Peter Scholze, Maryna Viazovska, Pierre Deligne, Manjul Bhargava, June Huh, and 19 other Fields Medal recipients. Their awards span from Deligne’s 1978 medal to Yu Deng’s 2026 medal.

That range matters. The statement does not represent one generation defending its established position. It brings together mathematicians working across different fields, institutions, and stages of their careers.

Tao wrote that the declaration grew from discussions among the medalists during the preceding week. He also acknowledged that its accelerated release allowed less consultation than an earlier community initiative, the Leiden Declaration.

The central claim in the joint declaration is blunt: the goals of AI companies and those of the mathematical community are “severely misaligned.” However, the statement does not reject AI or deny its recent progress.

Instead, it distinguishes mathematical output from mathematical understanding. A proof establishes that a proposition follows from accepted assumptions. Research mathematics also asks why the result holds, which ideas make it possible, and how those ideas connect to other areas.

Famous open problems often function as landmarks. Their importance comes partly from the tools, abstractions, and new questions developed while people try to solve them. A final true-or-false answer captures only one part of that value.

The signatories warn that companies can invert this relationship. An open problem becomes a target because solving it produces an impressive demonstration. The benchmark then replaces the understanding that made the problem scientifically useful.

This concern extends to how research enters the mathematical record. A result usually moves through seminars, private discussions, written drafts, expert checking, peer review, simplification, and teaching. During that process, researchers identify the essential idea and connect it to earlier work.

Machine-generated results do not bypass that need. They can increase it.

A long proof produced by thousands of AI agents might be formally valid yet difficult for specialists to interpret. Formal verification, where software checks logical steps against explicit rules, can improve confidence in correctness. It does not automatically identify the most valuable concept or produce a clear explanation.

The declaration calls ideas and students the profession’s most precious resources. Problems given to students often serve as training environments, not merely unfinished tasks waiting for answers. Students learn how to explore, fail, reformulate questions, and communicate an argument.

A system optimized to close the problem as quickly as possible can remove that environment. The result might still be useful, but the scientific and educational tradeoff becomes harder to ignore.

The Fields Medalists AI warning therefore targets a model of production. It challenges the assumption that more solved problems, generated faster, necessarily represent healthier mathematical progress.

Why the Conflict Erupted Now

The warning arrived when AI mathematics moved from competition results toward claims about open research problems.

For several years, AI systems improved on school mathematics, Olympiad questions, theorem proving, and formal verification. Those tasks supplied clear scoring systems and allowed researchers to compare model performance.

Research mathematics creates a different test. The answer is not known in advance, the literature can be incomplete, and even experts may disagree about whether an argument is genuinely new.

AI labs nevertheless have strong reasons to pursue these problems. Mathematics offers unusually legible evidence of reasoning capability. A difficult theorem sounds more meaningful than a small gain on an abstract model benchmark.

OpenAI has described its systems as research collaborators. In a January 2026 report, the company said GPT-5.2 contributed to several open Erdős problems with tools including Aristotle and Lean. Tao validated some resulting arguments.

That scientific collaboration report also drew an important boundary. OpenAI said current models could sometimes combine known techniques or connect existing fields. It acknowledged that inventing an entirely new kind of mathematics remained beyond them.

The document presented formal checking as one answer to plausible but incorrect proofs. Lean, an interactive theorem prover, forces a mathematical argument into mechanically checkable steps. That process can catch gaps that polished natural language might conceal.

Yet correctness is only one dimension of the current dispute. Attribution, disclosure, timing, and explanation remain human governance problems.

Those tensions became concrete around OpenAI’s claimed solution involving the Navier-Stokes equations. These equations describe fluid motion and sit at the center of a famous Millennium Prize problem concerning existence and smoothness.

OpenAI said an internal model, described as more capable than its publicly available system, produced a relevant proof. The company reportedly devoted more than 1,000 agents to the effort for over 50 hours, later scaling the search to as many as 10,000 agents.

OpenAI research leader Mark Chen said the required computing cost reached millions of dollars. Those figures came from the company and do not provide an independent measure of the proof’s scientific importance.

The technical claim quickly became entangled with a credit dispute. NYU mathematician Tristan Buckmaster and Anthropic researcher Levent Alpöge had been working on closely related fluid-dynamics research.

According to the reported dispute, OpenAI learned that the researchers were nearing a result and then directed substantial resources toward the problem. Buckmaster questioned whether interactions with OpenAI products had influenced that effort.

OpenAI denied that its researchers or agents accessed the pair’s prompts or proof. However, the company said it could not completely exclude the possibility that de-identified usage data had indirectly helped improve its models.

That qualification exposed a broader trust problem. A mathematician might use an AI tool to explore unpublished work while the tool’s provider also competes to announce major mathematical results.

Even without direct access to private prompts, the perceived conflict can change behavior. Researchers may withhold ideas, avoid certain systems, or delay collaboration because they cannot confidently separate tool use from corporate research.

The argument is therefore larger than one proof or one company. When an AI provider becomes both research infrastructure and a participant in the same discovery race, ordinary assurances about privacy may not resolve concerns about incentives.

Fields Medalists AI Warning Challenges Benchmark-Driven Mathematics

The central tradeoff pits rapid problem completion against the slower creation of shared understanding.

AI companies measure progress through outputs that can be demonstrated. Mathematics values solved problems, but it also values theories, explanatory methods, reusable techniques, and the transmission of knowledge.

These goals overlap when an AI system helps a mathematician test examples, search literature, formalize an argument, or explore possible proof strategies. In those cases, automation supports an existing research process.

They diverge when the solution itself becomes a corporate trophy. The incentive then favors famous targets, speed, secrecy, and announcement priority.

A company can benefit from saying its model solved a recognized problem before competitors. The community still must determine whether the proof is correct, novel, properly attributed, and conceptually useful.

That verification work is difficult to scale. Specialists may need weeks or months to inspect a dense argument. They must compare it with existing literature, locate hidden assumptions, and decide which parts deserve preservation.

A model can generate candidate proofs faster than experts can evaluate them. That imbalance creates what the declaration describes as mass production of true-or-false statements.

The phrase does not mean every result is wrong. It describes a volume problem. Even correct outputs can overwhelm the institutions responsible for checking and interpreting them.

Academic publishing developed around relatively scarce human manuscripts. Reviewers donate time because the pace remains manageable and the resulting work supports a shared professional system.

If AI labs release many technically sophisticated claims in rapid succession, the reviewing burden shifts outward. Universities and individual researchers absorb the cost, while the lab receives attention for the initial announcement.

The same asymmetry affects attribution. Mathematical credit rarely belongs only to the person who writes the final line. It reflects conjectures, partial results, failed approaches, informal conversations, and techniques developed across generations.

A generated proof can obscure that chain. Models synthesize patterns from extensive training material and may not reliably reveal which sources shaped a particular step. A technically correct solution can therefore create unresolved questions about intellectual priority.

The medalists also object to treating famous problems as disposable benchmarks. A difficult conjecture can organize decades of work because partial progress reveals useful structure. Once an answer appears, attention and funding can move elsewhere, even if the generated proof teaches researchers very little.

That possibility creates a reversal. AI is promoted as a tool for accelerating mathematical discovery, yet benchmark-driven deployment can reduce the human understanding that gives discovery its lasting value.

There is still a strong case for AI-assisted mathematics. Models can retrieve obscure references, generate examples, translate informal arguments into formal language, and help researchers coordinate large projects.

OpenAI mathematician Sébastien Bubeck has compared possible future systems with particle accelerators. In that model, substantial computational infrastructure helps teams investigate questions no individual could manage alone.

He has also argued that people must remain central because mathematics matters when mathematicians learn from it. That position is closer to the declaration than the public conflict might suggest.

The unresolved question concerns control. If researchers choose the questions, manage attribution, and receive usable explanations, AI can expand the field. If corporate demonstrations determine the agenda, the mathematical community becomes a verification service for somebody else’s product strategy.

The Fields Medalists AI warning asks institutions to recognize that distinction before benchmark incentives become standard practice.

The Caltech Mathathon Made the Incentive Problem Visible

A planned 40-hour event turned an abstract disagreement into a dispute over funding, labor, and early-career researchers.

Caltech was scheduled to host a Mathathon beginning October 30, 2026. Participants would use large language models to attack open research problems during an accelerated competition.

Anthropic and OpenAI were initially expected to supply a combined $2 million in AI credits. Critics said that structure rewarded rapid production before the community had developed standards for evaluating the results.

An open letter opposing the event was received on September 8 and published two days later. It had 771 signatories at publication, drawn from Caltech and the broader mathematical research community.

The Mathathon letter argued that 40 hours could not support deep understanding, careful verification, and responsible communication of open-problem research. It asked organizers to suspend the event.

Its authors also highlighted an unusual distribution of resources. Participants could receive significant computing credits, but external mathematicians would likely perform much of the later checking without payment or formal recognition.

This is not simply a complaint that computing is expensive. Large computational projects already play legitimate roles across science. The concern is that access follows an event and marketing structure rather than a durable research process.

The format also placed undergraduates and other early-career participants in uncertain territory. A student might generate an interesting result without having the expertise to verify it or understand its relationship to prior work.

If the result proved wrong, the student could face reputational consequences. If it proved valuable, the platform provider and event organizers might receive more visibility than the people who completed the mathematical and editorial work.

Supporters can reasonably answer that intensive events often generate collaboration. Hackathons encourage experimentation, bring new users into technical fields, and make expensive infrastructure accessible.

The Mathathon organizers could also require disclosure, publication, expert mentorship, and post-event review. A short event does not inherently prevent serious work from developing afterward.

That counterargument identifies a weakness in the broader declaration. The Fields medalists describe urgent risks but offer few operational rules. They do not define an acceptable release process, a disclosure standard, or a system for compensating verification.

Some mathematicians commenting on Tao’s publication argued that the profession must adapt its own incentives. Journals could credit explanatory rewrites, create dedicated review groups, or require machine-readable provenance from AI-assisted submissions.

The criticism is fair. AI companies cannot independently decide what academic recognition should reward. Universities, publishers, funders, conferences, and professional societies also shape the outcome.

Still, the event showed why voluntary adaptation may not be enough. A heavily funded company can generate new outputs much faster than universities can build norms around them.

OpenAI withdrew its sponsorship following the criticism, according to public statements surrounding the event. That response reduced immediate pressure but did not settle the underlying questions.

Anthropic’s position also complicates a simple company-versus-academia story. One researcher involved in the Navier-Stokes dispute worked at Anthropic, while the company also backed the Mathathon. AI labs contain researchers whose professional values do not always align neatly with corporate competition.

The most useful dividing line is therefore not academic versus commercial. It is whether an AI project treats understanding, attribution, and community review as design requirements or as cleanup tasks after the announcement.

Correct Proofs Still Need Human Institutions

Formal verification can establish consistency, but it cannot decide whether a result has been responsibly integrated into science.

This distinction is essential because the AI mathematics debate often collapses several questions into one.

The first question is whether a proof is logically correct. Formal systems such as Lean can help answer it by checking each step within a specified framework.

The second is whether the formal statement accurately represents the original problem. A flawless proof of the wrong formalization does not resolve the intended question.

The third is whether the result is new. Establishing novelty requires literature review, expert knowledge, and careful comparison with previous work.

The fourth is whether credit has been assigned fairly. A proof can be correct and novel while still depending on improperly disclosed ideas or data.

The fifth is whether humans can learn from it. A machine-checked object may establish truth without identifying the concepts that make the result reusable.

Melanie Matchett Wood, a Harvard mathematician who helped prepare a human-written account of an AI-assisted result, described a related communication problem. She said top models often overexplain easy material while passing too quickly over the hardest steps.

Her observation captures why exposition is not cosmetic. A good proof tells expert readers where the real difficulty lies. It allows them to transfer the method to another problem.

The emerging research norms remain unsettled because existing institutions divide these functions across authors, referees, editors, seminar audiences, and teachers.

AI systems disrupt that arrangement by combining search, drafting, computation, and formalization at unprecedented scale. They can create outputs before anyone has agreed who assumes each downstream responsibility.

A constructive response should preserve the advantages of that scale while making the responsibilities explicit.

AI-assisted papers could disclose the models, tools, prompts, and computational methods that materially shaped the result. Providers could publish enough technical detail for independent replication without exposing unrelated user data.

Projects could involve domain experts before a public announcement. Those experts would help determine whether the result is novel, whether the formal statement matches the original question, and how to cite prior contributions.

Companies could also fund verification without controlling its conclusion. Independent review pools, administered through universities or professional societies, would reduce the uncompensated labor problem.

Journals may need new contribution categories. A researcher who converts a huge generated proof into an understandable argument can add substantial scientific value, even if the final theorem was already announced.

Teams also need durable records. A searchable knowledge base can preserve drafts, source material, discussions, and decisions across a long verification process. The tool matters less than the provenance discipline behind it.

None of these measures eliminates the tradeoff. Fast generation and slow understanding operate on different timescales.

Nor should the declaration be treated as proof that AI-generated mathematics is inherently destructive. Its signatories explicitly recognize AI’s potential to enhance genuine study.

The sharper claim is institutional. Without standards for attribution, verification, and explanation, technical capability will set the research agenda by default.

That outcome would not result from an autonomous machine choosing to damage mathematics. It would result from people rewarding the easiest achievements to publicize.

What to Watch After the Fields Medalists AI Warning

The next test is whether the declaration produces concrete research rules instead of another temporary controversy.

The first signal will come from AI lab publication practices. Future announcements should reveal whether companies provide complete proofs, reproducible methods, model details, and clear accounts of human contributions.

A lab that invites independent review before making a major claim would strengthen the case that AI can function as shared scientific infrastructure. Another rushed announcement followed by an attribution dispute would reinforce the medalists’ warning.

The second signal concerns the Mathathon and similar events. Organizers can redesign competitions around mentorship, disclosure, review, and post-event explanation.

That would mean judging participants on more than whether a model reaches an answer. Evaluation could include the clarity of the method, the quality of citations, reproducibility, and the result’s value to future researchers.

If events continue to prioritize the most prominent solved problem within the shortest period, the benchmark model will remain dominant. If they reward interpretation and stewardship, they can test a more collaborative path.

The third signal will be institutional adoption. Journals, conferences, funders, and mathematical societies need policies for AI-assisted research that address more than plagiarism detection.

Useful rules would define authorship, model disclosure, data provenance, formal verification, reviewer access, and credit for explanatory reconstruction. They would also clarify what evidence supports a public claim that an open problem has been solved.

The Fields Medal’s purpose includes recognizing both existing achievement and promise for future achievement. That formulation reflects a field interested in intellectual development, not only finished outputs.

AI does not make that value obsolete. It makes the distinction between an answer and a research contribution more important.

Developers should care because mathematics is becoming a test case for AI-assisted knowledge work. The same conflicts can emerge when models generate software, scientific hypotheses, legal analysis, or design concepts.

Enterprise buyers should ask whether a system preserves provenance and supports independent checking. Speed matters, but accountability determines whether generated work can be trusted after the demonstration ends.

Knowledge workers should watch who benefits when an AI system completes the visible task. The people who verify, explain, maintain, and teach the result may create much of its lasting value.

The Fields Medalists AI warning is ultimately not a demand to stop using AI. It is a demand to decide what the technology should optimize before its metrics harden into institutions.

The coming months will show whether labs, universities, and publishers can establish that shared direction. If they cannot, every new machine-generated proof will deepen the same conflict: a faster answer, followed by a slower argument over whether knowledge actually advanced.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page