top of page

OpenAI Keeps Bulldozing Mathematicians, Even When It Tries to Listen

1 hour ago
13 min read

OpenAI keeps bulldozing mathematicians despite creating a nine-member advisory group intended to repair relations with them. The company says an unreleased model has resolved more than 100 long-standing problems, yet its announcement immediately confused researchers about who controls the group.

That confusion matters because the Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, is supposed to help evaluate and communicate emerging AI results. Its members insist that the organization is independent, unpaid, and free to criticize OpenAI. However, OpenAI framed the relationship so prominently that some researchers initially assumed the company had appointed the group.

The episode extends a conflict that intensified after OpenAI announced a machine-generated solution to the Navier-Stokes existence and smoothness problem. That claim displayed the scale of the company’s mathematical ambitions. The resulting disputes over credit, publication standards, and human understanding exposed how poorly its operating style fits academic mathematics.

OpenAI now needs mathematicians to interpret the results its systems produce. Mathematicians need access, time, and reliable disclosure before they can assess those results. The new advisory group sits between those needs, but it has advice rather than authority.

OpenAI’s Math Advisory Group Arrived With an Identity Problem

AGMAI was designed as an independent bridge, but OpenAI’s rollout made it look uncomfortably close to the company it may need to challenge.

AGMAI announced its formation on September 21, 2026. The group consists of nine mathematicians and is hosted by the Institute for Advanced Study in Princeton, New Jersey.

Its initial members are François Charles, Camillo De Lellis, Timothy Gowers, Martin Hairer, Nikhil Srivastava, Ulrike Tillmann, Ravi Vakil, Edward Witten, and Melanie Matchett Wood. Their affiliations span institutions including Harvard, Stanford, Oxford, Cambridge, Berkeley, Imperial College London, and EPFL.

According to its independent group statement, AGMAI can advise OpenAI and other frontier AI laboratories. Its remit includes reviewing emerging results, coordinating communication, and developing standards for AI-assisted mathematical research.

The group says it can work with any AI laboratory. Martin Hairer, an AGMAI member and Fields Medal recipient, told The Verge that conversations with other frontier laboratories had already started. He did not identify those companies.

Hairer also emphasized that AGMAI receives no financial or technical support from OpenAI. Its members have not signed agreements restricting their public comments, apart from confidentiality conditions needed for advance access to unpublished work.

OpenAI’s announcement included similar safeguards. The company said members could publish their advice, criticize its impact, and offer guidance that OpenAI had not requested. It also said members would remain unpaid and control the group’s membership.

Those provisions establish meaningful organizational distance. They do not erase the confusing circumstances of AGMAI’s birth.

OpenAI originally approached some mathematicians about forming an advisory board. Those mathematicians instead created an independent organization and recruited additional members. Neither side has publicly identified which AGMAI members received the initial approach.

OpenAI then gave the group far more visibility than its own launch channels could provide. Many observers first encountered AGMAI through the company’s announcement, not the group’s statement or Terence Tao’s mathematics blog.

That sequencing encouraged an inaccurate but predictable impression. AGMAI looked like an OpenAI panel because OpenAI had initiated the conversations and dominated the public presentation.

Hairer described the days following the announcement as frustrating and intense. He told The Verge that the wording had not helped the group establish its independence.

The timing added to the disorder. When asked about AGMAI’s plans shortly after its creation, Hairer noted that the organization was only two days old. It did not yet have an extensive operating framework for the demands heading toward it.

The confusion is more than a branding inconvenience. AGMAI needs credibility among researchers who already distrust OpenAI’s publication practices. Every ambiguous phrase makes that credibility harder to earn.

OpenAI’s advisory announcement explicitly acknowledges concerns raised by mathematicians. Yet it also promotes the company’s claimed results before independent reviewers have explained their scope or importance.

That mixture captures the central problem. OpenAI wants outside judgment, but it also wants to announce the scale of its achievement immediately.

Why OpenAI Keeps Bulldozing Mathematicians

OpenAI treats solved problems as evidence of model capability, while mathematicians treat proofs as contributions requiring context, attribution, review, and durable understanding.

OpenAI says it began training its latest internal model on August 28. By September 21, the company claimed that the model had resolved more than 100 long-standing open problems across most areas of mathematics.

The company has not publicly released a complete catalog of those problems. Independent researchers therefore cannot evaluate the overall difficulty, novelty, correctness, or significance of the collection.

The number itself is also difficult to interpret. An “open problem” might be a famous question that has resisted specialists for decades. It might instead be a narrow unresolved statement whose solution connects existing methods without transforming its field.

OpenAI’s earlier work shows that this distinction matters. In a January report, the company said GPT-5.2 contributed to solutions for several problems associated with Paul Erdős. Terence Tao validated work on four listed problems, although the company acknowledged that Erdős problems differ greatly in difficulty.

The same report drew a useful boundary around present capabilities. OpenAI said its models can sometimes assemble known methods into a correct argument. It also said inventing an entirely new mathematical framework remains beyond current systems.

That distinction rarely survives a headline. A model resolving an obscure conjecture through a known technique and a model creating a new branch of mathematics can both become “AI solves unsolved math.”

For academic researchers, a proof must do more than reach a formally correct endpoint. It should situate the result within previous work, identify the underlying ideas, recognize relevant contributors, and help others build upon it.

OpenAI’s communications have often emphasized the endpoint instead. That approach makes sense when open problems function as model benchmarks. It clashes with a community that values how a proof expands collective understanding.

Álvaro Lozano-Robledo, a University of Connecticut mathematics professor, criticized the company’s language about having many results without knowing how to release them. He argued that academics do not normally advertise a stockpile before establishing its significance.

Researchers also object to claims of “substantial progress” on unsolved problems. In mathematics, an incomplete argument might contain valuable insights. It can also collapse because of one hidden gap.

That sensitivity is not pedantry. Mathematical proofs derive their force from each logical step. A persuasive narrative cannot substitute for a valid argument.

Formal verification offers one response. Systems such as Lean encode a proof in a language that software can check step by step. This process catches gaps that might hide inside fluent, plausible prose.

However, formal correctness does not settle every scholarly issue. A mechanically checked proof can still obscure its central insight. It can overlook related literature or present a result in a form that human researchers struggle to reuse.

The challenge therefore has two layers. The first is determining whether the proof is correct. The second is deciding whether it improves mathematical knowledge rather than adding another verified artifact.

OpenAI has accelerated the first part of producing results. It has not shown that the surrounding academic system can process its output at the same speed.

That imbalance explains why OpenAI keeps bulldozing mathematicians even when its technical claims prove substantial. The company can generate, fund, package, and publicize results faster than researchers can examine their meaning.

The Navier-Stokes Dispute Destroyed the Benefit of the Doubt

The advisory group entered a relationship already damaged by allegations involving scooping, attribution, rushed publication, and opaque revisions.

The immediate background is OpenAI’s September announcement about the Navier-Stokes existence and smoothness problem. Navier-Stokes equations describe fluid motion, while the famous open question asks whether certain three-dimensional solutions remain smooth.

The problem is one of the Millennium Prize Problems named by the Clay Mathematics Institute in 2000. Only the Poincaré conjecture had previously received an accepted solution among the original seven.

OpenAI said a system involving 10,000 agents produced a solution in less than four days. An agent in this context is a model instance assigned to pursue, evaluate, or refine part of a larger task.

The company’s claim remains consequential even apart from the surrounding dispute. A valid solution would mark a major change in AI’s ability to participate in research-level mathematics.

However, mathematicians Tristan Buckmaster and Levent Alpöge had been working on related ideas for roughly a year. Buckmaster alleged that OpenAI learned about their progress and raced to publish first.

He also raised concerns that interactions with OpenAI systems might have exposed information about their approach. OpenAI said it did not inspect Buckmaster’s inputs when pursuing its result.

The conflicting accounts have not been independently resolved. They should therefore be treated as allegations, not settled findings.

OpenAI said it began targeting Millennium Prize Problems after hearing rumors that Anthropic had solved two of them. It reportedly allocated extensive computational resources across the set before concentrating on Navier-Stokes.

That origin story alarmed mathematicians because it sounded like a competition between AI laboratories, not a research program organized around scholarly need. A famous problem supplied prestige, measurable difficulty, and an attention-grabbing finish line.

The resulting proof reportedly ran 166 pages. Researchers then faced the slow work of evaluating the argument, tracing its relationship to earlier methods, and extracting ideas humans could understand.

A proof’s value can extend far beyond its final theorem. The tools developed while solving Fermat’s Last Theorem, for example, affected other areas of number theory and later applications.

A long machine-generated proof may also contain reusable ideas. Yet those ideas do not become useful merely because the final pages reach the correct conclusion. Specialists must find, explain, and connect them to existing knowledge.

That labor creates an asymmetry. An AI company pays to generate the output and receives attention for the claim. The academic community then supplies much of the unpaid work needed to establish what the output means.

The Navier-Stokes controversy intensified because researchers viewed this asymmetry alongside the attribution dispute. Critics did not see only an impressive proof. They saw a company converting community knowledge and computational scale into publicity.

OpenAI’s earlier releases had already caused concern. Researchers told The Verge that some manuscripts were poorly written and engaged too lightly with relevant literature.

They also alleged that OpenAI changed documents after publication without clearly recording every revision. Hairer characterized aspects of that practice as sloppy scholarship.

Transparent revision histories are especially important for AI-generated research. They allow readers to distinguish the original claim from corrections made after outside criticism.

Without that record, reviewers cannot easily determine whether a flaw was minor, whether the argument changed materially, or whether public descriptions still match the latest document.

OpenAI created AGMAI partly to prevent another failure of this kind. Yet the company’s own announcement revived the pattern by promoting more than 100 claimed results before the group had developed its process.

That was the reversal. The repair effort became another demonstration of the behavior it was supposed to repair.

OpenAI’s 100-Problem Claim Transfers the Burden to Academia

A backlog of more than 100 claimed solutions is not just a technical milestone. It is a review queue that universities never agreed to process.

Peer review already depends on scarce expert attention. Frontier results require specialists who understand the relevant field, its literature, and the subtle failure modes within a proof.

OpenAI’s claim spans most areas of mathematics. Evaluating it would therefore require many communities, not one centralized panel of nine distinguished researchers.

AGMAI can help triage the output. It can identify appropriate reviewers, recommend disclosure practices, and advise OpenAI on whether a result deserves broad attention.

It cannot personally validate every proof. Nor can nine members absorb the interpretive work produced by an industrial system operating at machine speed.

This creates a scale mismatch. OpenAI can run thousands of agents simultaneously. Mathematical judgment remains distributed among people with teaching, mentoring, publication, and institutional responsibilities.

The volume may also distort research priorities. If OpenAI announces results in a particular field, specialists may feel compelled to pause their own work and inspect the company’s output.

Colva Roney-Dougal, a mathematics professor at the University of St Andrews, described the pending announcements as agonizing. She said researchers face uncertainty about whether to accelerate papers, wait for OpenAI’s disclosures, or continue normally.

That uncertainty reaches graduate students as well as established professors. A student might spend years developing expertise around one problem. A vague corporate hint can suddenly cast doubt over the project without providing enough information to change course intelligently.

The danger is not simply that AI solves a researcher’s chosen problem. Scientific work has always faced the possibility that another person will publish first.

The difference is the combination of scale, secrecy, and publicity. OpenAI can privately pursue many targets, hint at dramatic progress, and release selected results under a globally recognized brand.

Researchers outside the company cannot know which areas face immediate competition. They also cannot match the available compute or communications reach.

The mathematics open letter signed by 25 Fields Medal recipients describes this incentive structure as severely misaligned. It argues that using open problems as AI benchmarks can harm mathematics and its community.

The concern is not a demand to stop AI-assisted research. Many mathematicians already use models for literature searches, experimentation, formalization, and proof development.

Instead, the letter challenges the practice of treating prestigious problems as trophies. That objective rewards speed and visibility, while academic mathematics also depends on attribution, exposition, mentorship, and cumulative understanding.

AI laboratories have incentives to announce the largest recognizable result. Researchers have incentives to spend time checking whether the result deserves that description. Those incentives push work in opposite directions.

OpenAI says AGMAI will help it assess significance before communication. That would be a meaningful improvement if the company consistently follows the group’s advice.

However, AGMAI has no decision-making power inside OpenAI. The company remains responsible for what it releases, how quickly it releases it, and how prominently it frames each claim.

OpenAI also says the group will not advise on the pace of its internal mathematical progress. That boundary protects the company’s research autonomy, but it leaves the largest source of pressure untouched.

The output can keep accelerating. The advisory group can only recommend how the resulting wave reaches shore.

Independence Without Authority Is the Central Tradeoff

AGMAI can provide credibility only by challenging OpenAI, but OpenAI can accept its credibility without surrendering control.

The group’s independence is real in several practical respects. Its members are unpaid, can publish their opinions, and can work with competing laboratories.

Those conditions distinguish AGMAI from a conventional corporate advisory board. They reduce the risk that OpenAI can directly control its membership or public conclusions.

Independence still does not guarantee influence. OpenAI can consult the group without accepting its recommendation. It can also decide what information members receive and when they receive it.

Confidential early access creates another tension. Reviewers need access before publication if they are expected to prevent misleading announcements. Yet confidentiality can temporarily limit their ability to warn the wider community.

The arrangement therefore depends on process. AGMAI needs clear rules for recusal, disclosure, publication, corrections, and disagreements with participating laboratories.

It also needs a way to separate preliminary screening from endorsement. If AGMAI reviews a manuscript, outsiders should not assume every member has verified every line or approved OpenAI’s public framing.

OpenAI’s announcement says the group will help assess significance. Significance is not a mechanical property like whether a proof passes a checker.

It depends on context, novelty, methods, and consequences for a field. Reasonable mathematicians can disagree about those factors even when they agree that an argument is correct.

AGMAI must communicate those distinctions precisely. Otherwise, its name risks becoming a general badge attached to claims that received only limited consultation.

The group also faces a legitimacy problem inside mathematics. Only one initial member reportedly signed the open letter before AGMAI’s launch, although Hairer later added his name to the signatories.

That does not make the membership unrepresentative by itself. However, it highlights the gap between advising a company and representing a broad, decentralized research community.

Mathematics has no single governing body capable of authorizing AGMAI to speak for everyone. The group must earn trust through its actions, not the prestige of its roster.

OpenAI also needs to avoid using the members’ reputations as a substitute for transparent evidence. An advisory relationship cannot validate undisclosed proofs.

This is where the skeptical case becomes strongest. OpenAI may genuinely want better academic standards while also benefiting from the appearance of external oversight.

Both statements can be true. The company can pursue a serious consultation process and still frame that process in ways that advance its public narrative.

Hairer appeared to recognize this risk. He accepted that AI companies would try to present the group’s involvement favorably, while insisting AGMAI could still protect the broader mathematical community.

The arrangement deserves a chance to work. OpenAI’s capabilities are advancing quickly enough that some coordination structure is better than a sequence of surprise announcements.

Still, the measure of success cannot be whether OpenAI avoids another embarrassing headline. It must be whether researchers receive enough notice, documentation, and credit to understand each result without absorbing unreasonable costs.

For developers and enterprise AI buyers, this dispute offers a broader lesson about evaluation. A system’s output is not automatically useful because it passes a correctness test.

Organizations must preserve sources, revisions, assumptions, and human decisions around generated work. A searchable AI knowledge base can support that record, but software cannot replace accountable review.

The same principle applies whether the output is a proof, a legal memo, a market analysis, or production code. Generation speed matters only when verification and institutional understanding can keep pace.

What Happens Next Will Show Whether the Group Matters

Three signals will determine whether AGMAI changes OpenAI’s behavior or merely helps package the next wave of claims.

The first signal is the release protocol for OpenAI’s reported backlog. The company says its model resolved more than 100 open problems, but researchers still lack a full inventory.

A credible protocol would categorize the results by field, difficulty, verification status, novelty, and dependence on prior work. It would also avoid presenting the entire collection as uniformly important.

Advance coordination should give affected specialists time to inspect manuscripts and resolve attribution issues. It should not become a private preview followed by another rushed announcement.

If OpenAI publishes results in manageable batches with clear documentation, that will strengthen the case that AGMAI has practical influence. A sudden publicity campaign built around the total count will weaken it.

The second signal is how the company handles corrections. AI-assisted manuscripts will contain errors, unclear passages, and incomplete citations, just as human manuscripts do.

The important question is whether OpenAI records changes publicly. Each revision should identify what changed, why it changed, and whether the correction alters the main claim.

That practice would answer one of the most concrete complaints about earlier releases. It would also give researchers confidence that they are reviewing a stable and traceable document.

Watch whether AGMAI publishes standards that participating laboratories can follow. Public standards would make the group’s work useful beyond one company and allow outsiders to judge compliance.

The third signal is AGMAI’s response to disagreement. An independent group does not prove its value by agreeing with OpenAI.

Its strongest evidence of independence would be a public objection to exaggerated framing, inadequate attribution, premature release, or an unsupported claim of significance.

That test does not require hostility. It requires visible separation between the group’s analysis and the company’s communications.

Other AI laboratories will also shape the outcome. Anthropic has already become part of the competitive context around research-level mathematics. Google and specialized formal-mathematics projects continue developing their own approaches.

If AGMAI works with several laboratories, it can become shared infrastructure for the field. If its public role remains tied mostly to OpenAI, doubts about its independence will persist.

The underlying capability will not disappear while these governance questions develop. OpenAI’s models have already contributed to work that expert mathematicians consider legitimate and interesting.

The central uncertainty is whether the company can translate that capability into scholarship without overwhelming the people whose judgment gives the results meaning.

For now, OpenAI keeps bulldozing mathematicians because it controls the speed, scale, and publicity of the operation. AGMAI can advise from beside the controls, but it cannot take the wheel.

Readers should watch the next batch of proofs rather than the next promise. Are the documents readable, attributed, independently assessed, and accompanied by transparent revision histories?

Those details will reveal whether OpenAI has changed its research conduct. They will also show whether AI-generated mathematics becomes shared knowledge or simply a growing pile of claims awaiting human repair.

Keep a record of the announcements, manuscripts, corrections, and outside responses. Compare the company’s first description with what specialists conclude weeks later. That gap, more than the headline count, will measure whether this uneasy experiment is working.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page