top of page

ICM Mathematicians Challenge the 2026 AI Replacement Story

Aug 11
12 min read

Kai Williams interviewed more than 20 mathematicians at ICM 2026, and their answers complicated the Techmeme 2026 narrative about AI replacing experts.

Several researchers already use AI to search unfamiliar literature, test examples, review drafts, and fill gaps in proofs. Yet most stopped short of treating these systems as independent mathematical colleagues.

The sharpest disagreement concerns what happens next. One camp expects AI to absorb technical work while people choose questions and develop theories. Another believes that boundary will disappear as models improve.

That dispute matters beyond mathematics. Mathematical proofs offer unusually clear tests of machine reasoning because conclusions can sometimes be checked precisely. The field is becoming an early laboratory for understanding how AI changes expert work.

The ICM Interviews Revealed a Profession in Transition

Mathematicians are no longer debating whether AI will affect their profession. They are debating which parts of the profession will remain distinctly human.

Williams attended the International Congress of Mathematicians in Philadelphia, held from July 23 through July 30. The official congress page identifies the Pennsylvania Convention Center as the venue.

The congress takes place every four years and brings together leading researchers from across mathematical disciplines. Its program combines major awards, research lectures, public talks, and discussions about the profession’s future.

Williams published his 20 mathematician interviews on August 4. The resulting picture was more nuanced than either an automation victory story or a defense of traditional academic work.

The opening example captured that tension. Jacob Tsimerman announced at a July 23 press conference that he was joining OpenAI’s safety team. He had just received a Fields Medal.

Tsimerman told Williams that he expected AI to become “robustly superhuman” at work currently performed by professional mathematicians. He also said he wanted his new visibility to direct attention toward AI safety.

That position came from someone at the profession’s highest level, not an outside forecaster. It made the possibility of broad automation harder for the mathematical community to dismiss.

However, the interviews did not reveal a consensus that mathematicians were about to become obsolete. Williams found that many participants remained optimistic about their own work, especially over the near term.

Some had already received practical help from AI. Their examples were narrower than the public image of an autonomous system discovering entire theories without human direction.

Researchers described using models to locate relevant papers, explain unfamiliar techniques, construct specialized examples, and examine pieces of larger arguments. These applications accelerated projects that still had human-defined goals.

This distinction changes how the Techmeme 2026 story should be read. The immediate shift is not wholesale replacement. It is the arrival of useful systems inside serious research workflows.

That change is still substantial. A tool that compresses days of literature exploration into a focused starting point alters which questions researchers can investigate.

It can also lower the cost of moving between mathematical specialties. Modern mathematics contains many deep subfields whose terminology and methods are difficult for outsiders to penetrate.

A model that helps researchers navigate those boundaries can increase the value of broad curiosity. It can connect a problem in one field with techniques developed somewhere else.

Williams reported that literature discovery was the most common application mentioned by his interviewees. That pattern resembles an advanced research assistant more than an autonomous theorem factory.

Other uses moved closer to mathematical production. Alonso Castillo-Ramirez said ChatGPT constructed a cellular automaton with properties needed for his work.

The example would have been difficult to find through ordinary programming alone, he said. Still, it remained one component within a larger project conceived and directed by a person.

Graduate student Vasiliy Neckrasov described using AI for literature searches, draft review, proof completion, and other mathematical tasks. He first forms a project-level picture before asking the model to pursue details.

That workflow places human judgment before machine execution. The mathematician selects the problem, decides which sources matter, and evaluates whether the generated steps serve the intended argument.

The interviews therefore reveal two changes at once. AI has become useful enough to influence real research, while its most credible uses still depend on substantial human expertise.

That combination explains the mixed mood in Philadelphia. Researchers can welcome faster exploration while worrying about what happens when models begin choosing directions themselves.

Why Techmeme 2026 Is Tracking AI Mathematics So Closely

Mathematics has become a high-value test because recent systems are moving from classroom problems into open research questions.

Three years earlier, leading language models frequently failed at basic arithmetic. By 2025, frontier systems were producing results comparable with top performers on elite high-school competitions.

Research mathematics presents a harder target. Problems often require long arguments, specialized knowledge, creative reformulation, and judgment about which paths deserve attention.

A correct answer is not enough. Mathematicians need an argument that survives detailed scrutiny and connects meaningfully with established knowledge.

Recent claims suggest that AI systems are beginning to cross that boundary. Williams cited work on the Erdős unit distance conjecture and a reported counterexample concerning the Jacobian conjecture.

These examples require careful qualification. Some results involve unreleased systems, company-provided accounts, significant computational resources, or proofs that still need extended human review.

One strong data point comes from Google DeepMind researchers. Their agentic math system was designed as a stateful workspace for literature search, experimentation, theorem proving, and theory building.

The system coordinates multiple agents, maintains working documents, tracks failed approaches, and sends outputs through review cycles. It resembles a research environment more than a single chatbot conversation.

On FrontierMath Tier 4, the researchers reported 23 correct answers among 48 non-public problems, producing an accuracy of 48 percent. Epoch AI conducted the blind evaluation through the system’s interface.

The paper also discloses an important limitation. The system used more inference work than ordinary benchmark harnesses and had no fixed limit on generated tokens or model calls.

That caveat does not erase the result. It shows that system design, tools, parallel investigation, and repeated review contribute significantly to performance.

The same paper describes early collaborations with professional mathematicians. Researchers used the system to investigate open questions, locate overlooked literature, perform numerical experiments, and develop candidate proofs.

The authors also report uneven satisfaction. Some mathematicians found the system less effective for their work, and certain proofs remained under detailed human review.

These details clarify why AI mathematics keeps appearing across Techmeme 2026 coverage. The field offers visible milestones, but each milestone raises questions about evaluation and attribution.

Who selected the problem? Who supplied the key formulation? How much computation was used? Did an independent expert inspect every step?

A headline can collapse those distinctions into the phrase “AI solved a problem.” Mathematical practice demands a more careful accounting of the complete research process.

The strongest systems now combine several capabilities. They search literature, generate conjectures, write code, test examples, draft arguments, and ask formal proof tools to verify specific steps.

Formal verification means expressing an argument in a language that a proof checker can validate mechanically. Lean is one prominent system used for this purpose.

A proof checker can confirm that formal steps follow its logical rules. It does not automatically establish that the theorem is important, the assumptions are appropriate, or the explanation is insightful.

Informal proofs create a different challenge. They communicate ideas in the style used by working mathematicians, but they can hide gaps behind confident or compressed language.

This is why expert review remains essential. A plausible argument can contain a subtle error that survives several rounds of superficial checking.

Google’s researchers describe a related failure mode in their system. Reviewer agents can sometimes converge on a flawed argument whose remaining errors they can no longer detect.

That admission is important. Repeated machine review does not guarantee correctness when the reviewers share similar blind spots.

The benchmark progress is real, but it does not support a simple conclusion that research mathematics has been automated. It shows that sophisticated AI workflows can now contribute to work previously reserved for trained specialists.

The Main Divide Is Human Direction Versus Autonomous Discovery

The defining contest is not mathematicians against machines. It is human-directed augmentation against systems that increasingly choose their own mathematical goals.

Yu Deng offered the more optimistic view at the ICM press conference. He expected AI to handle technical details while mathematicians developed new ideas, theories, and frameworks.

That division resembles earlier waves of automation. Calculators reduced the value of manual arithmetic, while symbolic software automated many algebraic manipulations.

Those tools did not end mathematics. Researchers moved toward questions requiring different concepts, abstractions, and forms of judgment.

Jordan Ellenberg described this process years before the current wave. Once computers make an operation routine, mathematicians often reclassify it as computation and focus elsewhere.

Many ICM interviewees appeared to extend that historical pattern to generative AI. They treated current models as instruments that enlarge the range of work a human can manage.

The problem is that generative systems are moving beyond fixed calculation. They can search for strategies, connect remote ideas, critique drafts, and produce lengthy arguments.

These activities occupy territory that mathematicians previously regarded as creative. The boundary between technical execution and intellectual direction is becoming unstable.

Tsimerman challenged the assumption that theory building provides a permanent human refuge. Earlier predictions said language models would never handle mathematics or research-level problems.

Those predictions weakened as systems improved. He sees the new emphasis on theory building as another moving boundary rather than a durable limitation.

Greg Burnham of Epoch AI raised a similar concern. Evaluating only current abilities can obscure the direction and speed of capability growth.

Still, extrapolation has limits. Current evidence does not establish that scaling will produce systems that consistently formulate valuable theories or select consequential research programs.

The ability to generate a novel definition differs from recognizing that the definition reorganizes a field. Importance depends partly on relationships with other ideas and communities.

Mathematics is filled with questions that are difficult but not especially useful. Choosing a worthwhile problem requires technical taste, historical awareness, and a sense of shared priorities.

Models can learn patterns associated with those judgments. Whether they can develop reliable mathematical taste remains unresolved.

There is also a difference between proposing possibilities and sustaining a research agenda. A successful program can require years of reformulation, communication, mentoring, and coalition building.

Human direction therefore remains central to the most credible AI impact on mathematics. Experts supply goals, interpret failures, and decide whether a result deserves further attention.

The pressure will rise if models become reliable at those tasks. At that point, augmentation could shift toward autonomous discovery.

That transition would affect more than individual productivity. It could change hiring, graduate training, grant allocation, publication standards, and the influence of technology companies.

A university might fund fewer exploratory projects if administrators believe commercial systems can produce results more cheaply. Students could lose the manageable problems traditionally used to develop research judgment.

Private AI laboratories might also gain influence over which mathematical questions receive large computational budgets. Their model access and infrastructure already exceed those of many academic groups.

These concerns shaped the AI research tensions identified by Nature Machine Intelligence. Its editorial emphasized human choice, verification, and the risk of skewed incentives.

The opponent map is therefore institutional as well as technical. Human-directed research values openness, understanding, education, and community development.

Autonomous systems are commonly evaluated through solved problems, benchmark scores, and measurable output. Those metrics capture only part of what the profession values.

The central question is who sets the objective. If mathematicians direct the tools, AI can expand human inquiry.

If benchmark performance and corporate priorities set the agenda, the field could produce more answers while weakening its capacity for shared understanding.

Faster Proofs Can Create a Knowledge Bottleneck

A system that generates correct proofs faster than people can understand them creates abundance at one stage and scarcity at another.

Terence Tao focused on this tension during his public ICM lecture. Mathematics serves several goals, including solving problems, building explanations, educating people, and sustaining intellectual communities.

Historically, those goals often reinforced one another. Solving an important problem could create new techniques, train researchers, and deepen collective understanding.

AI can separate them. A machine-generated proof might settle a conjecture without producing an explanation that humans can absorb or reuse.

Tao warned that mathematics was approaching a situation where a major result could be proved and verified without any person understanding it.

Formal correctness would still matter. Yet a checked proof is not identical to a useful human theory.

The distinction can sound philosophical until it affects the research pipeline. Mathematicians need understandable concepts to teach students and connect results across fields.

If machine output expands faster than expert comprehension, journals and universities face an attention problem. They must decide which proofs deserve scarce human review.

William Thurston’s account of mathematical progress offers a historical warning. He argued that the profession’s real work involves helping people understand and think about mathematics.

Thurston once produced major results in foliations faster than he communicated their underlying ideas. Other researchers left the area, and its supporting community diminished.

AI could reproduce that dynamic at much greater scale. A flood of results might reduce motivation to explore problems whose answers already exist inside opaque machine output.

Timothy Gowers has described a related possibility. The mathematical literature could expand while the community lacks shared expertise about large portions of it.

That would be an unusual form of technical success. Humanity would possess more validated statements but less practical command of the structures connecting them.

The risk begins earlier than fully autonomous superhuman mathematics. Graduate education already relies on problems that are difficult enough to develop skill but limited enough for students to solve.

If AI handles those problems instantly, instructors must design new ways to teach conjecture formation, error detection, and mathematical taste.

Avoiding AI entirely would not solve the problem. Students entering research need experience with tools they will encounter throughout their careers.

Uncritical dependence creates the opposite danger. A student who delegates too much reasoning might complete tasks without developing the judgment needed to evaluate model output.

Verification infrastructure also remains uneven. Some areas have mature formal libraries, while others lack the definitions and supporting theorems needed for efficient formalization.

Informal evaluation relies on a small number of specialists. Their time can become the limiting resource as systems generate more sophisticated claims.

The bottleneck resembles software maintenance. Generating code is cheap, but understanding dependencies, validating behavior, and accepting responsibility remain costly.

Mathematics raises the stakes because errors can sit unnoticed inside long technical arguments. Confidence and polished prose provide weak evidence of correctness.

The Leiden Declaration attempts to establish community principles for this transition. It covers research, education, publishing, funding, policy, and relationships with AI companies.

The Leiden Declaration argues that mathematicians retain choices about whether and how they adopt AI. It identifies understanding, openness, attribution, and human agency as values worth preserving.

The International Mathematical Union endorsed the initiative in June. It was presented and discussed at ICM on July 26.

The declaration does not offer a complete operating model for AI-era mathematics. It provides a starting point for institutions deciding what evidence, disclosure, and access standards to require.

Practical policies could include documenting model involvement, separating machine-generated conjectures from verified results, and preserving human-readable explanations alongside formal proofs.

Researchers also need reproducible access where possible. A result produced by an unreleased model cannot be examined under the same conditions by independent teams.

This matters when technology companies publicize research milestones. Company demonstrations can be valuable evidence without serving as final validation.

The cautious position is not that machine-generated mathematics lacks value. It is that proof production must not become the only measure of mathematical progress.

Human understanding determines whether a result can support education, further discovery, and responsible application. Without that layer, an expanding proof library risks becoming inaccessible inventory.

What the Next AI Mathematics Results Must Prove

The next phase will be judged by independent replication, durable human understanding, and evidence that AI can shape research rather than only complete assigned tasks.

The first signal to watch is independent validation of major open-problem claims. Experts need access to complete arguments, assumptions, computational records, and clear attribution.

When an unreleased system produces a result, outside mathematicians should determine whether the proof is correct and whether the method contains genuinely new ideas.

Successful replication would strengthen the view that AI has crossed into dependable research mathematics. Repeated corrections or withdrawn claims would weaken it.

The second signal is whether AI systems can build useful theories. Solving a well-posed problem begins after someone has already selected the destination.

Theory building requires definitions that expose hidden structure, conjectures that organize evidence, and explanations that other researchers can extend.

A single unusual proposal will not settle this question. The stronger test is whether human experts adopt machine-generated concepts across multiple projects.

That adoption must extend beyond curiosity. Researchers should find that the proposed framework clarifies old results, suggests productive questions, or connects previously separated fields.

The third signal concerns institutions. Universities, journals, funders, and mathematical societies must decide how they treat AI-assisted work.

Disclosure standards will show whether the community can distinguish routine assistance from substantial machine contribution. Review policies will reveal who carries responsibility for correctness.

Training practices deserve equal attention. Programs need ways to preserve deep reasoning while teaching students to use AI as a research instrument.

Funding decisions will expose the balance of power. If resources flow primarily toward proprietary systems, academic mathematics could become dependent on tools it cannot inspect.

Broader access would support the augmentation model. Researchers could test systems, compare methods, and adapt workflows without surrendering their agendas.

Restricted access would reinforce concerns about corporate influence. A few laboratories could determine which projects receive advanced machine support.

The outcome matters to developers and other knowledge workers. Mathematics offers a preview of what happens when AI becomes competent at tasks central to professional identity.

Experts may become more productive while losing control over entry-level work. Verification and judgment may grow more valuable even as raw production becomes cheaper.

The same pattern can emerge in software, law, science, and financial analysis. Organizations will need systems that preserve sources, decisions, revisions, and human context.

A personal knowledge base can support that work by keeping human reasoning connected with the material an AI retrieves or transforms.

The lesson from Techmeme 2026 is not that AI has finished mathematics. It is that credible researchers have moved from dismissing the possibility to negotiating its consequences.

Some see a faster route to discovery. Others see pressure on careers, education, funding, and shared understanding.

Both reactions can be justified by the same evidence. Systems are useful enough to alter current work, but their long-term ceiling remains unknown.

The practical response is to test bold claims without ignoring the trajectory. Researchers should demand reproducibility, preserve understandable explanations, and make human values explicit.

Watch the next independently reviewed proof, the first broadly adopted AI-generated theory, and the policies governing attribution and access. Those signals will reveal whether AI expands mathematical culture or merely its output.

The question for every knowledge-intensive field is now similar: which parts of expert work should machines accelerate, and which parts must remain understandable to people?

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page