top of page

Mantic AI Forecasting Gets $25 Million After Beating Every Human Entrant

5 days ago
14 min read

Mantic AI forecasting secured $25 million after its system outperformed every human entrant in a recent forecasting tournament. The result gives the London startup a compelling claim, but not a universal victory over human judgment.

The seed round was led by Radical Ventures, with backing from Microsoft’s M12, Thinking Machines Lab, Balderton Capital, and other investors. Mantic now has capital to turn a tournament result into a product for businesses, governments, and financial institutions.

That transition creates the real tension. Forecasting competitions reward calibrated probabilities across clearly defined questions. Organizations must make costly decisions amid incomplete information, shifting objectives, and consequences that a leaderboard cannot measure.

Mantic’s closest opponent is therefore not another startup. It is the existing human forecasting workflow, which combines analysts, domain experts, executive judgment, and sometimes prediction markets.

Mantic AI Forecasting Turns a Tournament Win Into $25 Million

Mantic converted a measurable forecasting result into one of the clearest commercial bets yet on automated judgment.

The company announced the seed round on September 18, 2026. Radical Ventures led the financing, while M12, Thinking Machines Lab, Balderton Capital, and other investors participated.

The company did not disclose its valuation in the announcement covered by Reuters. An earlier report placed the expected valuation near $90 million, but that figure was attributed to an unnamed source.

Mantic was founded in London in 2024 by CEO Toby Shevlane and CTO Ben Day. Shevlane previously worked as a research scientist at Google DeepMind, including research connected to AI capabilities and geopolitical developments.

Day holds a machine-learning doctorate from the University of Cambridge. His prior work included industrial optimization and AI applications in drug development, according to Mantic’s published team information.

The company previously raised $4 million in pre-seed financing. Episode 1 led that round, with participation from trading firm DRW and angel investors connected to major AI laboratories.

The new financing followed Mantic’s performance in the Summer 2026 Metaculus Cup. Metaculus runs tournaments where participants assign probabilities to political, economic, technological, and cultural outcomes.

According to the reported seed financing details, Mantic finished ahead of every human contestant. Only one automated entrant, called laertes, recorded a higher result.

That distinction matters. Mantic did not beat every possible human forecaster, nor did it settle whether machines possess better judgment across all circumstances.

It beat the humans who entered one structured competition under that competition’s questions, schedule, and scoring system. This remains an important result, but its boundaries deserve equal attention.

The questions covered events whose answers became observable during the tournament. Participants had to state probabilities instead of offering vague predictions that could later be reinterpreted.

One example involved Colombia’s presidential election. Shevlane told Reuters that Mantic assigned Abelardo De La Espriella roughly a 40% winning probability early in the tournament.

The prevailing forecast was closer to 30%, and De La Espriella ultimately won. That example suggests the system did not merely reproduce the crowd’s consensus.

Another question concerned whether Shakira’s “Dai Dai” would overtake “Waka Waka” on the Billboard Hot 100. Human participants heavily favored one outcome, while Mantic placed less confidence in that consensus.

The less popular outcome occurred. Shevlane presented the result as evidence that the system resisted herd behavior, although one example cannot establish a general pattern.

The financing turns those outcomes into a larger proposition. Investors are betting that Mantic can repeat its performance across more questions and connect probabilities to consequential decisions.

That is why this is more than a routine seed announcement. Mantic has gained funding because forecasting produces an unusually direct test of an AI system’s real-world reasoning.

A forecast eventually resolves. The system can be scored, compared with human judgment, and improved using outcomes that were unknown when the prediction was made.

The harder question begins after that score. Mantic must show that accurate probabilities can improve decisions inside organizations where incentives, timing, and accountability remain complicated.

Why Forecasting Became an AI Investment Target Now

AI forecasting is attracting capital because newer systems can research, update, and score predictions at a scale that human teams cannot sustain.

Traditional machine learning performs well when organizations have abundant structured data and stable relationships between inputs and outcomes. Judgmental forecasting presents a different problem.

Judgmental forecasting involves events that require contextual reasoning rather than repeated measurements alone. Elections, regulatory decisions, armed conflicts, executive departures, and product launches fit this category.

These events rarely follow clean historical patterns. Relevant evidence appears across news reports, public statements, economic releases, social behavior, and analogous cases.

Human forecasters can combine those signals. However, good forecasters are scarce, and their attention cannot expand across thousands of continuously changing questions.

Large language models create a potential scale advantage. They can gather information, compare historical cases, generate competing arguments, and update numerical probabilities whenever new evidence arrives.

That capability does not automatically make an AI accurate. It changes the economics of testing and iteration, which helps explain investor interest.

A human forecaster may wait months before learning whether a technique worked. An automated system can be evaluated across large collections of previously resolved questions.

Mantic says its system handles predictions from one week to one year ahead. Its stated domains include geopolitics, business, policy, technology, and culture.

These are areas where organizations already pay analysts to interpret weak signals. They are also areas where a wrong assumption can redirect capital, staffing, inventory, or policy.

The company’s public forecasting system targets financial analysts, corporate strategists, risk teams, researchers, and policymakers. Its examples include interest rates, sanctions, litigation, leadership changes, and supply disruptions.

For a hedge fund, an improved probability can inform a trade. For a manufacturer, it can influence inventory or sourcing decisions before a disruption becomes obvious.

Governments could use forecasts to prioritize contingency planning. A corporate strategy group could track the probability of a competitor launching a product within a defined period.

Radical Ventures partner Aaron Rosenberg told Reuters that companies and government agencies had expressed interest. He also said some organizations had integrated Mantic’s AI, although the company declined to identify customers.

The lack of named customers limits independent assessment. Still, the range of interested buyers illustrates why forecasting can become a commercial AI category.

The product is not another conversational interface waiting for a user to ask the right question. It promises a persistent probability layer that can monitor risks and opportunities.

That promise arrives as general-purpose models become better at research and multi-step reasoning. Systems can now retrieve current evidence, compare sources, and repeat workflows without constant human prompting.

Earlier evidence was much less favorable to machines. A real-world forecasting study found GPT-4 significantly less accurate than human-crowd forecasts in a live tournament.

The important word is “live.” When answers remain unknown, a model cannot rely on memorized solutions from familiar benchmark material.

The new Mantic result suggests that specialized systems can close or cross that gap. The improvement appears to come from the complete forecasting pipeline, not one foundation model operating alone.

That distinction also explains the funding thesis. Frontier models are becoming inputs, while startups compete through research orchestration, evaluation data, updating policies, calibration, and customer integration.

If those layers produce consistently better forecasts, their value can persist even as the underlying models become widely available.

The investment target is therefore not prediction as spectacle. It is an operating system for continuously measuring uncertainty around decisions that organizations already need to make.

The Real Opponent Is the Human Forecasting Workflow

Mantic is competing against a slow and fragmented decision process, not against human intelligence in the abstract.

Organizations rarely rely on one forecaster. They combine reports, meetings, spreadsheets, outside consultants, expert opinions, and executive intuition.

Each input may contain valuable knowledge. The weakness lies in converting those inputs into consistent probabilities that can be tracked and scored over time.

An analyst might describe a regulatory approval as likely. Another might call it plausible. A third might recommend waiting without stating any probability.

Those expressions are difficult to compare. They also allow people to reinterpret their earlier confidence after the outcome becomes known.

A forecasting platform requires a sharper commitment. A 70% prediction should occur about seven times across ten comparable cases.

This property is called calibration, meaning the stated probability should match the observed frequency over many forecasts. Calibration turns confidence into something an organization can audit.

Human forecasting teams can do this well. Research associated with the Good Judgment Project showed that trained generalists could outperform conventional expert processes on geopolitical questions.

The strongest forecasters decomposed problems, used comparison classes, updated frequently, and resisted ideological certainty. Those habits are teachable, but maintaining them across an enterprise requires discipline.

Mantic’s advantage is repeatability. Software can apply the same process to every tracked question, update on a schedule, and preserve a record of each probability change.

The system can also cover more questions than a small analyst group. This scale matters because decision-makers often need broad monitoring before they know which issue deserves deeper attention.

For example, a company could track leadership changes across an entire market. It could then assign human analysts only when a probability crosses an internal threshold.

That division of labor creates a more credible commercial path than fully replacing analysts. Machines provide breadth and continuous updating, while people investigate consequences and decide what action follows.

Mantic nevertheless faces a strong incumbent in the aggregated human crowd. A crowd can combine independent information and cancel some individual biases.

Prior competition analysis found that the Metaculus community prediction remained one of the platform’s most consistent performers. That aggregate previously ranked above Mantic in a quarterly event.

Prediction markets offer another route. Platforms such as Kalshi and Polymarket use financial incentives and market prices to collect dispersed beliefs.

Markets can update quickly when participants encounter new information. They also provide transparent probabilities when sufficient trading activity exists.

Their weakness is uneven coverage. Traders concentrate on questions that attract attention, liquidity, or entertainment value, rather than every question an organization needs answered.

A company may care deeply about a narrow supplier regulation. That question might never attract enough outside participation to produce a meaningful market price.

Human experts present the opposite tradeoff. They offer deep context and can understand institutional details, but their work remains expensive and difficult to scale.

Mantic’s central bet is that an automated system can occupy the space between those approaches. It can provide tailored coverage while retaining some research depth and probabilistic discipline.

The claim becomes commercially useful even before the system beats the world’s best human forecaster. Matching a good crowd across thousands of specialized questions would already change the workflow.

An enterprise buyer should therefore avoid framing adoption as human versus machine. The better test compares complete decision processes under the same conditions.

One group might use conventional analysis. Another might receive Mantic probabilities with explanations and updates. Both groups would face equivalent decisions with measurable outcomes.

That comparison would reveal whether the system improves timing, confidence, and resource allocation. It would also show when human experts add value beyond the numerical forecast.

The pressure falls most heavily on research processes that do not preserve predictions or evaluate past performance. Mantic makes unmeasured judgment harder to defend.

It also pressures analysts to separate two tasks. Estimating what will happen is different from deciding what an organization should do about it.

Even a perfectly calibrated forecast cannot choose an organization’s risk tolerance. It cannot decide whether a low-probability outcome deserves preparation because its consequences would be severe.

Mantic can challenge the forecasting layer. Responsibility for action still belongs to people and institutions.

How Mantic Trains a Forecasting System

Mantic’s mechanism combines frontier models, historical backtesting, repeated scoring, and training against outcomes that eventually became known.

The company does not claim to have built a foundation model entirely from scratch. Shevlane told Reuters that Mantic specializes models from other AI laboratories for forecasting tasks.

The process begins with questions that have precise resolution criteria. Instead of asking whether an economy looks healthy, a question must identify a measurable event and deadline.

The system then researches available evidence and assigns probabilities to possible outcomes. It can revisit those probabilities as new information arrives.

Once the event resolves, the forecast receives a score. Probabilistic scoring penalizes confidence in wrong outcomes more heavily than cautious uncertainty.

A common example is the Brier score, which measures the squared difference between a predicted probability and the actual result. Lower scores indicate better predictions.

Repeated scoring provides a training signal. The system can learn which research methods, evidence types, reasoning patterns, and updating rules correlate with better outcomes.

Mantic says it can replay historical periods while restricting information to material available at the simulated date. That allows developers to test strategies without waiting months for new events.

The company reported building a dataset containing more than 10,000 forecasting questions from 2024 and 2025. It also reported backtesting on 348 questions from the second quarter of 2025.

Those numbers come from Mantic’s own technical explanation. They have not been independently audited in the materials reviewed for this article.

Backtesting creates a major speed advantage. Developers can modify a system, rerun historical questions, and compare scores without waiting for every event to resolve again.

Mantic also says it uses reinforcement learning, which adjusts behavior using rewards derived from forecasting outcomes. The objective is not eloquent prose but better probabilistic accuracy.

This changes how a language model’s reasoning gets evaluated. A convincing explanation receives no special credit if the associated probability performs poorly.

The system can also test different components separately. One model might gather evidence, another might challenge assumptions, and an aggregation step might combine several estimates.

Mantic has not publicly disclosed every component of its current production architecture. External readers should not assume one model independently produces each final probability.

Shevlane has described the broader approach as specializing frontier models. The company’s defensible advantage may therefore reside in orchestration, proprietary evaluation data, and training methods.

The process has three attractive properties for investors. It produces measurable outcomes, supports rapid iteration, and can scale across many questions.

It also has an unusual feedback loop. Every resolved question adds evidence about the system’s calibration and decision rules.

However, backtesting demands careful controls. A model may have encountered the outcome or related reporting during its original training.

If historical answers leak into a test, the resulting score measures memory alongside forecasting skill. That problem becomes harder when the foundation model’s training data remain undisclosed.

Recent leakage research argues that backtest scores alone cannot determine how much hidden information influenced a result. Researchers need external reference points and explicit assumptions.

Live tournaments reduce that particular risk because outcomes are unknown at prediction time. This makes Mantic’s Summer 2026 result more informative than a private retrospective benchmark.

Live evaluation introduces different complications. Question selection, update frequency, participation rules, and scoring formulas can affect rankings.

A system that monitors news continuously has a natural coverage advantage over a person forecasting during limited hours. That advantage is commercially useful, even if it complicates scientific comparisons.

This distinction should not be treated as a flaw. Businesses often want the always-on system precisely because humans cannot update hundreds of probabilities every hour.

The important requirement is transparent evaluation. Buyers need separate measurements for accuracy, calibration, coverage, timeliness, and decision impact.

A single tournament rank compresses those dimensions into one result. Product adoption requires opening them again.

What the Superhuman Label Does Not Prove

“Superhuman” accurately describes one contest result, but it remains too broad as a general claim about Mantic’s abilities.

The strongest verified statement is narrow. Mantic outperformed every human contestant in the Summer 2026 Metaculus Cup and finished behind one other automated system.

That result matters because the questions were live. It does not establish superiority across every domain, time horizon, user group, or decision environment.

Tournament participants also differ in experience and activity. Some humans may answer fewer questions or update less frequently than an automated entrant.

Earlier reporting about forecasting competitions noted that scoring can reward coverage. Systems benefit when they answer early, cover more questions, and update regularly.

Those behaviors are valuable in practice. However, a ranking that combines coverage and accuracy cannot reveal whether the machine made better judgments on every matched prediction.

The comparison class also matters. Beating all tournament participants is not identical to beating a carefully selected team of professional superforecasters.

A system might dominate general questions yet struggle with obscure regional politics, technical regulation, or decisions shaped by private information.

Organizations also ask questions that differ from public tournaments. Internal forecasts may concern confidential product schedules, negotiations, customer behavior, or operational failures.

Mantic must show that its methods transfer when public web research provides only part of the evidence. Enterprise deployments may require secure access to internal documents and structured company data.

That integration creates governance questions. A probability can influence hiring, capital allocation, insurance, trading, public policy, or security planning.

Users need to know when evidence is missing, when sources conflict, and when a forecast shifts because one uncertain report entered the system.

Explanations help, but fluent explanations can produce false confidence. A plausible narrative does not guarantee that the underlying probability is calibrated.

Forecasts can also affect the events they describe. A public prediction about a bank, election, or conflict might influence behavior and alter the outcome.

Private forecasts introduce another concern. Customers may act on predictions that outside researchers cannot inspect, making independent performance evaluation difficult.

The undisclosed customer list leaves several important questions unanswered. Readers cannot yet examine how organizations use Mantic or whether the system improved their decisions.

Accuracy does not equal usefulness by itself. A forecast must arrive early enough to change an action, and the potential benefit must exceed implementation costs.

Consider a supply-chain team facing a 20% probability of new export restrictions. The probability matters only when paired with response options and their costs.

The organization might diversify suppliers, increase inventory, or wait. Mantic can inform that choice, but it cannot supply the company’s risk preferences.

There is also a danger of automation bias. Decision-makers may defer to a numerical forecast because it appears objective, even when the model lacks essential context.

The opposite problem remains possible. Executives may ignore a calibrated forecast whenever it conflicts with intuition or organizational politics.

Neither failure is primarily technical. They concern how institutions assign authority, record decisions, and review mistakes.

Mantic’s own examples support a measured interpretation. Its useful contribution is not certainty about the future but disciplined probabilities that update with evidence.

A 70% forecast still fails three times in ten when calibrated correctly. Users who treat that figure as a guarantee will misunderstand the product.

The same principle applies to rare events. A system can assign a low probability responsibly, yet the event can still occur and cause extensive damage.

Buyers should ask for performance broken down by domain, horizon, question type, and information availability. Aggregate scores can conceal weak categories.

They should also compare frozen forecasts with continuously updated ones. An accurate last-minute update serves a different purpose from an early warning delivered months ahead.

Confidence intervals and sample sizes matter as well. A run of correct predictions can reflect skill, favorable question selection, or chance.

Mantic’s financing gives it resources to produce broader evidence. Until that evidence arrives, “superhuman forecasting” should remain a testable product claim.

Three Signals That Will Test Mantic’s Bet

Mantic’s next phase will be judged by live performance, disclosed adoption, and evidence that forecasts improve decisions outside tournaments.

The first signal is continued performance in prospective competitions. Prospective means the outcome remains unknown when every forecast is recorded.

Mantic needs to perform well across several tournaments, question sets, and time horizons. One victory can create attention, but repeated results establish reliability.

The most useful disclosures would include raw accuracy, calibration, coverage, update frequency, and comparisons against the same human forecasts. Domain-level results would reveal where the system struggles.

The current Metaculus tournament offers one venue for continued observation. Future results can strengthen or weaken the claim that the Summer performance was repeatable.

Finishing near the top again would support Mantic’s technical thesis. A sharp decline would suggest sensitivity to question selection, competitors, or tournament conditions.

The second signal is named customer adoption with measurable workflows. Mantic and its investors say companies and government agencies have shown interest.

Interest is not the same as deployment. A meaningful case would explain what questions the customer tracked, how forecasts entered decisions, and what changed.

Financial institutions are an obvious early market because probability improvements can connect to measurable returns. They are also unlikely to disclose proprietary strategies.

Corporate risk teams might offer more visible evidence. They could report better warning times, broader coverage, or fewer missed developments without revealing confidential decisions.

Government adoption would demonstrate relevance to geopolitics and policy. It would also raise stronger requirements for auditability, procurement, security, and accountability.

A named deployment would strengthen the commercial case even without revealing revenue. Continued secrecy would leave outsiders dependent on investor statements.

The third signal is product evidence that separates prediction quality from decision quality. Mantic must show more than a stream of probabilities.

Useful products should preserve forecast histories, explain major updates, identify influential evidence, and show calibration across comparable questions.

They should also support decision thresholds. A customer might specify what action follows when a risk rises above 30%, 50%, or 70%.

That connection would make forecasts operational instead of decorative. It would also allow organizations to audit whether users acted consistently.

Mantic’s system could become most valuable as an attention allocator. It might scan hundreds of questions, flag meaningful changes, and direct specialists toward emerging risks.

This model preserves a role for human expertise. People evaluate consequences, challenge missing context, and choose actions, while software maintains broad probabilistic coverage.

The company’s $25 million round gives it room to test that model. The capital does not resolve whether automated forecasting will earn lasting institutional trust.

Mantic AI forecasting now has a verified tournament result, well-known investors, and a mechanism designed for rapid improvement. Its next evidence must come from repeated live tests and real decisions.

For developers, the lesson is to evaluate the entire forecasting pipeline, including retrieval, calibration, updating, and leakage controls. A capable base model is only one component.

Enterprise buyers should begin with a recorded pilot rather than broad automation. Select defined questions, preserve human baselines, set decision thresholds, and score every resolved outcome.

Knowledge workers should treat probabilities as structured evidence, not instructions. Keep the supporting sources and assumptions available through a searchable knowledge base.

The decisive question is not whether an AI can win a forecasting contest. It is whether your organization can test its forecasts, act appropriately, and learn from every result.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page