News analysis · 19 September 2026

Mantic AI Forecasting: Require a Decision Receipt

By the ELYMENT AI editorial team · Free to read

Mantic’s strong result in the summer 2026 Metaculus Cup shows that specialised AI can produce useful probabilities for messy real-world events. It does not prove that one tournament winner can improve every corporate decision. Before using an AI forecast for capital, inventory, hiring, risk or strategy, require a decision receipt: the exact question, resolution rule, probability and timestamp, evidence available then, comparator, scoring method, decision threshold, human owner and eventual outcome. Forecast quality must be measured before business impact is claimed.

A luminous probability signal passes through evidence and decision gates beneath the headline AI Forecasts Need Decision Receipts.
Original ELYMENT.AI editorial illustration.

What changed in AI forecasting

Reuters reported on 18 September 2026 that London-based Mantic raised US$25 million in seed funding after its system outperformed every human participant and all but one bot in the summer 2026 Metaculus Cup. Radical Ventures led the round, with Microsoft’s M12, Thinking Machines Lab, Balderton Capital and others participating. Mantic was co-founded in 2024 by Toby Shevlane and Ben Day.

Mantic says it specialises frontier models for forecasting, tests them on historical questions, grades their predictions and improves the system. Its website positions the product for corporate, financial and government decisions spanning demand, regulation, litigation, supply chains and geopolitical risk. That is commercially significant because probabilities can be updated and scored after outcomes resolve, unlike a persuasive but untestable narrative.

A tournament win is evidence, not a universal guarantee

The Metaculus result is a useful proof point, but the scope matters. Reuters says the competition covered political, economic and cultural events. Performance can change with question selection, time horizon, information access, resolution rules and comparison group. Mantic’s own site presents selected examples; those cases help explain the approach but do not establish performance across every customer domain.

Businesses should therefore avoid translating ‘beat humans in a forecasting tournament’ into ‘will improve our decisions’. A forecast may be statistically strong yet operationally irrelevant, delivered too late, based on inaccessible evidence or unable to change an approved action. Conversely, a modest forecasting gain can be valuable when it changes a high-cost decision at the right threshold.

Require a forecast decision receipt

For every consequential forecast, preserve a compact record that makes both prediction quality and business use auditable:

  • the exact, time-bounded question and objective resolution rule;
  • the probability, timestamp, update history and evidence available at each update;
  • the model or service version, configuration and any human adjustment;
  • a named baseline, such as the existing plan, expert estimate or market consensus;
  • the proper scoring rule and the portfolio of questions used for evaluation;
  • the decision threshold, action taken, accountable owner and cost of being wrong; and
  • the resolved outcome, realised value and post-decision review.

Score probabilities before measuring ROI

Metaculus explains that a proper scoring rule rewards a forecaster for stating its sincere probability. Its platform uses a logarithmic score, which penalises confident errors sharply, and its Peer score compares one prediction with others on the same question. This is more informative than counting correct calls after converting probabilities into yes-or-no labels.

Run a shadow evaluation before operational use. Pre-register representative questions, lock resolution criteria, capture forecasts before outcomes are known and compare them with the organisation’s current baseline. Review calibration, score, coverage, update timing and performance by domain. Keep the evaluation separate from vendor-selected highlights, then test whether better probabilities would actually have changed decisions.

What leaders should do next

Start with one decision class where uncertainty is material and outcomes resolve often enough to learn. Examples include demand bands, supplier disruption, regulatory timing or project delivery. Set action thresholds in advance, keep a human owner and prohibit the system from silently changing approved plans. Revalidate after model, data-source or scoring changes.

ELYMENT AI helps organisations turn AI claims into acceptance tests, evidence and governed workflows. For forecasting, the decisive question is not whether a model sounds prescient. It is whether the organisation can show what was predicted, when, against which baseline, how it scored and which decision changed as a result.

Sources

Continue learning

Frequently asked questions

Did Mantic beat human forecasters?

Reuters reports that Mantic outperformed every human participant and all but one bot in the summer 2026 Metaculus Cup.

What is a forecast decision receipt?

It records the question, probability, timestamp, evidence, model, baseline, scoring rule, decision threshold, owner, action and resolved outcome.

How should a business test an AI forecasting tool?

Use pre-registered, representative questions; fixed resolution rules; proper scores; a current baseline; shadow decisions; and outcome-based review before operational reliance.

Explore ELYMENT AI