News analysis · 19 September 2026
Mantic AI Forecasting: Require a Decision Receipt
By the ELYMENT AI editorial team · Free to read
Mantic’s strong result in the summer 2026 Metaculus Cup shows that specialised AI can produce useful probabilities for messy real-world events. It does not prove that one tournament winner can improve every corporate decision. Before using an AI forecast for capital, inventory, hiring, risk or strategy, require a decision receipt: the exact question, resolution rule, probability and timestamp, evidence available then, comparator, scoring method, decision threshold, human owner and eventual outcome. Forecast quality must be measured before business impact is claimed.

What changed in AI forecasting
Reuters reported on 18 September 2026 that London-based Mantic raised US$25 million in seed funding after its system outperformed every human participant and all but one bot in the summer 2026 Metaculus Cup. Radical Ventures led the round, with Microsoft’s M12, Thinking Machines Lab, Balderton Capital and others participating. Mantic was co-founded in 2024 by Toby Shevlane and Ben Day.
Mantic says it specialises frontier models for forecasting, tests them on historical questions, grades their predictions and improves the system. Its website positions the product for corporate, financial and government decisions spanning demand, regulation, litigation, supply chains and geopolitical risk. That is commercially significant because probabilities can be updated and scored after outcomes resolve, unlike a persuasive but untestable narrative.
A tournament win is evidence, not a universal guarantee
The Metaculus result is a useful proof point, but the scope matters. Reuters says the competition covered political, economic and cultural events. Performance can change with question selection, time horizon, information access, resolution rules and comparison group. Mantic’s own site presents selected examples; those cases help explain the approach but do not establish performance across every customer domain.
Businesses should therefore avoid translating ‘beat humans in a forecasting tournament’ into ‘will improve our decisions’. A forecast may be statistically strong yet operationally irrelevant, delivered too late, based on inaccessible evidence or unable to change an approved action. Conversely, a modest forecasting gain can be valuable when it changes a high-cost decision at the right threshold.
Require a forecast decision receipt
For every consequential forecast, preserve a compact record that makes both prediction quality and business use auditable:
- the exact, time-bounded question and objective resolution rule;
- the probability, timestamp, update history and evidence available at each update;
- the model or service version, configuration and any human adjustment;
- a named baseline, such as the existing plan, expert estimate or market consensus;
- the proper scoring rule and the portfolio of questions used for evaluation;
- the decision threshold, action taken, accountable owner and cost of being wrong; and
- the resolved outcome, realised value and post-decision review.
Score probabilities before measuring ROI
Metaculus explains that a proper scoring rule rewards a forecaster for stating its sincere probability. Its platform uses a logarithmic score, which penalises confident errors sharply, and its Peer score compares one prediction with others on the same question. This is more informative than counting correct calls after converting probabilities into yes-or-no labels.
Run a shadow evaluation before operational use. Pre-register representative questions, lock resolution criteria, capture forecasts before outcomes are known and compare them with the organisation’s current baseline. Review calibration, score, coverage, update timing and performance by domain. Keep the evaluation separate from vendor-selected highlights, then test whether better probabilities would actually have changed decisions.
What leaders should do next
Start with one decision class where uncertainty is material and outcomes resolve often enough to learn. Examples include demand bands, supplier disruption, regulatory timing or project delivery. Set action thresholds in advance, keep a human owner and prohibit the system from silently changing approved plans. Revalidate after model, data-source or scoring changes.
ELYMENT AI helps organisations turn AI claims into acceptance tests, evidence and governed workflows. For forecasting, the decisive question is not whether a model sounds prescient. It is whether the organisation can show what was predicted, when, against which baseline, how it scored and which decision changed as a result.
Sources
- Reuters, AI startup Mantic raises $25 million for superhuman forecasting (18 September 2026) - Independent reporting on the funding round, founders, system approach and summer 2026 Metaculus Cup result.
- Mantic, The world’s most accurate AI predictions (Accessed 19 September 2026) - Primary source for Mantic’s product positioning, target decision areas, team and worked forecasting examples.
- Metaculus, Scores FAQ (Accessed 19 September 2026) - Primary documentation for proper scoring rules, logarithmic scores, peer comparison and time-averaged evaluation.
Continue learning
Frequently asked questions
Did Mantic beat human forecasters?
Reuters reports that Mantic outperformed every human participant and all but one bot in the summer 2026 Metaculus Cup.
What is a forecast decision receipt?
It records the question, probability, timestamp, evidence, model, baseline, scoring rule, decision threshold, owner, action and resolved outcome.
How should a business test an AI forecasting tool?
Use pre-registered, representative questions; fixed resolution rules; proper scores; a current baseline; shadow decisions; and outcome-based review before operational reliance.