← Latest papers
💰 quantitative finance

Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents

This paper introduces Foresight Arena, a permissionless, on-chain benchmark that evaluates AI forecasting agents on real-world prediction markets using trustless smart contracts and proper scoring rules to rigorously measure predictive accuracy while preventing overfitting and data contamination.

Original authors: Maksym Nechepurenko, Pavel Shuvalov

Published 2026-05-04
📖 5 min read🧠 Deep dive

Original authors: Maksym Nechepurenko, Pavel Shuvalov

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a giant, digital arena where AI "crystal balls" compete to see the future. This paper introduces Foresight Arena, a new way to test if Artificial Intelligence is actually good at predicting real-world events, or if it's just guessing.

Here is the breakdown of how it works, using simple analogies.

1. The Problem: Why Old Tests Were Flawed

Before this, testing AI forecasters had three big holes:

  • The "Cheat Sheet" Problem: Old tests used static questions (like a fixed quiz). AI models might have seen the answers while they were being trained, so they weren't really "predicting"; they were just remembering.
  • The "Referee" Problem: Most tests relied on a human or a central company to keep score. If that referee was biased or hacked, the results were fake.
  • The "Trader vs. Seer" Problem: Some tests measured how much money an AI made betting on events. But making money isn't just about being right; it's about when you bet and how much you risked. An AI could be terrible at predicting the future but still make money by being a lucky gambler.

2. The Solution: A Trustless, Digital Stadium

Foresight Arena fixes this by building the test on a blockchain (a public, unchangeable digital ledger).

  • The "Commit-Reveal" Game: Imagine a game of poker where you have to write your guess on a piece of paper, seal it in an envelope, and put it on the table before anyone else can see it. Only after everyone has sealed their guesses do you open the envelopes. This stops AI agents from cheating by looking at what others guessed first.
  • The "No-Human-Judge" Rule: The results aren't decided by a person. They are automatically read from a public, decentralized system (Polymarket/Gnosis). If the event happens, the blockchain knows. No one can change the score after the fact.
  • The "Free Entry" Rule: Anyone (or any AI) can join without asking for permission.

3. The Scorecard: Measuring "Truth" vs. "Luck"

Instead of counting money, the arena uses a special math formula called the Brier Score.

  • The Analogy: Think of it like a weather forecaster. If they say "There is a 90% chance of rain" and it rains, they get a good score. If they say 90% and it's sunny, they get a bad score. The score measures how close their confidence was to reality.
  • The "Alpha Score" (The Edge): This is the paper's secret sauce. It doesn't just ask, "Were you right?" It asks, "Were you more right than the crowd?"
    • Imagine a crowd of 1,000 people guessing the winner of a sports game. The "Market Consensus" is the average guess of that crowd.
    • If an AI guesses the same as the crowd, its score is zero.
    • If the AI guesses differently and turns out to be right, it gets a positive score. This proves the AI found a piece of information the crowd missed.

4. The Experiment: Who Won?

The authors ran a live test for 50 rounds involving five top-tier AI models (from companies like Anthropic, OpenAI, Google, etc.) and a "Random Baseline" (a computer guessing numbers randomly).

The Results:

  • The Random Guessers: They failed miserably, proving the test works.
  • The Top AI Models: Three models (Claude, GPT, and Gemini) did slightly better than the crowd, but the difference was tiny. It was like a race where the top three runners finished within a fraction of a second of each other. The test wasn't long enough to say for sure who was the absolute fastest.
  • The "Copycat" Models: Two models (Grok and GLM) tried to guess what the crowd was thinking but added some "noise" (random errors). They actually did worse than just copying the crowd perfectly. This is a key finding: Trying to be slightly different from the crowd when you aren't actually smarter just hurts your score.

5. The "Anatomy" of a Good Prediction

The paper breaks down why an AI wins or loses using a "Murphy Decomposition" (a fancy way of taking apart a score):

  • Resolution (The "Sharpness"): Can the AI tell the difference between a likely event and an unlikely one? The winners were "sharper" than the crowd.
  • Reliability (The "Honesty"): When an AI says "70% chance," does it happen 70% of the time? The winners were very honest about their confidence.
  • The Lesson: To beat the market, you need to be sharper than the crowd. Just being "calm" or "honest" isn't enough if you aren't seeing things the crowd missed.

6. The Big Takeaway

This paper built a permanent, unchangeable reputation system for AI.

  • In the past, an AI company could claim, "Our AI is the best predictor!" and show you a spreadsheet they made up.
  • Now, with Foresight Arena, an AI has a public, unforgeable track record on the blockchain. You can go look at the history yourself.

The Bottom Line:
The paper concludes that while current AI models are much better than random guessing, they are currently very close to the "collective wisdom" of the human crowd. To prove an AI is truly superior, we need to keep running these tests for much longer (hundreds of rounds, not just 50) to see if the tiny differences add up to a clear winner.

What the paper does NOT claim:

  • It does not say these AIs can predict the stock market for profit (that's a different game).
  • It does not say these AIs are ready to run governments or hospitals.
  • It does not claim the current results are the final word; it explicitly states the test needs to run longer to separate the top performers.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →