Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters
This paper introduces "Hindcast," a rigorous evaluation framework that prevents data leakage in LLM forecasting by replaying prediction markets against a frozen, time-cutoff snapshot of public information to ensure models are graded on genuine foresight rather than post-event recall.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to test how good a detective is at solving a mystery. You give them a case file and ask, "Who stole the cookie?" If the detective simply opens the file and reads the answer written at the very bottom, they aren't really showing off their detective skills; they are just showing off their ability to read. This is the tricky problem scientists face when testing Artificial Intelligence (AI) models on predicting the future.
To see if an AI is truly smart, researchers usually ask it to predict things that have already happened, like "Who won the 2022 World Cup?" The problem is that today's AI models have been trained on the entire internet, including articles written after the game ended. So, when the AI says, "Argentina won," it might not be because it figured out the clues beforehand; it might just be remembering the final score from its training data. It's like a student taking a history test who has secretly memorized the answer key. To fix this, scientists need a way to test the AI's "foresight" (guessing what will happen) without letting it cheat by looking at the "answer key" (what actually happened).
This is where a new study called HINDCAST comes in. The researchers wanted to see if AI could actually predict the future if we forced it to stand in the past, with no knowledge of what happened next. They set up a time-travel simulation using prediction markets (places where people bet on outcomes) and a frozen archive of the internet. They found that when you stop the AI from cheating, it can still get better at guessing by reading old news, but only if that news actually had useful facts. If the old news was just people guessing and shouting opinions, reading it actually made the AI worse at predicting the truth.
The Time-Travel Test
The researchers built a system they call HINDCAST (a play on "forecast," but looking backward). Imagine a time machine that freezes a specific moment in history, say, one month before a big sports tournament starts. At this frozen moment, the AI is allowed to look at a library of news articles and forum posts, but only the ones written before that date. It cannot see anything written after that moment, even if it's sitting in the library today.
To test the AI, the researchers used real-world betting markets called Polymarket. These are like digital casinos where people bet on whether something will happen (like "Will Team X win?"). Because these markets have already finished, the researchers know the true answer. They also know what the "market price" was at that frozen moment in the past, which represents what a crowd of humans thought the odds were back then.
The setup works like this:
- Freeze Time: Pick a past date () before a market resolved.
- Lock the Library: The AI can only search a frozen snapshot of Reddit (a popular internet forum) from before that date.
- The Prediction: The AI reads the available posts and guesses the outcome.
- The Score: The AI gets graded on two things: how close it was to the actual final result, and how close it was to what the human bettors thought at that exact moment.
The Big Discovery: Facts vs. Hype
The team tested nine different AI models to see if giving them access to this "frozen library" helped them predict better than just guessing from their own memory.
The Good News: For most of the AI models, reading the old posts did help. In fact, for eight out of the nine models, the AI made fewer mistakes when it could read the archive. The biggest improvement was seen in a model called Qwen3-32B, which reduced its error rate by 23%. This suggests that when there are real facts available in the past, AI can use them to make smarter guesses.
The Bad News (and the Catch): The help wasn't universal. It depended entirely on what the AI was reading.
- Where it worked: In topics like Sports, Awards, and Trading, the internet posts before the event were full of concrete facts (e.g., "The team's star player is injured" or "The price of gold is rising"). In these cases, the AI read the facts and got better at predicting the winner.
- Where it failed: In topics like Entertainment (like guessing which song will be #1) or Politics, the internet was full of loud opinions, fan hype, and speculation. When the AI read these posts, it got confused. It started believing the "hype" was a fact. For example, if fans were screaming that a new song would be #1, the AI would bet on it, even though most songs never actually make it to the top. In these cases, reading the archive made the AI worse than if it had just ignored the internet and stuck to its own logic.
The "Over-Reading" Problem
The study found a funny pattern: the more time a market stayed open, the more likely the AI was to get confused by the noise. If a betting market was open for a long time, the internet filled up with endless chatter and conflicting opinions. The AI, trying to be helpful, read all of it and started to overthink, eventually making bad guesses based on loud opinions rather than solid facts.
One specific model, R1-Distill-Qwen-7B, actually got worse overall when it was allowed to read. The researchers think this happened because the model got so bogged down in reading long, confusing chains of reasoning that it forgot the actual evidence. It's like a student who reads so many different opinions on a topic that they forget the simple facts and end up guessing wrong.
What This Means for the Future
The main takeaway from HINDCAST is that an AI's ability to predict the future is only as good as the quality of the information it can find before the event happens. If the internet is full of facts, the AI can be a great forecaster. If the internet is full of noise and hype, the AI might just become a confident fool, believing the loudest voice instead of the truth.
The researchers also showed that this "time-travel" testing method is a powerful tool. Because the library is frozen, you can test new AI models against the same old questions without them cheating. This means we can keep improving our tests as AI gets smarter, ensuring that we are really measuring how well they think, not just how well they remember.
In short, HINDCAST proves that while giving AI a library of facts helps it see the future, giving it a library of rumors might just make it blind.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.