← Latest papers
💬 NLP

WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

This paper presents a leakage-free evaluation of six frontier large language models during the 2026 FIFA World Cup, revealing that while the models achieve bookmaker-level accuracy on match outcomes, they exhibit shared limitations such as over-reliance on favorites, poor differentiation among themselves, and a tendency to under-predict draws and specific scorelines.

Original authors: Zhenran Wang, Zhonghan Bian, Jinsong Li, Zhangyang Qi

Published 2026-08-05
📖 7 min read🧠 Deep dive

Original authors: Zhenran Wang, Zhonghan Bian, Jinsong Li, Zhangyang Qi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a detective trying to solve a mystery, but the crime hasn't happened yet. In the world of artificial intelligence, we often test "smart" computer programs by asking them questions about things that have already occurred, like "Who won the 2022 World Cup?" The problem is that these programs might have just memorized the answer from the internet, rather than actually figuring it out. To truly test if a brain is smart, you need to ask it to predict the future—something that doesn't exist yet and can't be Googled. This is the frontier of "forecasting," where we see if a machine can reason its way to an answer that no one knows yet.

This paper, titled WORLDCUP ARENA, takes that idea and runs a massive, real-time experiment. Instead of asking a computer to guess the winner of a game that finished years ago, the researchers set up a live tournament: the 2026 FIFA World Cup. They invited six of the world's most advanced AI models to act as sports analysts. Every day, before a single match kicked off, these AIs had to look at the teams, the weather, the referees, and the latest news, and then make seven different predictions for the game (like who will win, how many goals will be scored, and the exact score). The catch? The games hadn't happened yet. The answers didn't exist anywhere on the internet. This meant the AIs couldn't access the score; they had to actually think.

The researchers wanted to see if these super-smart computers could beat the odds, or if they would just follow the crowd. They found that while the AIs were impressive, they weren't magic. On average, they got about 63.9% of the match outcomes right. That sounds good, but it's almost exactly the same as just blindly betting on the team the bookmakers thought would win. In fact, the AIs were so cautious that they rarely predicted a tie, even though ties happened often, and they kept guessing the same boring score (2–1) over and over again.

Here is the most surprising part: the six different AI models were almost identical in their thinking. They agreed with each other 76% of the time. If you asked them to vote on the winner, the "majority vote" didn't do any better than the best single AI. It turns out that even the most advanced "thinking" models of 2026 are all following the same playbook: they are very good at looking at the big picture (like predicting which country wins the whole tournament) but struggle with the tiny, messy details of a single, close game. The paper concludes that right now, these AI systems aren't sharply different from one another; they are all just very good at mimicking human betting patterns, but they haven't figured out how to truly "see" the future any better than a human expert with a calculator.

The Experiment: A Live Tournament of Minds

The researchers set up a "leakage-free" arena. In normal tests, scientists have to worry that the AI might have "leaked" the answer from its training data. But here, because the matches were happening in real-time in 2026, the answers literally didn't exist when the questions were asked. The setup was rigorous:

  • The Players: Six frontier AI models (Claude, GPT, Gemini, Kimi, GLM, and Seed).
  • The Rules: Each model had to use its "extended thinking" mode (the deepest level of reasoning available) and its own built-in web search tool to read the latest news, injury reports, and odds.
  • The Task: For every one of the 104 matches in the tournament, the AI had to fill out a "prediction card" with seven markets: who wins, the exact score, total goals, and more.
  • The Data: The AIs were given a "dossier" for each team that updated daily, growing from a small summary to a massive file of over 1 million characters by the end of the tournament.

What They Found: The "Herd" Mentality

The results painted a clear, slightly funny picture of how these AIs think.

1. They Follow the Crowd, Not the Crystal Ball
The AIs didn't outsmart the bookmakers. They simply sided with the bookmakers' favorite team 85.6% to 90.4% of the time. Their average accuracy of 63.9% was statistically the same as just mechanically backing the favorite (which gets 64.4%). They weren't discovering new insights; they were echoing the market.

2. The "Draw Blindness"
The AIs were terrible at predicting ties. In the tournament, 27 matches ended in a draw. The models only predicted 8 to 14 draws. They were so afraid of the tie that they barely considered it, even though it happens frequently in football.

3. The "2–1" Obsession
When guessing the exact score, the models got stuck in a rut. They placed 28.0% of all their score predictions on a single result: 2–1. They ignored the vast variety of possible scores and crowded their guesses onto one "prototypical" result.

4. Big Picture vs. Small Details
The models were "macro-strong but micro-weak."

  • Macro: They did great at predicting the tournament winner and group standings.
  • Micro: They fell apart in the closest games. In the Round of 32 (where teams were well-matched), they got 86.5% right. But in the Semi-Finals, where the dossiers were richest with information and the teams were evenly matched, their accuracy collapsed to 8.3%. The more they knew about the teams, the worse they did at picking the winner of a tight game.

5. The "Herd" Effect
The six models agreed with each other 76.0% of the time. Because they were all thinking so similarly, asking them to vote didn't help. A "majority vote" of the six models added nothing to the accuracy. They were all stuck in the same loop.

The Leaderboard: A Tight Race

The final scores were incredibly close, showing that the current generation of AI isn't sharply differentiated.

  • Claude took first place with 897 points.
  • GLM came in last with 813 points.
  • The gap between first and last was only 84 points (about 10.3%).

When the researchers ran statistical tests (called "bootstrap intervals") to see if the winner was truly better, the results were shaky. Claude was the winner in 53.6% of the simulated re-runs, but the difference between the top models was often within the margin of error. For example, the difference between the first and second place was so small that if you ran the tournament again, the order might flip.

The Takeaway

This study suggests that while today's most advanced AI models are powerful tools for gathering information and reasoning, they haven't yet learned to break free from the patterns of the past to truly predict the future. They are excellent at summarizing what we know, but when faced with a truly uncertain event where the answer doesn't exist yet, they tend to play it safe, follow the crowd, and guess the most "average" outcome.

The paper concludes that the current generation of frontier systems is not sharply different from one another. They all share the same strengths (big-picture thinking) and the same weaknesses (fear of draws, obsession with average scores). The "magic" of AI forecasting, it seems, is still just a very sophisticated version of following the betting odds.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →