← Latest papers
💬 NLP

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

The paper introduces WorldCupArena, a dynamic benchmark for evaluating language models and deep-research agents on football forecasting by testing their ability to predict match outcomes, scores, and events using real-time information before matches occur, with initial results showing modest gains over betting and human baselines in exact accuracy but clearer improvements in granular scoreline predictions.

Original authors: Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang

Published 2026-07-21
📖 5 min read🧠 Deep dive

Original authors: Zhaokai Wang, Tianlin Gui, Jiayuan Rao, Shangzhe Di, Yihong Tang, Dingli Liang

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to guess the outcome of a big event before it happens. In the world of artificial intelligence, we have a lot of smart computer programs called "Large Language Models" (LLMs). Think of them as super-readers who have memorized almost everything written on the internet. Usually, we test them by asking questions with answers that already exist, like "Who was the first president?" or "What is the capital of France?" But what if we want to know if these computers can actually predict the future? That's where "forecasting" comes in. It's like asking a weather forecaster to predict tomorrow's rain, or a sports fan to guess who will win the big game. The tricky part is that the computer has to make its guess using only the information available right now, without peeking at the final score. This paper steps into the chaotic, exciting world of football (soccer) to see if these AI models can act like expert sports analysts, or if they just get lucky.

The researchers behind this study, titled "WorldCupArena," decided to put 13 different AI systems to the ultimate test: the 2026 FIFA World Cup. They didn't just ask the computers "Who will win?" They asked for a full "match report" before the game even started. Imagine a psychic who, before a match, predicts not just the winner, but the exact score, which players will start, who will get a yellow card, how many times the ball will be kicked, and even who will win the entire tournament. The AI had to make these guesses 24 hours before kickoff, using either a pre-packaged list of facts or by searching the web for clues themselves. Once the real games were played, the researchers compared the AI's crystal ball predictions against the actual results to see who was the best forecaster.

Here is what they found, and it's a bit of a surprise. First, just guessing the winner correctly isn't enough to tell the models apart. Many of the AI systems got the "win or lose" part right about 68% of the time, which is pretty good, but it didn't show us who was truly the smartest. The real differences showed up in the details. Some models were great at predicting the exact score, while others were better at guessing which players would be on the field. The paper discovered that adding a "search engine" feature to the AI didn't always make it better at predicting the game. In fact, for some of the models, searching for more info actually made their predictions slightly worse compared to just using the facts they were given. It's like a detective who finds too many clues and gets confused, rather than solving the case faster.

The study also looked at how close the predictions were, even when they weren't perfect. If a model predicted a 2-1 score and the game ended 3-1, it got partial credit for being "close." On this "closeness" scale, the best AI systems did show a clear improvement over human fans and even professional betting markets, but only in a specific way: they were better at guessing scores that were near the real result, rather than hitting the exact number every time. They didn't magically become perfect fortune-tellers.

When it came to the whole tournament, the results were fascinating. Four different AI systems correctly predicted that Spain would be the champion. However, only two of those four could also correctly guess that Spain would play against Argentina in the final. This shows that getting the big picture right doesn't mean you got the whole story right. The paper also highlighted a funny flaw: when the AI models all agreed on a favorite team (like Spain beating a smaller team), they sometimes failed spectacularly together if the game turned into a boring 0-0 draw. They all followed the crowd and got it wrong at the same time.

In the end, WorldCupArena proves that while AI is getting better at sports forecasting, it's not there yet. It can't consistently beat the odds or predict the exact score of every game. But by breaking down the predictions into tiny pieces—like who gets a card, how many shots are taken, and the exact scoreline—the researchers found that these models have different strengths and weaknesses. Some are good at the big picture, others at the details. The paper suggests that simply giving an AI more information to search for doesn't automatically make it a better predictor. Instead, the best forecasters seem to be the ones that can balance the facts they know with a good guess, without getting too confident when they are wrong. This benchmark isn't just about football; it's a new way to test if our smart computers can really think ahead, or if they are just good at guessing what we want to hear.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →