The FIFA World Cup 2026 Master Dataset: A Normalized 3NF Relational Benchmark of the Expanded 48-Team International Football Tournament
This paper introduces the FIFA World Cup 2026 Master Dataset, a rigorously verified, 3rd Normal Form relational benchmark containing comprehensive match, player, and tactical statistics for the expanded 48-team tournament, designed to advance sports analytics and predictive modeling.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Football has long been a game of passion and instinct, but for the last few decades, a quiet revolution has been turning the sport into a science. Analysts and coaches now rely on vast amounts of structured information to understand what happens on the pitch, moving beyond simple scores to track every pass, shot, and substitution. This field, known as sports analytics, seeks to find patterns in the chaos of a match, using data to explain why a team won or lost, how a player performed, and what might happen next. While commercial companies often hold the most detailed records behind paywalls, the scientific community has long needed open, reliable datasets to test new ideas and improve our understanding of the game. The challenge has always been that available public data is often messy, unconnected, or missing the specific details needed for deep study, such as the exact location of a stadium or the precise market value of a player.
In the summer of 2026, a researcher named MD Mominul Islam from Shahjalal University of Science and Technology in Bangladesh addressed this gap by creating a comprehensive digital record of the FIFA World Cup. This tournament was a historic event, expanding from 32 to 48 national teams and featuring 104 matches played across 16 venues in Canada, Mexico, and the United States. The final match saw Spain defeat Argentina 1–0, but the true significance of the event for data science lies in the sheer volume of activity it generated. The researcher compiled information following the 39-day tournament, gathering details on every single game, the 48 qualified teams, and the 1,248 registered players. The result is a massive, organized collection of facts that links player biometrics, venue geography, and match events into a single, coherent system.
The core achievement of this work is the creation of a "normalized" database, which is a specific way of organizing information so that every piece of data has a unique place and is connected logically to everything else. Instead of having scattered spreadsheets where information might be repeated or contradictory, the researcher built a system with 12 distinct tables that fit together like a puzzle. One table holds the details of the 16 stadiums, including their capacity and, crucially, their elevation above sea level, which can affect how players breathe and perform. Another table tracks the 1,248 players, recording their height, date of birth, international experience, and their estimated market value in euros. A third set of tables captures the 104 matches themselves, linking them to the specific referees, the teams involved, and the final scores. This structure allows a researcher to ask complex questions, such as how the altitude of a stadium in Mexico City influenced the performance of a team from a coastal nation, without having to manually cross-reference dozens of different files.
What makes this dataset particularly valuable is the inclusion of metrics that are often missing from free public data. The collection includes "expected goals," a statistical measure that estimates the likelihood of a shot becoming a goal based on where it was taken and the angle of the shot. It also provides minute-by-minute records of who was playing, when they were substituted, and how long they stayed on the field. For every match, the dataset offers a summary of team tactics, such as how much time a team controlled the ball and how many shots they took. Furthermore, the researcher included pre-calculated features designed to help computers learn from the data, making it easier for others to build models that can predict match outcomes. The data covers the entire tournament, from the opening group stage matches to the final, capturing the full arc of the competition.
To ensure the information was trustworthy, the researcher did not simply copy and paste numbers from the internet. They built an automated system to check the data against official reports from FIFA and other authoritative sources. This process involved running a series of nine different tests to verify that the numbers made sense, alongside a manual spot-check of 30 matches. For instance, the system checked that the total number of teams and players matched the official count, that no match had a score without a corresponding status, and that the sum of goals scored by a player in the detailed event logs matched their total goals in the player statistics. This rigorous verification process confirmed that the data was accurate to a very high degree, with an empirical accuracy rate of 99.87 percent.
The findings presented in this work are not just a list of scores but a verified foundation for future discovery. The data reveals a strong connection between the expected goals a team generates and the actual goals they score, showing that the statistical models used to predict performance align closely with reality. It also highlights the economic disparities between nations, showing the total market value of the squads for the top 15 qualified teams. While the dataset does not include the high-speed, second-by-second tracking of every player's movement due to licensing restrictions, it provides a level of detail that is rare for open-access resources. It captures the physical conditions of the tournament, the tactical decisions of the coaches, and the individual contributions of every player in a format that is ready for analysis.
This dataset is now available to anyone interested in sports science, provided in formats that can be used by standard database software and machine learning tools. It serves as a benchmark for researchers who want to study how expanding a tournament affects the game, how different environments impact player performance, or how to build better models for predicting the future of football. By making this information open and reliable, the work removes a significant barrier for scientists and analysts, allowing them to focus on discovery rather than data cleaning. The result is a clear, concrete record of one of the world's most popular sporting events, preserved in a way that allows the story of the 2026 World Cup to be told not just through highlights and headlines, but through the precise, interconnected facts of the game itself.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.