An air quality, meteorological and traffic spatio-temporal annotated data set
This paper presents a six-year, hourly spatio-temporal dataset from Madrid that integrates air quality, meteorological, and traffic data from over 5,000 sensors to support urban environmental and mobility research.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the air around us as a giant, invisible soup. Sometimes this soup is clear and fresh, but other times, it gets thick with invisible ingredients like car exhaust, dust from tires, and smoke from factories. This is what scientists call "air pollution." But this soup doesn't just sit there; it changes constantly. The wind might blow the smog away, or a heavy rain might wash it out. The sun might heat it up, making it react in new ways. This is "meteorology," or the study of weather. And what stirs the soup the most? Us. Specifically, our cars. When traffic is heavy, the soup gets thicker with pollution; when the roads are empty, it clears up. This is the connection between "traffic" and "air quality."
For a long time, scientists have tried to understand this relationship, but they often had to look at the ingredients separately. They might have a recipe for the weather, a list of the pollution, and a log of the traffic, but these lists were often written in different languages, at different speeds, and for different places. It's like trying to bake a cake when your flour is measured in cups, your sugar in grams, and your eggs in dozens, all written on different scraps of paper from different years. To really understand how our cities breathe, we need to mix these ingredients together into one perfect, synchronized recipe. That is exactly what this new study has done.
The paper introduces a massive new digital collection called METRAQ, which acts like a giant, time-traveling library for the city of Madrid, Spain. Think of it as a super-organized detective's notebook that brings together three different stories that were previously told separately: the story of the air, the story of the weather, and the story of the traffic. The researchers gathered data from the city's official open data portal, digging through records that go back as far as 2001 for air quality, 2015 for traffic, and 2019 for weather.
The magic of METRAQ isn't just that it has a lot of data; it's how it fits the pieces together. The team took millions of measurements and forced them to speak the same language. They aligned everything to the exact same hour and the exact same location. Imagine a grid of thousands of tiny sensors scattered across the city. At every single hour of the day, METRAQ tells you exactly what the air smelled like, what the weather was doing, and how many cars were zooming by at that specific spot. The result is a dataset that covers 72 months (six years) where all three stories overlap perfectly.
Inside this digital library, you can find 14 different types of air pollutants (like the invisible gases from car exhaust and the tiny dust particles from tires), 7 different weather variables (like wind speed, temperature, and how much it rained), and traffic information derived from over 5,000 sensors. The researchers didn't just copy-paste; they cleaned the data, checked for errors, and even filled in the missing gaps. If a sensor broke for an hour, they used math to guess what the value likely was, marking those guesses clearly so no one gets confused. They even figured out how to estimate traffic levels at spots where there were no traffic sensors, using the data from nearby sensors to paint a complete picture.
The paper doesn't just say, "Here is the data." It also proves that the data works. The authors ran several tests to see if this new dataset could help computers learn to predict the future. They tried to guess what the air quality would be tomorrow based on today's data, and they tried to guess what the air quality was at a specific spot when the sensor there was broken. In these tests, the dataset performed well, showing that it is a reliable tool for scientists and computer programs.
So, what is the big takeaway? The paper suggests that by combining air, weather, and traffic data into one synchronized, high-quality package, we can finally study urban air quality with much greater precision. It doesn't claim to have solved air pollution forever, but it provides the best possible map and compass for anyone trying to navigate the complex relationship between our cars, our weather, and the air we breathe. It turns a messy pile of separate notes into a single, coherent story that can help us build cleaner, healthier cities.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.