EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards
This paper introduces EarthVerse, a comprehensive benchmark comprising 405 reproducible tasks across 19 natural hazard families that evaluates the end-to-end scientific reliability of AI agents in dynamic Earth systems, revealing a significant gap between individual step accuracy and consistent, holistic reasoning in current models.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
When scientists try to understand how the Earth works, they rarely have a single, perfect picture of what is happening. Instead, they must piece together a story from many different sources: a weather report from a local station, a satellite image taken from space, a government agency's summary of a flood, and a reanalysis of historical climate data. Each of these sources tells part of the truth, but they often speak different languages, cover different areas, and record events at different times. The real challenge of Earth science is not just reading these reports, but figuring out which ones belong together, how to compare them fairly, and whether they tell a consistent story about a disaster or a changing climate. If a scientist picks the wrong data or mixes up the timing, the entire conclusion about a hazard can be wrong, with serious consequences for how communities prepare for or respond to danger.
A new study called EarthVerse asks a difficult question: can artificial intelligence systems do this kind of scientific detective work on their own? The researchers built a massive test to see if computer agents could act like real scientists. They created 405 specific investigations based on 199 real-world disasters and extreme weather events, ranging from heat waves to earthquakes. Each investigation was a package of about 34 different files, including raw data, maps, and reports. The computer agents were given a scientific question, such as determining the severity of a heat wave, but they were not told which files to look at. They had to search through the entire package, find the right records, check if the data matched up in time and space, perform the necessary calculations, and then write a final answer that was fully supported by the evidence they found.
The results revealed a significant gap between what these systems can do in simple tests and how they perform in complex, real-world scenarios. When the researchers tested 25 different AI systems, the best ones were able to get about 85 percent of the small, individual facts correct. They could often find the right numbers and perform the math correctly. However, when the researchers looked at whether the entire investigation was complete and reliable, the success rate dropped dramatically. Only about 35 percent of the investigations were done well enough to be considered fully trustworthy. This means that even the most advanced systems often made a critical mistake somewhere in the chain of reasoning. They might have found the right temperature record but used the wrong time window, or they might have calculated a statistic correctly but failed to check if the data came from a compatible source.
The study showed that the problem is not a lack of knowledge or an inability to do math. The systems knew the facts and could calculate the numbers. The failure happened when they had to decide which evidence to trust and how to keep all the pieces of the puzzle aligned. When the researchers gave the AI systems extra help by pointing them directly to the correct files, their performance jumped significantly, proving that the main bottleneck was finding the right information, not understanding it. The study also found that simply making the AI think longer or harder did not help if it was stuck looking at the wrong data. The systems needed to be able to change their mind, go back and find a better source, and update their conclusion based on new evidence.
This research suggests that while artificial intelligence is becoming very good at answering questions when the answer is already provided, it is still struggling with the harder task of building the evidence base itself. In the world of Earth science, where a single missing piece of data can change the understanding of a disaster, this gap is critical. The study concludes that for AI to be truly reliable in scientific fields, it must be able to maintain a consistent link between every claim it makes and the specific evidence that supports it, ensuring that the entire story holds together from the first data point to the final conclusion.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.