TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
This paper introduces TerraBench, a comprehensive benchmark and executable framework designed to evaluate and advance the ability of AI agents to perform grounded, multi-step reasoning across heterogeneous Earth-system data by integrating language models with specialized scientific tools for environmental analysis, simulation, and verification.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to solve a complex mystery about the weather, but you have two very different kinds of detectives on your team:
- The "Word Wizard" (Large Language Model): This detective is amazing at reading books, understanding language, and planning a strategy. They can tell you what steps to take to solve the problem. However, they have never seen a weather map, can't read a satellite photo, and if you ask them to calculate a number, they might just guess.
- The "Data Cruncher" (Scientific Tools): This detective is a robot with a supercomputer. It can instantly read satellite images, run complex climate simulations, and calculate exact numbers. But it doesn't speak human language, can't understand a question, and needs someone to tell it exactly what to do.
The Problem:
Right now, these two detectives don't work well together. The Word Wizard can't use the Data Cruncher's tools effectively, and the Data Cruncher can't understand the Word Wizard's plans. This makes it very hard to build an AI that can actually help scientists make real decisions about climate change, farming, or disaster response.
The Solution: TerraBench and TerraAgent
The authors of this paper built a new testing ground called TerraBench and a new "manager" called TerraAgent to see if AI can finally learn to make these two detectives work as a team.
Think of TerraAgent as a project manager. When you ask a question like, "How will a hurricane affect the crops in Florida next week?", the manager:
- Plans: Breaks the question down into steps.
- Calls Tools: Tells the Data Cruncher to look up satellite images, check weather history, and run a simulation.
- Checks Work: Looks at the results the robot gives back.
- Reports: Writes the final answer in a clear, structured format.
The Test (TerraBench)
To see if this manager is good, the authors created a massive exam called TerraBench. It's not a simple multiple-choice test. It's more like a series of 403 complex "missions" that require the AI to:
- Fundamentals: Read a map and a weather chart to answer a basic question.
- Simulation: Pretend a disaster happens (like a flood) and use a simulator to predict the damage.
- Verification: Check if the AI's answer matches real scientific papers and data.
The test is incredibly strict. It doesn't just care if the AI picked the right tool; it cares if the AI used the right settings on that tool and if the final number is correct within a tiny margin of error.
The Results: A Reality Check
The authors tested the smartest AI models available (like Claude Sonnet 4.6 and Qwen3.5) on this exam. The results were surprising:
- Good at Planning, Bad at Math: The top AI models were decent at figuring out which tools to use and in what order (scoring about 59% on the "process" part).
- Terrible at Precision: However, when it came to getting the final numbers right, they failed miserably. The best model only got the final answer correct about 23% of the time.
- The "Almost" Trap: Most of the time, the AI picked the right tool but used the wrong settings (like looking at the wrong date or the wrong location), leading to answers that were close but scientifically useless.
The Analogy of the "Wrong Recipe"
Imagine you ask a chef (the AI) to bake a cake using a specific recipe (the scientific data).
- The chef knows exactly which ingredients to grab (Tool Selection).
- But when they measure the flour, they use a cup instead of a gram scale (Wrong Arguments).
- The result is a cake that looks like a cake but tastes like sand.
The paper shows that current AI is great at knowing what to do, but it struggles to do it precisely enough to be trusted with real-world science.
Conclusion
The paper concludes that to build a truly useful AI for Earth science, we can't just give it access to tools. We need to teach it how to coordinate those tools with extreme precision, handle complex workflows, and ensure the final numbers are scientifically accurate. TerraBench is the first tool to measure exactly how far we are from that goal.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.