No One Knows the State of the Art in Geospatial Foundation Models
This paper argues that the geospatial foundation model (GFM) community lacks a clear state of the art due to inconsistent evaluation standards, data, and model releases, and proposes six concrete expectations to establish necessary community norms for meaningful comparison and progress.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a massive, chaotic marketplace where hundreds of vendors are selling "magic maps" (Geospatial Foundation Models). These maps are supposed to help us solve big problems like predicting floods, tracking deforestation, or finding crops. Everyone claims their map is the best, but if you try to compare them, you realize nobody actually knows who is winning.
This paper argues that the field of "Geospatial Foundation Models" is currently in a state of confusion because the vendors aren't following the same rules of the game. Here is the breakdown of the problem and the proposed solution, using simple analogies.
The Problem: A Marketplace Without a Rulebook
The authors audited 152 research papers (like checking the receipts of 152 vendors) and found three major problems that make it impossible to tell which model is actually the best.
1. The "Secret Recipe" Problem (Missing Weights)
- The Issue: 39% of the papers don't share the actual "magic map" (the model weights). It's like a chef saying, "My soup is the best in the world," but refusing to let anyone taste it or see the recipe.
- The Consequence: If you can't see the soup, you can't verify if it's actually good, and you can't cook it yourself to see if it works.
2. The "Different Playgrounds" Problem (No Shared Benchmarks)
- The Issue: Imagine a sports league where Team A plays basketball on a court in New York, Team B plays soccer on a field in London, and Team C plays tennis in Tokyo. Then, they all claim to be the "best athletes."
- The Reality: The papers use 401 different testing datasets (playgrounds). The top 3 most common ones are used in only 10% of the tests. Most papers test on unique, one-off datasets that no one else uses.
- The Consequence: You can't compare Team A's basketball score to Team B's soccer score. The authors found that 35% of papers don't even test on the most common datasets, making a fair ranking impossible.
3. The "Magic Numbers" Problem (Inconsistent Results)
- The Issue: This is the most shocking finding. When two different papers test the exact same model on the exact same dataset, they get wildly different scores.
- The Analogy: Imagine two people testing the same car on the same track. One says it goes 60 mph; the other says it goes 120 mph. Both used the same car and the same track, but they didn't agree on how to measure it.
- The Reality: The authors found 46 cases where the same model got a score difference of at least 10 points (sometimes as high as 56 points!) just by changing how the test was run. This means if you read a paper claiming a model is "90% accurate," you have no idea if that number is real or just a fluke of how they ran the test.
4. The "Blended Smoothie" Problem (Confused Variables)
- The Issue: Sometimes a paper says, "Our new model is better!" But they also changed the data they trained it on.
- The Analogy: It's like a baker saying, "My new cake is delicious because I changed the oven temperature!" But they also secretly added a secret ingredient (better data) that made it taste good. You don't know if the oven or the secret ingredient was the real hero.
- The Reality: Papers rarely separate "better architecture" from "better data," so we don't know what actually caused the improvement.
The Solution: Six Rules for a Fair Game
The authors aren't saying the models are useless; they are saying the reporting is broken. They propose six simple rules to fix the marketplace:
- Share the Recipe (Release Weights): If you claim your model is reusable, you must share the actual model files with a clear license. If you can't (due to privacy or laws), say why clearly.
- Play on the Same Field (Shared Benchmarks): Authors must test their models on a small, shared set of standard datasets (like the top 10 most popular ones) so everyone is playing the same game.
- Label Your Numbers (Copied vs. Rerun): If you use a score from another paper, label it "Copied." If you ran the test yourself, label it "Rerun" and show your settings. Don't mix them up.
- Show the Variance (Report Uncertainty): If you run a test once, say so. If you run it five times, show the average and the range. Don't pretend a single lucky run is a guaranteed fact.
- Build a Universal Tester (Shared Harness): The community needs one single, automated tool (like a universal referee) that everyone uses to run tests. This ensures the "rules" (like how to measure accuracy) are identical for everyone.
- Isolate the Changes (Control Experiments): If you change the model and the data, you must run a test where you keep the data the same to prove the model change actually helped.
The Bottom Line
The paper concludes that no one knows the current "State of the Art" because the evidence is too messy to compare. It's a coordination failure, not a failure of individual scientists.
If the community follows these six rules, we can move from a chaotic bazaar of conflicting claims to a clear, trustworthy leaderboard. Then, when a farmer needs a flood map or a city planner needs a building detector, they can look at the leaderboard and actually know which tool is the best, rather than guessing based on confusing numbers.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.