Now We Know? A Systematic Comparison of TerraMind and THOR
This paper presents a systematic comparison of two European Space Agency Geospatial Foundation Models, THOR and TerraMind, across ten diverse use cases to demonstrate that architectural design choices—specifically patch size and decoder complexity—explain more performance variance than model identity, revealing complementary investment strategies and establishing a diagnostic methodology for future GFM evaluation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are trying to teach a computer to understand the Earth from space. For years, scientists have been building "Geospatial Foundation Models" (GFMs). Think of these models as super-smart, pre-trained brains that have already looked at billions of satellite photos. They are like a student who has read every geography textbook in the library before ever taking a specific test. Because they've seen so much, they can learn new tasks—like finding floods or tracking icebergs—much faster than a model starting from scratch. But here's the tricky part: we have many of these "super-brains," and they all get different scores on tests. Sometimes one is better at finding floods, and another is better at spotting fires. The big question is: Why? Is one model just smarter because of its brain architecture? Is it because it has a better "decoder" (the part that translates brain signals into a map)? Or is it just lucky with the specific pictures it was tested on? Until now, we've mostly just looked at the final scores, like a sports leaderboard, without understanding the mechanics behind the win.
This paper, titled "Now We Know?", is like a detective story where two rival detectives, THOR and TerraMind, are put through a series of controlled experiments to see what actually makes them tick. Both were built by teams funded by the European Space Agency, but they have very different philosophies. THOR is the "chameleon": it can change its "patch size" (how it slices up an image) on the fly, trading computer power for sharper details when it needs to. TerraMind is the "generative artist": it was trained to imagine missing pieces of a puzzle, allowing it to fill in gaps or switch between different types of sensors (like turning a radar image into an optical one) using a trick called "Thinking-in-Modalities."
The researchers didn't just let them fight it out on a leaderboard. Instead, they ran a massive "ablation study," which is a fancy way of saying they took the models apart and reassembled them piece by piece to see which part mattered most. They tested these models on ten different real-world missions, from tracking methane leaks to mapping snow cover.
Here is what they found, and it's a bit of a plot twist. First, they discovered that the "leaderboard" score is often a liar. The biggest factor determining who wins isn't necessarily which model (THOR or TerraMind) you pick, but how you slice the image and how complex your decoder is. If you slice an image into tiny pieces (small patch size), THOR becomes incredibly powerful, especially for spotting small, tricky things like icebergs or flood zones, because it sees the fine details. However, this comes at a high cost: it requires a lot of computer power. TerraMind, on the other hand, uses a fixed slice size. It shines when you need to process optical images (like standard photos) efficiently, but it struggles a bit more with radar-only data unless you give it a very powerful decoder.
The study also ruled out a few common assumptions. For instance, many people thought that simply fusing data from two different satellites (like combining radar and optical) would always make the model smarter. The researchers found that, with full datasets, this wasn't always true; sometimes, just using the optical data alone was better because the fusion method they used was too simple to handle the extra noise. They also found that "freezing" the model's brain (not letting it learn new things during the test) worked surprisingly well for TerraMind on optical tasks, but THOR often needed to be fully "unfrozen" (allowed to learn) to reach its full potential, especially on radar tasks.
Ultimately, the paper suggests there is no single "winner." It's not about which model is the best in the world; it's about matching the tool to the job. If you have a powerful computer and need to find tiny, specific objects in radar images, THOR with a tiny patch size is your hero. If you need a fast, efficient model for general optical mapping, TerraMind might be the better choice. The authors conclude that the future of these models isn't just about building bigger brains, but about understanding the specific trade-offs between how we slice the data, how we decode the answers, and what kind of "brain" we are using. It's a reminder that in the world of AI, context is everything.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.