← Latest papers
🤖 AI

MOMO: Mars Orbital Model Foundation Model for Mars Orbital Applications

The paper introduces MOMO, the first multi-sensor foundation model for Mars remote sensing, which leverages a novel Equal Validation Loss strategy to merge representations from HiRISE, CTX, and THEMIS sensors, achieving superior performance across nine downstream tasks compared to existing baselines.

Original authors: Mirali Purohit, Bimal Gajera, Irish Mehta, Bhanu Tokas, Jacob Adler, Steven Lu, Scott Dickenshied, Serina Diniega, Brian Bue, Umaa Rebbapragada, Hannah Kerner

Published 2026-04-06
📖 5 min read🧠 Deep dive

Original authors: Mirali Purohit, Bimal Gajera, Irish Mehta, Bhanu Tokas, Jacob Adler, Steven Lu, Scott Dickenshied, Serina Diniega, Brian Bue, Umaa Rebbapragada, Hannah Kerner

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a robot to understand the surface of Mars. The problem is that we don't have just one pair of eyes looking at the Red Planet; we have three very different cameras, each seeing the world in a completely different way.

  • HiRISE is like a super-powerful magnifying glass. It sees tiny rocks and cracks from space (0.25 meters per pixel), but it can only look at a tiny speck of the planet at a time.
  • CTX is like a standard drone camera. It sees the whole landscape clearly (5 meters per pixel) and covers almost the entire planet.
  • THEMIS is like a thermal night-vision goggles. It sees heat patterns from very far away (100 meters per pixel) and covers the whole globe, but the details are blurry.

For a long time, scientists had to build three separate "brains" (AI models) to understand these three cameras. If they wanted to study a landslide using the thermal camera, they had to use one brain. If they wanted to find a tiny boulder, they had to use another. It was inefficient and confusing.

Enter MOMO (Mars Orbital Model). Think of MOMO as the ultimate "Swiss Army Knife" AI for Mars. It's the first "Foundation Model" (a super-smart, pre-trained brain) designed specifically to understand all these different Martian cameras at once.

How Did They Build It? (The "Model Merging" Recipe)

Usually, to train an AI on three different cameras, you'd try to feed all the pictures into one giant pot and hope it learns everything. But that's like trying to cook a steak, a soup, and a cake in the same pan at the same time—it gets messy.

Instead, the researchers used a clever two-step recipe:

  1. Train Separately: They first trained three separate, specialized chefs. One chef learned only from the magnifying glass (HiRISE), one from the drone (CTX), and one from the thermal goggles (THEMIS). Each became an expert in their own domain.
  2. The "Equal Validation Loss" (EVL) Merge: This is the paper's secret sauce. Imagine you have three chefs who are cooking their own dishes. You want to combine them into one master chef, but you can't just grab them at random times.
    • If you grab Chef A when they are just starting (undercooked) and Chef B when they are burnt (overcooked), the result will be terrible.
    • The researchers invented a strategy called Equal Validation Loss (EVL). Think of this as a perfectly synchronized dance. They watched the three chefs and waited until they all reached the exact same level of "mastery" (measured by how well they were doing on a test).
    • Once they were all at that perfect "sweet spot," they merged their knowledge into one single model.

Why Is This a Big Deal?

1. It's a "One-Size-Fits-All" Solution
Before MOMO, if a scientist wanted to study a crater, they had to know which camera took the picture and pick the right AI model. Now, they just give the picture to MOMO. Whether it's a high-res photo of a single rock or a low-res thermal map of a continent, MOMO understands it.

2. It's Better Than "Earth" Models
Scientists often try to use AI models trained on Earth (like Google Maps or weather satellites) to study Mars. It's like trying to use a map of New York City to navigate the streets of Tokyo. The terrain, lighting, and rocks are totally different. MOMO was trained specifically on Mars data, so it understands Martian dust, craters, and ice much better than Earth-trained models.

3. It's Future-Proof
If NASA launches a new satellite with a brand-new camera next year, they don't have to retrain the whole system. They just train a small "expert" on that new camera and use the same "dance" (EVL strategy) to merge it into the existing MOMO brain. It's like adding a new tool to a toolbox without having to buy a whole new toolbox.

What Can MOMO Do?

The researchers tested MOMO on 9 different tasks, and it crushed them:

  • Finding Boulders: Locating tiny rocks for future rovers to avoid.
  • Mapping Craters: Counting and classifying craters to understand the planet's history.
  • Spotting Landslides: Identifying where the ground is moving.
  • Detecting Frost: Seeing where ice forms on the surface.

In almost every test, MOMO was more accurate than models trained from scratch, models trained on Earth data, or models trained on just one type of camera.

The Bottom Line

MOMO is like giving planetary scientists a universal translator for Mars. Instead of struggling to speak three different languages (HiRISE, CTX, and THEMIS), they now have one fluent speaker who understands the nuances of the entire Red Planet. This makes discovering new things on Mars faster, cheaper, and more accurate, paving the way for future human exploration.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →