← Latest papers
🤖 AI

Mind the Gap: Quantifying the Domain Gap in Cross-Sensor Diffusion Super-Resolution

This paper presents the first systematic study demonstrating that current diffusion-based super-resolution models suffer from a significant domain gap when transitioning from synthetic training data to real cross-sensor satellite imagery, revealing that while synthetic training fails on real data, training on real data introduces optimization challenges that highlight the need to disentangle super-resolution from domain adaptation.

Original authors: Dawid Kopeć, Katarzyna Jabłońska, Wojciech Kozłowski, Maciej Zięba

Published 2026-06-29
📖 4 min read☕ Coffee break read

Original authors: Dawid Kopeć, Katarzyna Jabłońska, Wojciech Kozłowski, Maciej Zięba

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to teach a chef how to cook a perfect steak.

The Problem: The "Fake" vs. "Real" Kitchen
In the world of satellite images, we have two types of cameras:

  1. Sentinel-2: A free, wide-angle camera that takes clear but slightly blurry photos (like a photo taken from a high hill).
  2. PlanetScope: A paid, zoomed-in camera that takes incredibly sharp, detailed photos (like a photo taken from a drone).

Scientists want to use AI to turn the blurry Sentinel photos into sharp PlanetScope photos. This is called Super-Resolution.

However, there is a catch: We don't have a "perfect" pair of photos taken at the exact same second by both cameras. So, to train the AI, researchers usually take a sharp photo, artificially blur it with a computer filter (like smearing it with a finger), and tell the AI, "Here is the blurry version, now guess what the sharp version looks like."

The Experiment: The Taste Test
The authors of this paper asked a simple question: Does an AI chef trained on "computer-smudged" photos know how to cook a real, naturally blurry photo taken by a different camera?

They set up three scenarios:

  1. The Practice Run (Synthetic): The AI learns on computer-blurred photos and is tested on computer-blurred photos. (Easy mode).
  2. The Real-World Test (The Gap): The AI learns on computer-blurred photos but is tested on real, naturally blurry photos from the Sentinel satellite. (Hard mode).
  3. The Direct Challenge: The AI learns directly on real blurry photos and is tested on real sharp photos. (Expert mode).

The Results: A Shocking Failure
The findings were quite dramatic, like a chef who can perfectly replicate a fake steak but ruins a real one.

  • The Synthetic Trap: When the AI trained on "computer-smudged" photos was given a real satellite photo, it failed miserably. It didn't just fail to sharpen the image; it actually made it look worse. It tried to force the "computer-smudge" rules onto a real photo, creating weird, fake textures. It was like a chef trying to cook a real steak using instructions for a plastic toy steak.
  • The Real-World Struggle: When they tried to train the AI directly on real satellite photos, it didn't fail as badly, but it still couldn't match the perfection of the "Practice Run." The real world is too messy. The lighting changes, the clouds move, and the cameras see colors slightly differently. The AI got confused by this "noise" and couldn't learn a consistent recipe.

The Downstream Test: Does it Help the Firefighters?
To see if this mattered in real life, the researchers used the AI-enhanced photos to help a computer program detect burned forest areas (a task called segmentation).

  • The AI trained on "fake" data (Practice Run) actually helped the fire detection program work a little better. It added some useful detail.
  • The AI trained on "real" data (Direct Challenge) actually hurt the fire detection program. Because the AI was trying so hard to adapt to the messy real world, it started inventing fake details (hallucinations) that confused the fire detector.

The Big Takeaway
The paper concludes that we are currently stuck in a "Mind the Gap" situation.

  • If we train AI on fake, computer-generated data, it works consistently but can't handle the real world.
  • If we try to train it on real data, it gets confused by the messiness of reality and creates artifacts that break downstream tasks.

The authors suggest that the solution isn't just building a "smarter" AI chef. Instead, we need to teach the AI two separate skills: one to learn how to sharpen images, and another to learn how to translate between different camera types. Right now, we are asking the AI to do both at once, and it's failing at the combination.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →