← Latest papers
🤖 machine learning

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

This paper introduces the "Perception-Physics Paradox" to highlight how Vision Foundation Models often rely on visual correlations rather than physical invariants in scientific domains, proposing "scientific alignment" via structural isomorphism and the new TC-Bench dataset to demonstrate that current models fail to reason correctly in extreme regimes despite high predictive performance.

Original authors: Dingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller, Francesco Locatello

Published 2026-05-26
📖 5 min read🧠 Deep dive

Original authors: Dingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller, Francesco Locatello

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: Looking Good vs. Knowing the Truth

Imagine you have a student who is taking a test about the weather. The student is incredibly good at recognizing pictures of clouds. If you show them a photo of a fluffy white cloud, they say, "That's a nice day." If you show them a dark, swirling storm, they say, "That's a storm." They get a perfect score on the test.

But here is the catch: The student doesn't actually understand the physics of the storm. They are just memorizing that "dark swirl = storm."

This paper argues that modern AI models (specifically "Vision Foundation Models") are like this student. They are amazing at recognizing what things look like (perception), but they often fail to understand the underlying rules that make those things happen (physics). The authors call this the Perception–Physics Paradox.

The Problem: The "Visual Saturation" Trap

The researchers studied Tropical Cyclones (hurricanes/typhoons). These storms have a specific physical rule: the lower the air pressure in the center, the stronger the wind.

  • Moderate Storms: A storm with a pressure of 1000 hPa looks different from one with 980 hPa. The AI can easily tell them apart.
  • Intense Storms: When a storm gets extremely strong (pressure drops below 980 hPa), the clouds form a perfect, tight "eye." At this point, a storm with a pressure of 915 hPa looks almost identical to a storm with a pressure of 905 hPa.

The Analogy: Imagine two cars driving at 100 mph and 101 mph. To a camera, they look exactly the same. If your AI is just looking at the picture, it can't tell the difference. It might guess "100 mph" for both.

The paper found that while these AI models look smart and get good scores on average, they collapse when the storms get really intense. They can't distinguish between a "very strong" storm and a "super strong" storm because the visual clues have run out. The AI is essentially guessing based on the average, not the actual physics.

The Solution: TC-BENCH (The Stress Test)

To prove this, the authors built a new tool called TC-BENCH.

  • What it is: A massive, global database of hurricane satellite images paired with real scientific data (like exact pressure and wind speed).
  • Why it's special: Previous tests mostly looked at "average" storms. TC-BENCH is designed to specifically stress-test the AI with the most extreme, dangerous storms where the visual clues are tricky.
  • The Pipeline: It's like a factory that automatically gathers, cleans, and organizes hurricane data from all over the world so scientists can run the same tests again and again without bias.

The "Scientific Alignment" Test

The authors didn't just ask, "Can the AI guess the wind speed?" They asked a deeper question: "Does the AI's internal brain (its 'latent space') actually map to the laws of physics?"

They introduced a concept called Scientific Alignment. Think of it like a map:

  • Good Map: If you move one inch North on the map, you move one inch North in the real world. The structure matches.
  • Bad Map: If you move North on the map, sometimes you go North, sometimes East, and sometimes you stay still. The map looks okay from a distance, but it's useless for navigation.

They used three specific "probes" (tests) to check the map:

  1. Static Fidelity (The Snapshot Test): Can the AI look at a single photo and accurately figure out the pressure?
    • Result: In strong storms, the AI's internal map gets "blurry." It can't tell the difference between 915 hPa and 905 hPa.
  2. Dynamic Coherence (The Movie Test): If the storm changes over time, does the AI's internal map change in a way that matches reality?
    • Result: No. The AI's internal representation of time gets "stuck" or collapses when the storm gets intense.
  3. Manifold Consistency (The Logic Test): Does the AI understand that storms near the equator need stronger winds to have the same pressure as storms near the poles? (This is a real physics rule).
    • Result: The AI fails this logic test in extreme storms. It forgets the rules of the game.

The Conclusion: "Looking Right" Isn't "Being Right"

The paper concludes that scaling up AI models (making them bigger) does not automatically make them understand science.

Even the smartest, most powerful AI models currently available are relying on "visual shortcuts." They are great at recognizing patterns in normal weather, but when the weather gets extreme and the pictures look similar, their internal understanding of physics breaks down.

The Takeaway:
If you want to use AI for life-or-death scientific decisions (like predicting a hurricane's path), you can't just trust that it "looks" smart. You need to test if its internal logic actually matches the laws of physics. The authors provide the tools (TC-BENCH and the probing framework) to do exactly that, showing that right now, these models are not yet "scientifically aligned."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →