ThermEval: A Structured Benchmark for Evaluation of Vision-Language Models on Thermal Imagery
This paper introduces ThermEval, a comprehensive benchmark comprising 55,000 visual question answering pairs and a novel dataset with dense temperature maps, to demonstrate that current vision-language models significantly struggle with thermal imagery reasoning due to their reliance on RGB-centric priors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a brilliant student who has read every book in the library and watched every movie ever made. This student is an expert at describing what they see: "That's a red apple," "That's a sunny beach," "That's a happy dog." They are a master of RGB vision—seeing the world through the lens of color, texture, and light, just like our human eyes.
But now, imagine you hand this student a pair of night-vision goggles (thermal cameras). Suddenly, the world looks completely different. There are no red apples or blue skies. Instead, everything is a glowing map of heat. A hot engine looks like a blazing sun; a cold stone looks like a frozen moon.
This paper, ThermEval, is essentially a report card for our brilliant student when they try to use those night-vision goggles. The authors found out that while the student is great at describing pictures, they are terrible at understanding heat.
The Problem: The "Colorblind" AI
Current AI models (Vision-Language Models or VLMs) are trained mostly on photos of the world. They know what a "person" looks like in daylight. But thermal images don't show a person; they show a heat signature.
The authors asked a simple question: If you show an AI a thermal image of a person, can it tell you how hot their forehead is? Can it tell if one person is hotter than another?
The answer, unfortunately, was a resounding "No."
The Solution: ThermEval (The "Heat Test")
To prove this, the researchers built a giant exam called ThermEval. Think of it as a specialized driving test, but instead of driving a car, the AI has to "drive" through the world of heat.
They created 55,000 questions covering seven different skills:
- The "Is it a photo?" Test: Can the AI tell the difference between a normal photo and a heat map? (Most could).
- The "Color Shift" Test: Thermal cameras often use different color palettes (like "Magma" or "Viridis") to show heat. If you change the colors, does the AI get confused? (Yes, many did).
- The "Crowd Count" Test: How many people are in this heat map? (A bit shaky).
- The "Thermometer" Test: Can the AI read the little color bar on the side of the image that acts like a thermometer? (Many failed miserably, hallucinating numbers like "335 degrees" instead of "33.5").
- The "Hot vs. Cold" Test: Who is hotter, the person on the left or the right? (The AI often guessed based on what it thought humans usually feel like, rather than looking at the image).
- The "Exact Temperature" Test: What is the exact temperature of this nose? (The AI often just guessed "37°C" because that's the average human body temperature, ignoring the actual image).
- The "Distance" Test: Does the AI know that heat looks different when you are far away? (Most failed).
The Results: The AI is "Hallucinating" Heat
The results were surprising. Even the biggest, most powerful AI models (some with hundreds of billions of parameters) failed these tests.
- The "Book Smarts" Trap: The AI wasn't looking at the heat map. It was guessing based on what it knew from its training data. If asked, "How hot is a human forehead?", it would say "37°C" because that's what it learned in a textbook, even if the image showed a feverish 40°C or a cold 30°C. It was ignoring the visual evidence and relying on language priors (what it expects to be true).
- The "Colorbar" Confusion: Many models couldn't even read the thermometer scale on the side of the image. It's like giving someone a map with a legend but asking them to ignore the legend and guess the terrain.
- The "Fine-Tuning" Fix: The researchers tried teaching the AI specifically how to read heat maps (Supervised Fine-Tuning). This helped a lot! The AI got much better, almost matching human performance. But it still wasn't perfect. It's like teaching a student to pass a specific test, but they still don't truly understand the physics of heat.
Why Does This Matter?
You might think, "So what? AI can't read heat maps yet." But this is a big deal for real life. Thermal cameras are used for:
- Search and Rescue: Finding people lost in the woods at night.
- Medical Screening: Detecting fevers without touching people.
- Self-Driving Cars: Seeing pedestrians in total darkness.
- Industrial Safety: Finding overheating electrical wires before they catch fire.
If an AI driving a car in the dark can't tell the difference between a warm rock and a warm human, or if a medical AI can't accurately measure a fever, people could get hurt.
The Takeaway
The paper concludes that current AI is like a super-smart artist who has never felt the sun. They can describe a picture of the sun perfectly, but they don't understand what heat feels like.
To fix this, we can't just make the AI bigger or smarter. We need to retrain them from the ground up to understand physical signals like heat, not just pretty colors. The ThermEval benchmark is the first step in creating a new generation of AI that can truly "see" the invisible world of temperature.
In short: We built a test to see if AI can read a thermometer. They failed. Now we know exactly what we need to teach them to save lives in the dark.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.