← Latest papers
🤖 machine learning

HorusEye: Language as Dynamic Attention for Emergency Visual Analysis

The paper introduces HorusEye, a framework evaluating VLMs on the RefCOCO-Degraded dataset across five stages of emergency visual analysis, revealing that language feedback effectiveness is highly model-dependent and that standard degradation mitigation strategies like cropping can catastrophically fail in thermal imagery.

Original authors: Armel Yara

Published 2026-06-16
📖 4 min read☕ Coffee break read

Original authors: Armel Yara

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are a search-and-rescue team trying to find a lost person in a chaotic, dangerous environment. Sometimes it's foggy, sometimes there's thick smoke, and sometimes you have to rely on night-vision thermal cameras that see heat instead of color.

This paper, titled "HorusEye," asks a simple but crucial question: If we talk to our AI cameras and give them verbal instructions like "Look over there!" or "Check the left side," will that help them find the person better when the view is blurry or weird?

The researchers call this idea "Language as Dynamic Attention." Think of it like a human rescuer shouting to a partner: "Don't look at the tree, look at the bush behind it!" The paper tests if AI can do the same thing.

Here is the breakdown of their experiment in everyday terms:

1. The Setup: The "Gloomy Gym"

The researchers took a standard dataset of photos (people in clear, sunny conditions) and artificially messed them up to simulate emergencies. They created 15,244 images in four states:

  • Clean: A normal, clear photo.
  • Fog: Like a thick morning mist.
  • Smoke: Like a fire scene with gray haze.
  • Thermal: A black-and-white heat map (like a night-vision camera).

They tested five different AI "brains" (Vision-Language Models) to see how well they could point to a specific person in these photos.

2. The Big Surprise: Not All AI Brains Are the Same

The most important finding is that giving verbal instructions works for some AIs but hurts others. It's not a universal superpower.

  • The Star Performer (Gemini): When the view was terrible (especially in Thermal/heat vision), giving Gemini verbal nudges like "Look more to the left" was like magic. Its accuracy jumped by 47%. It successfully used the words to "re-focus" its eyes.
  • The Confused Student (Qwen2-VL): When given the exact same verbal nudges in the same terrible conditions, this AI actually got worse (dropped 5%). The instructions seemed to confuse it rather than help it.

The Lesson: You can't assume that talking to an AI will fix its vision. It depends entirely on which AI you are talking to.

3. The "Thermal Paradox": Zooming In Can Be Dangerous

The researchers also tested a common trick: Cropping.
Usually, if you are trying to find a person, you zoom in (crop the image) to get a closer look. This usually helps.

  • In normal photos: Zooming in helped the AI see better.
  • In Thermal (heat) photos: Zooming in was a disaster. Accuracy dropped by 26%.

The Analogy: Imagine trying to guess if someone is sitting or standing in a dark room using only a heat camera. If you zoom in only on their body, you can't see the floor or the chair they are sitting on. You lose the context. In thermal images, the background clues are essential to figure out what the person is doing. Zooming in removes those clues, causing the AI to fail catastrophically.

4. The "Dangerous Liar" (BLIP-2)

The study found one AI that is particularly risky for emergency use: BLIP-2.
When the images got blurry or smoky, this AI didn't just get confused; it started hallucinating (making things up) more often.

  • It would confidently describe objects that weren't there.
  • It would never say, "I'm not sure," even when the picture was terrible.

The researchers call this "unsafe." In an emergency, an AI that confidently lies is more dangerous than an AI that admits it can't see.

Summary of the "HorusEye" Findings

  • Talking to AI helps, but only if the AI is smart enough to listen. (Gemini listened; Qwen2 got confused).
  • Zooming in is a trap for heat-vision cameras. You need to see the whole scene to understand what's happening in the dark.
  • Some AIs are liars. BLIP-2 makes up facts when things get messy, making it a bad choice for rescue missions.

The Bottom Line: If you are building a rescue system, you can't just pick any AI and hope verbal instructions will save the day. You have to test specifically how that AI reacts to fog, smoke, and heat, and you must be careful not to zoom in too much when using thermal cameras.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →