Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation
This paper proposes a novel token-level visual reliability measure that combines feature sensitivity and counterfactual signals to effectively detect and diagnose hallucinations in sign language translation by quantifying the extent to which models rely on visual evidence rather than language priors.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
The Big Picture: The "Blind" Translator
Imagine a translator who is trying to translate a video of a person signing (using hand gestures) into spoken text.
- The Goal: The translator should look at the hands and say exactly what they are saying.
- The Problem: Sometimes, the translator gets lazy or overconfident. Instead of looking at the hands, they just guess what the person probably said based on how sentences usually go. They might produce a sentence that sounds perfect and grammatically correct, but it has nothing to do with what the signer actually did. In AI terms, this is called a hallucination.
This paper argues that the newer, "smarter" translators (which skip an intermediate step called "glosses") are actually worse at this. They are so good at guessing based on language patterns that they often stop looking at the video entirely.
The Core Idea: "Grounding" vs. "Guessing"
The authors introduce a new way to check if the AI is actually looking at the video or just making things up. They call this Reliability.
Think of it like a detective checking an alibi:
- Grounding (Looking at the evidence): The AI says, "I am choosing this word because I see the signer's hand moving in a specific way."
- Guessing (Relying on memory): The AI says, "I am choosing this word because, in English, people usually say 'hello' after 'good morning'."
The paper's main job is to build a "lie detector" that can tell the difference between these two behaviors.
How the "Lie Detector" Works
The authors created a test to see how much the AI relies on the video. They run the translation three times in parallel:
- The Real Deal: The AI sees the actual video.
- The Blank Screen: The AI sees the video, but the screen is black (no visual data).
- The Wrong Video: The AI sees a completely different video (like a cat instead of a signer).
The Test:
- If the AI changes its answer significantly when the video is removed or swapped, it means it was Grounded (it was actually looking at the video).
- If the AI keeps saying the exact same thing even when the video is gone or wrong, it means it was Guessing (it was just relying on its internal language habits).
They combine these tests into a single score called Reliability. A high score means the AI is paying attention to the video. A low score means it's hallucinating.
The "Gloss-Free" Trap
The paper compares two types of translators:
- Gloss-Based (The Old Way): These models use a middleman. They translate the video into "glosses" (simple labels for hand movements) first, then into text. It's like having a strict teacher who forces the student to point to the object before naming it. This keeps them honest.
- Gloss-Free (The New Way): These models jump straight from video to text. They are like a student who is told, "Just tell me what you see," without the strict labels.
The Finding: The "Gloss-Free" models are much more prone to hallucinations. Because they aren't forced to pause and label the movements, they tend to skip looking at the video and just "fill in the blanks" with what sounds good. The paper proves that these models are "guessing" far more often than the older models.
Why This Matters
Usually, when we check if an AI is lying, we look at how "confident" it sounds. If it says something with high confidence, we assume it's true.
- The Paper's Twist: In Sign Language Translation, confidence is a trap. A model can be 100% confident that it's raining, even if the video shows a sunny day, because it's just guessing based on the weather forecast it read earlier.
The authors show that their new Reliability Score is much better at catching these lies than just checking confidence. It works by asking: "Did you actually look at the video to decide this, or did you just guess?"
Summary of Results
- It Works: The new "Reliability" score successfully predicts when the AI is hallucinating.
- It's Universal: It works on different datasets and different types of AI models.
- It Explains the "Why": It proves that the newer, gloss-free models hallucinate more because they stop using the visual information and start relying too much on their language training.
- It's a Safety Tool: By combining this visual check with standard text checks, we can get a much clearer picture of when an AI is making things up.
In short, the paper teaches us that for sign language translation, you can't just trust the AI's confidence; you have to check if it's actually watching the video.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.