EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports
The paper introduces EgoArgus, a human-annotated benchmark dataset designed to evaluate the ability of Vision-Language Models to act as reliable egocentric assistants by arbitrating between visual evidence and user dialogue in daily scenarios, revealing current limitations in modality trust assessment and the restricted effectiveness of existing bias mitigation methods.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine a pair of glasses that can see what you see, hear what you say, and offer helpful advice in real time. This is the promise of a new generation of artificial intelligence known as vision-language models. These systems are designed to act as daily assistants, watching the world through a camera mounted on your head or chest while listening to your conversation. They are meant to understand your situation, predict what you might do next, and step in with a warning or a suggestion when necessary. But for these assistants to be truly reliable, they must solve a difficult puzzle: when the video they see and the words you speak tell different stories, which one should they trust?
Researchers at National Yang Ming Chiao Tung University in Taiwan have built a new test to see if current artificial intelligence can handle this conflict. They created a dataset called EgoArgus, which simulates five different types of daily interactions between a user and an assistant. In some scenarios, the video and the conversation agree, working together to paint a clear picture. In others, the conversation might be completely unrelated to what is happening on screen, or it might be politely misleading. The most challenging cases occur when the user says one thing while the camera clearly shows something else, such as a person claiming to have turned off a faucet while the water is still running. The researchers wanted to see if the AI could ignore the misleading words and rely on the visual evidence, or if it would get confused and follow the wrong path.
To test this, the team gathered thousands of examples from real-life videos of people doing daily tasks, like cooking, shopping, or driving. They paired these videos with dialogue that fit into their five scenarios. They also created synthetic episodes where an AI assistant had to decide whether to speak up or stay silent. The results were revealing. When the user's words contradicted the video, the artificial intelligence models failed dramatically. Instead of trusting the camera, which showed the truth, the models blindly followed the user's incorrect statements. In these contradictory situations, the models performed worse than if they had simply guessed at random. They would confirm the user's mistake rather than correcting it, or they would suggest actions that made no sense given what was actually happening on screen.
The study also examined whether the AI could decide when to intervene. A good assistant knows when to speak and when to stay quiet. The researchers found that even the strongest models struggled with this balance. Some models were too eager, offering warnings when nothing was wrong, while others missed genuine safety hazards entirely. When they did intervene, their timing was often off, either acting too early before there was enough evidence or too late after the problem had already occurred. The team tried several methods to fix these errors, such as forcing the models to pay more attention to the video or training them to prefer visual evidence over text. These attempts provided only limited help. The models remained stubbornly biased toward the text, suggesting that simply adjusting their attention is not enough to make them reliable.
Further investigation into how these models think revealed that the problem runs deep. The researchers looked inside the models to see how they processed the video and the words. They found that the models kept the two types of information mixed together for a long time. It was only in the final stages of their processing that the models began to separate the visual evidence from the spoken words. By that point, however, the decision had often already been swayed by the misleading dialogue. This suggests that the models do not have a robust mechanism to weigh conflicting evidence in real time. They can distinguish between a picture and a sentence, but they cannot always decide which one is the truth when the two disagree.
The findings indicate that while these artificial intelligence systems are becoming better at understanding the world, they are not yet ready to be trusted as independent daily assistants. They lack the critical judgment needed to navigate a noisy environment where human speech might be mistaken, irrelevant, or even deceptive. For these assistants to become truly useful, they need to learn not just to see and hear, but to know when to trust their eyes over their ears. Until they can do that, they remain prone to following the wrong lead, even when the answer is right in front of them.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.