← Latest papers
💬 NLP

Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

The SHROOM-Visions 2026 shared task, hosted at the UncertaiNLP Workshop, successfully engaged 27 teams to develop model-agnostic systems for detecting and classifying fine-grained hallucinations in large vision-language models across four languages, achieving significant performance improvements over baselines using the SHEEP dataset.

Original authors: Raúl Vázquez, Aman Sinha, Chuyuan Li, Claudio Savelli, Eduardo Calò, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, Jörg Tiedemann, Timothee Mickus

Published 2026-08-27
📖 4 min read☕ Coffee break read

Original authors: Raúl Vázquez, Aman Sinha, Chuyuan Li, Claudio Savelli, Eduardo Calò, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, Jörg Tiedemann, Timothee Mickus

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the rapidly evolving world of artificial intelligence, a new generation of computers has learned to see and speak simultaneously. These systems, known as large vision-language models, can look at a photograph and describe it in fluent sentences, or answer questions about what they see. They are remarkably capable, yet they suffer from a specific kind of error called a hallucination. This is not a visual trick like a mirage, but a verbal one: the computer produces text that sounds confident and grammatically perfect but is factually wrong or completely unrelated to the image it is looking at. It might invent an object that isn't there, describe a color incorrectly, or misread a sign. As these tools become more common in daily life, from helping doctors read scans to assisting travelers, the ability to spot these confident lies has become a critical safety issue.

To address this challenge, researchers organized a global competition called SHROOM-Visions 2026. The goal was not to build a better image-generator, but to create a better detector of errors. Twenty-seven teams from around the world participated, submitting hundreds of different computer systems designed to scan an image, a question, and a generated answer, then pinpoint exactly which words were lies and what kind of lie they were. The competition used a massive collection of twenty thousand examples across four languages: English, Chinese, French, and Italian. Crucially, the researchers did not rely solely on errors made by other computers, which can be biased or repetitive. Instead, they included thousands of examples where humans deliberately wrote false descriptions of images, ensuring the test data was robust and independent of any specific machine's quirks.

The results revealed that while computers have made significant progress, the task remains difficult. The best systems in the competition managed to identify and categorize these errors with a moderate level of accuracy, significantly outperforming basic baseline methods by a margin of thirty to forty points. However, even the top performers did not achieve perfect scores. The most successful systems could correctly identify the location of a hallucinated word about half the time and classify the type of error with similar reliability. This suggests that while we are getting better at teaching machines to police their own output, the problem of distinguishing truth from fiction in a visual context is far from solved.

A surprising finding emerged from the analysis of how these systems performed. The researchers discovered that the ranking of the teams was not as stable as a simple list of scores might suggest. When the same systems were tested on different subsets of the data, their relative positions shifted considerably. A team that appeared to be the clear leader in one group of tests might drop significantly in another. This instability indicates that the performance of these detectors is highly sensitive to how the test data is constructed. Some systems performed exceptionally well on data generated by other machines but struggled when faced with human-written errors, while others showed the opposite trend. This suggests that a system's success is not just about its raw intelligence, but also about how closely its training data matches the specific type of error it is being asked to find.

The competition also highlighted the complexity of the errors themselves. The researchers categorized hallucinations into five distinct types: inventing something that does not exist, misdescribing something that is visible, misreading text within an image, miscounting objects, and other miscellaneous errors. The systems found that some types of errors were easier to catch than others. For instance, counting mistakes were generally easier to detect than subtle misdescriptions of an object's properties. Furthermore, the difficulty varied by language, with English proving to be the most challenging language for the systems to handle, while Chinese saw slightly higher average performance.

Ultimately, the SHROOM-Visions 2026 task demonstrated that detecting hallucinations is a nuanced problem that requires more than just a single metric of success. The researchers found that a system's ability to locate an error and its ability to correctly label what kind of error it was are related but distinct skills. A system might be good at finding the wrong words but poor at explaining why they are wrong. The study concludes that for these tools to be truly reliable, future evaluations must account for this uncertainty and use diverse, human-informed data rather than relying on data generated by machines alone. The path forward involves building benchmarks that are resilient to the rapid changes in technology, ensuring that the detectors we build today remain effective as the models they monitor continue to evolve.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →