From Plots to Words: Model-Aware Multimodal Explanations as a Foundation for Accessible, Non-Visual Interaction
This paper introduces a context-aware, multi-agent framework that converts visual time-series forecasting outputs into structured, model-aware textual explanations to enhance accessibility for blind and visually impaired users, demonstrating significant improvements in explanation quality and trustworthiness over numerical baselines in an initial LLM-based evaluation.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern digital world, vast amounts of information flow through computer systems every second, often visualized as complex charts and graphs. For people who can see, these images are a quick way to understand trends, spot problems, and predict what might happen next. However, for the millions of people who are blind or have low vision, these visual interfaces are often impenetrable walls. Screen readers, the tools that convert text into speech, struggle to convey the meaning of a curved line or a shaded area on a graph. This creates a significant gap in accessibility, leaving many users unable to interact with data-intensive systems that rely heavily on visual dashboards. To bridge this gap, researchers are turning to artificial intelligence, specifically large language models, which are powerful computer programs capable of understanding and generating human language. The challenge lies in teaching these models to not just describe what a chart looks like, but to explain why a prediction was made, using the internal logic of the computer system itself rather than just the visual pattern.
A team of researchers at the Institute of Theoretical and Applied Informatics in Poland has developed a new system designed to solve this problem by turning visual forecasts into clear, structured words. Their work focuses on the management of cloud computing infrastructure, where operators must constantly monitor massive amounts of data to predict future resource needs. In these environments, decisions are often made based on trend lines and uncertainty bands that are invisible to non-visual users. The researchers built a multi-agent framework, a system where several specialized artificial intelligence programs work together to process data, generate predictions, and then translate those predictions into detailed textual explanations. Instead of showing a user a graph, the system speaks to them, describing not only the numbers but also the specific reasons behind the forecast, including how confident the computer is in its own prediction.
The core of this research is a method that moves beyond simple description to true explanation. The team tested three different ways of generating responses to user questions about data. The first method, a baseline, relied only on raw numbers. The second method added visual descriptions, allowing the system to "see" the chart and describe the trends it observed. The third and most advanced method, which the researchers call "explainable," went a step further. It integrated signals from the computer model itself, such as which specific data points were most important for the prediction and how uncertain the model felt about the future. This approach ensures that the explanation is grounded in the actual decision-making process of the machine, rather than just a superficial description of the output. The system was tested using real-world data from a large cluster of graphics processing units, tracking hundreds of thousands of job submissions over several months to simulate a complex, high-stakes environment.
When the researchers compared the three methods, the results showed a clear progression in quality. The system that simply listed numbers was often vague and lacked depth. Adding visual descriptions helped the system provide more helpful and insightful answers, but the most significant improvement came from the method that included the model's internal reasoning. This advanced version produced explanations that were not only more trustworthy but also significantly more aware of the model's own limitations and confidence levels. In a rigorous evaluation, this explainable approach improved the overall quality of the responses by up to 32 percent compared to the basic numerical baseline. It was particularly effective at reducing errors where the system might invent facts, a problem known as hallucination, and at providing a clearer sense of uncertainty, which is crucial for making safe decisions.
The researchers emphasize that while their system creates a powerful foundation for accessible interaction, it has not yet been tested directly with blind or low-vision users. Their work serves as a proof of concept, demonstrating that it is possible to convert complex, visual forecasting outputs into reliable, model-aware text that can be consumed through speech or screen readers. By shifting the focus from showing a chart to explaining the logic behind the numbers, the team has created a pathway for non-visual users to engage with data-intensive systems in a way that was previously impossible. The study suggests that for artificial intelligence to be truly accessible, it must do more than translate images into words; it must translate the underlying reasoning of the machine into a narrative that any user can understand and trust. This approach positions language not just as a way to describe data, but as the primary interface for interacting with it, ensuring that the insights of modern computing are available to everyone, regardless of their ability to see.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.