BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
This paper introduces BEAR-Bench, a bilingual (English-Russian) benchmark of 1,000 human-annotated questions based on text-rich enterprise and academic documents to evaluate the complex reasoning capabilities and hallucination detection reliability of multimodal large language models, revealing significant performance gaps even among state-of-the-art systems.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the modern world, computers have become remarkably good at looking at pictures and reading the words inside them. If you show a machine a photo of a street sign, it can tell you what the sign says. If you show it a graph from a news report, it can often describe the trend the line is making. This ability, known as multimodal understanding, combines vision and language into a single skill. For a long time, the challenge was simply getting the machine to read the text clearly. But the next, harder step is reasoning. This is the ability to look at a complex document, find several different pieces of information scattered across a page, and connect them to answer a question that isn't written down anywhere. Imagine a financial report filled with tables, charts, and paragraphs of text. A human expert can look at a number in a table, compare it to a trend in a chart, and then read a sentence in the text to understand why those numbers changed. Teaching a computer to do this same kind of thinking, especially when the document is dense with information and written in languages other than English, has remained a difficult frontier.
A team of researchers from Yandex and the Applied AI Institute has taken a significant step toward solving this problem by creating a new test designed specifically for this kind of complex thinking. They call their creation BEAR-Bench. Unlike previous tests that focused on simple reading or required the computer to know outside facts, this new benchmark asks machines to solve problems using only the information present on a single page of a document. The test is bilingual, featuring both English and Russian, a choice that highlights a gap in current technology where machines often struggle with languages other than English or Chinese. The researchers gathered one thousand real-world documents, ranging from corporate financial reports and investor presentations to academic scientific papers. These pages are not simple; they are packed with text, tables, charts, diagrams, and mathematical formulas.
To build this test, the researchers did not just ask computers to generate questions. They hired thirteen human experts, each with a degree in a technical field, to write the questions and answers by hand. For every single page image, an expert had to create a question that required multiple steps to solve. For example, a question might ask the computer to find a specific value in a table, look at a related chart to see how that value changed over time, and then read a sentence in the text to explain the reason for that change. The computer was not allowed to use any outside knowledge; it had to find the answer entirely within the image provided. This ensured that the test measured the machine's ability to reason with the document in front of it, rather than its ability to memorize facts from the internet. The researchers also carefully filtered the documents to ensure they contained enough visual complexity, such as diagrams and equations, to make the task challenging.
When the researchers put sixteen different artificial intelligence models through this test, the results revealed a clear picture of where the technology stands today. The best-performing models, which are the most advanced systems currently available, managed to answer about seventy-five percent of the questions correctly. While this might sound impressive, it also means that even the smartest machines are still making mistakes on roughly one out of every four questions. The study found that these models perform noticeably worse when the documents are in Russian compared to English, confirming that language remains a significant hurdle. Furthermore, the type of document mattered; models tended to struggle more with business reports than with scientific papers, and they had particular trouble when the questions required them to count items or combine information from different parts of a page.
The researchers also looked closely at why the machines failed. They found that the most common errors were not due to a lack of logic, but rather to simple perception problems. The computers often misread a number in a table, failed to notice a specific label on a chart, or got confused about where an arrow was pointing in a diagram. Once the machine misread a single visual detail, the rest of its reasoning would fall apart. This suggests that before these systems can become true experts at analyzing documents, they need to get much better at seeing the small details clearly.
Beyond just testing the models, the researchers used the results to see if there is a way to tell when a machine is making a mistake. In many real-world situations, it is not enough to know that a computer might be wrong; you need a way to detect that error before it causes harm. The team tested several different methods for spotting these mistakes. They found that for the closed, proprietary systems where they could not see the internal workings, the best way to detect an error was to use another powerful AI model to act as a judge and check the answer. For the open systems where they could look inside the code, they found that checking the computer's own internal confidence signals worked well for short answers, but struggled when the computer produced very long, detailed explanations. The study concludes that while we are making progress, we still do not have a perfect way to catch every error, and the best systems still leave plenty of room for improvement.
This work matters because it moves the conversation beyond simple reading. It shows that while machines can read text, they are still learning how to think with it. By creating a test that is hard, bilingual, and focused on real documents, the researchers have provided a clear map of where the technology is strong and where it is weak. They have shown that the path forward requires not just bigger models, but better ways to handle the messy, detailed reality of professional documents in multiple languages. Until machines can reliably read a chart, cross-reference it with a table, and explain the connection without getting lost in the details, their use in high-stakes fields like finance and science will require careful human oversight.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.