MGAL: A Multilingual Granularity-Aware Long-Context Benchmark
The paper introduces MGAL, the first multilingual, granularity- and position-aware long-context benchmark constructed from UN reports, which reveals that current large language models struggle with coarse-grained tasks, exhibit performance gaps in lower-resource languages, and face new challenges related to local semantic crowding and fluency-consistency trade-offs.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine trying to read a thousand-page book, but instead of turning pages, you are forced to hold the entire text in your mind at once, searching for a single fact hidden somewhere in the middle. This is the challenge facing modern artificial intelligence systems known as large language models. These programs have become remarkably skilled at understanding human language, capable of writing essays, answering questions, and translating text with ease. However, their ability to process very long documents—spanning tens of thousands of words—remains a significant hurdle. While they can often recall a fact from the beginning or the end of a text, they frequently stumble when the information they need is buried deep within the middle, or when they must understand how different parts of a story connect to form a coherent whole. Researchers have long suspected that these systems struggle with the "long-context" problem, but until now, there has been no comprehensive way to test exactly where and why they fail, especially when dealing with languages other than English.
A team of researchers has now built a new testing ground called MGAL to solve this mystery. They constructed this benchmark using real-world documents from the United Nations, specifically reports that range from 8,000 to 128,000 tokens in length. These documents cover critical global topics like human rights, peacekeeping, and public health, and they are available in six official languages: English, Chinese, French, Russian, Spanish, and Arabic. The researchers did not just create a simple quiz; they designed a sophisticated set of tasks that test the models at four different levels of detail. At the finest level, the models must find specific words or numbers. At the next level, they must understand how a single sentence fits into a paragraph. Then, they must generate missing paragraphs that make sense in context, and finally, they must summarize or translate entire documents. Crucially, the researchers also tracked exactly where the information was located within the text—whether it appeared at the beginning, the middle, or the end—allowing them to see if the models had a bias toward certain parts of the document.
When the researchers tested twelve of the most advanced language models against this new benchmark, the results revealed a clear and consistent pattern. The models performed exceptionally well when the task required finding a specific word or number, showing they can act like efficient search engines for fine details. However, as the tasks became broader and required understanding larger chunks of text, such as summarizing a whole report or filling in a missing paragraph, their performance dropped significantly. The models struggled to maintain a coherent understanding of the document as a whole. Furthermore, the study highlighted a distinct gap between high-resource languages, like English and Chinese, and lower-resource languages. While the largest, most expensive models performed similarly across languages, smaller, open-source models showed a clear advantage for the major languages and struggled considerably with the others.
Perhaps the most revealing findings came from analyzing how the models failed. The researchers discovered that when sentences in a text shared similar topics or repeated the same names, the models tended to get confused. Instead of understanding the logical role a sentence played—whether it was providing background, offering a conclusion, or presenting a counter-argument—the models relied on shallow clues. They would pick the sentence that sounded most similar to the surrounding text or contained the same connecting words, even if it was the wrong choice. This is similar to a student who, when asked to summarize a story, simply repeats the names of the characters they remember most often rather than explaining what actually happened. Additionally, the study found that the models often produced text that sounded fluent and grammatically perfect but was factually wrong. They would confidently invent details, dates, or people that were not in the original document, creating a smooth-sounding narrative that drifted away from the truth.
The researchers also confirmed that the position of information matters deeply. In many tasks, the models were much better at finding answers located at the very beginning or the very end of a document, while performance dipped noticeably when the answer was in the middle. This suggests that the models have a hard time keeping their attention focused on the center of a long sequence of text. By using this new benchmark, the team has provided a clear map of the current limitations of artificial intelligence in long-document understanding. They have shown that while these systems are powerful tools for processing information, they still lack the deep, consistent comprehension required to truly "read" and understand complex, lengthy reports across different languages. This work sets a new standard for future research, guiding developers to build models that can look beyond surface-level clues and maintain a faithful, accurate understanding of the entire story, no matter how long it is.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.