PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding
This paper introduces PUMA, a novel 900-task benchmark designed to evaluate the capabilities of multimodal AI models in the Polish cultural and linguistic context, revealing significant performance gaps in audio and document understanding despite strong visual question-answering results.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
In the rapidly evolving field of artificial intelligence, computers are learning to see, hear, and read just as humans do. For years, these systems were trained primarily on text, learning to process words and sentences. Now, they are expanding their senses to include images, audio recordings, and complex documents. However, a significant gap remains in how well these machines understand the world beyond English. Just as a person can speak a language fluently yet miss the subtle jokes, historical references, or cultural traditions embedded within it, an artificial intelligence can process words correctly while failing to grasp their deeper meaning. Understanding a culture requires more than grammar; it demands familiarity with local history, geography, popular trends, and the unspoken rules of daily life. Without this cultural grounding, even the most advanced systems risk offering superficial or incorrect interpretations of the world they are trying to understand.
To address this blind spot, researchers at the National Information Processing Institute in Warsaw, Poland, have created a new testing ground called PUMA. This benchmark is not a single test but a collection of 900 carefully crafted challenges designed specifically to probe the limits of artificial intelligence within the Polish cultural and linguistic context. The team built these tasks by hand, drawing from real-world examples of Polish history, modern life, geography, and audio recordings. The goal was to see if machines could truly understand Poland, not just translate its words. The benchmark covers three distinct ways humans perceive the world: images, sound, and documents. In the image section, models are asked to identify historical landmarks, recognize contemporary pop culture figures, or distinguish between different types of Polish landscapes. The audio section tests whether a machine can transcribe speech from noisy recordings, understand the intent behind a spoken conversation, or even identify non-speech sounds like traditional folk instruments or specific environmental noises. The document section challenges the systems to read handwritten notes, extract data from complex tables, and interpret visually rich pages like menus or official forms.
When the researchers put dozens of the world's most advanced artificial intelligence systems through this rigorous test, the results revealed a stark divide between general capability and cultural fluency. The top-performing commercial models, particularly those from Google's Gemini family, achieved high scores, demonstrating a strong ability to handle visual questions and document analysis. These systems managed to score nearly 80 percent on the most difficult tasks, showing they can effectively navigate the visual and textual landscape of Poland. However, the study found that even these leading models struggle significantly when the task moves away from clear text and images into the realm of sound. Non-speech audio, such as identifying specific musical instruments or environmental sounds unique to the region, proved to be the most difficult category, with the best models scoring only around 56 percent. This suggests that while machines are becoming excellent at reading and seeing, their ability to "hear" and interpret the nuanced sounds of a culture remains underdeveloped.
The evaluation also highlighted a surprising weakness in how some systems handle the practical side of language. When asked to extract specific data from a document and format it perfectly, many models failed, often tripping over minor structural errors. Furthermore, the study noted that some of the most famous American models refused to answer questions about Polish historical figures or cultural works, treating them as sensitive topics when they were not. This refusal rate was much higher than in other model families, effectively lowering their scores on knowledge-based questions. In contrast, specialized models designed for specific tasks, such as converting speech to text or reading printed characters, often outperformed the massive, general-purpose systems in those narrow areas. This indicates that for specific, practical applications, smaller, focused tools may still be more reliable than the all-encompassing giants.
Ultimately, the PUMA benchmark serves as a mirror, reflecting the current state of artificial intelligence in a specific cultural context. It shows that while machines are making impressive strides in visual and textual understanding, they still lack the deep, intuitive grasp of local culture that comes from lived experience. The researchers found that progress in this field cannot be measured solely by how well a system performs on English-language tests or general knowledge quizzes. True understanding requires the ability to navigate the messy, rich, and often noisy reality of a specific culture, from the sound of a local dialect to the layout of a handwritten note. By opening their tools and data to the public, the team hopes to encourage the development of artificial intelligence that is not just globally capable, but locally grounded, ensuring that these powerful systems can serve and understand people in their own languages and cultural settings.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.