← Latest papers
💬 NLP

Multilingual Vision-Language Models, A Survey

This survey reviews 33 multilingual vision-language models and 23 benchmarks to highlight the critical tension between achieving language-neutral cross-lingual representations and maintaining cultural awareness, while identifying significant gaps between current training objectives and evaluation methodologies.

Original authors: Andrei-Alexandru Manea, Jindřich Libovický

Published 2026-05-14
📖 5 min read🧠 Deep dive

Original authors: Andrei-Alexandru Manea, Jindřich Libovický

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Picture: Teaching Robots to See and Speak Many Languages

Imagine you are trying to teach a robot to understand the world. You show it pictures and teach it to describe them or answer questions about them. This is what Vision-Language Models do. They are like a robot's eyes and brain combined, allowing it to "see" an image and "speak" about it.

For a long time, researchers only taught these robots English. It was like teaching a child to read only one language. But the world speaks many languages. This paper is a massive survey (a review of the field) that looks at how well these robots are learning to speak and understand many languages at once, from 2020 to 2025.

The authors, Andrei-Alexandru Manea and Jindřich Libovický, looked at 33 different robot models and 23 different tests (benchmarks) to see where we stand.

The Core Conflict: The "Universal Translator" vs. The "Local Guide"

The paper identifies a big tension, like a tug-of-war, between two goals:

  1. Language Neutrality (The Universal Translator):

    • The Idea: A picture of a dog is a dog, no matter if you call it dog, chien, or perro. The robot should see the same thing and give the same answer regardless of the language.
    • The Analogy: Think of this as a universal map. If you look at a map of the Eiffel Tower, it's the Eiffel Tower whether you are in Paris, Tokyo, or New York. The facts don't change.
    • The Problem: If the robot is too neutral, it might ignore the fact that in some cultures, a dog is a beloved family member, while in others, the word might be used as an insult.
  2. Cultural Awareness (The Local Guide):

    • The Idea: The robot needs to understand that context matters. A "family meal" in the US might mean turkey and pie, but in Japan, it might mean miso soup. The robot should adapt its answer to fit the culture.
    • The Analogy: Think of this as a local tour guide. A guide in Italy knows that "coffee" means a small espresso, while a guide in the US knows it might mean a huge mug of drip coffee. They know the local rules.
    • The Problem: It is very hard to teach a robot these subtle cultural rules. You can't just translate a sentence; you have to change the meaning based on who is listening.

The Paper's Finding: Currently, it is much easier to teach the robot to be a Universal Translator (Language Neutrality) than a Local Guide (Cultural Awareness). We have good math for the first, but we are still figuring out how to teach the second.

How the Robots Are Built (The Architecture)

The paper explains that these models have evolved, much like how smartphones have gotten bigger and more powerful.

  • The Old Way (Encoders): Early models were like librarians. They could look at a picture and a sentence and say, "Yes, these match," or "No, they don't." They were great at sorting things but couldn't write new stories.
  • The New Way (Decoders): Newer models are like storytellers. They can look at a picture and write a whole paragraph about it, or answer complex questions creatively.
  • The Size: These models have grown huge. Early ones had a few hundred million "brain cells" (parameters). The newest ones have up to 2 trillion. That's like going from a bicycle to a supersonic jet.

The Problem with the Tests (Benchmarks)

To see if the robots are smart, we give them tests. The paper found a major flaw in how we test them:

  • The "Translation Trap": About two-thirds of the tests are built by taking an English test, translating it into other languages, and seeing if the robot gets the same score.
    • Analogy: Imagine testing a chef by giving them a recipe in English, then translating it to French, and asking if they can cook the exact same dish. This tests if they can follow instructions, but it doesn't test if they know how to cook a traditional French meal that should be different from the English one.
  • The Missing "Local" Tests: There are very few tests that use pictures and questions specific to local cultures. It's hard to find a test where the picture of a "home" looks different for a German user versus a Nigerian user, and the robot is expected to know the difference.

What the Paper Found

  1. Progress is Real: We have moved from simple models that only spoke English to massive models that can handle dozens of languages.
  2. The Gap: Even though a model might claim to support 100 languages, we usually only test it on 5 or 10. We don't really know if it works for the other 90.
  3. The Transparency Issue: The biggest, most powerful models come from big tech companies. They tell us how they built the robot (the architecture), but they are very secretive about what they fed it (the training data). It's like a chef saying, "I made a great cake," but refusing to tell you what ingredients were in it. This makes it hard to know if the robot is biased or if it truly understands different cultures.
  4. The Future Challenge: We need to stop just translating English tests. We need to create new tests that are born in different cultures, with local images and local questions, to truly see if the robots are culturally aware.

Summary in One Sentence

This paper argues that while we have built amazing robots that can see and speak many languages, we are currently too focused on making them "neutral" (translating facts) and not enough on making them "culturally aware" (understanding local nuances), and our tests need to change to measure the latter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →