← Latest papers
💬 NLP

"Intelegi Româneste?'' A Recipe for Romanian Vision-Language Models

This paper presents a systematic study on building Romanian-specific Vision-Language Models by translating English corpora, curating the culturally grounded HoraVQA benchmark, and demonstrating that models adapted with Romanian language backbones outperform larger multilingual counterparts across all evaluated tasks.

Original authors: Mihai Masala, Marius Leordeanu, Mihai Dascalu, Traian Rebedea

Published 2026-06-01
📖 4 min read☕ Coffee break read

Original authors: Mihai Masala, Marius Leordeanu, Mihai Dascalu, Traian Rebedea

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a brilliant, multilingual robot chef. This chef is famous for cooking amazing meals using English recipes (English text) and looking at pictures of food (images). But if you ask the chef to cook a traditional Romanian dish using a Romanian recipe, they might get confused, burn the food, or just serve you a generic American burger instead.

This paper, titled "A Recipe for Romanian Vision-Language Models," is essentially a guide on how to teach that robot chef to cook authentic Romanian cuisine. The authors, researchers from Romania, didn't just tell the chef "try harder"; they built a whole new kitchen, translated the recipes, and even changed the pictures on the menu to match the local culture.

Here is the breakdown of their work in simple terms:

1. The Problem: The "English-Only" Kitchen

Most advanced AI models (called Vision-Language Models) are like chefs who only know English. They are great at describing a picture of a dog in English, but if you show them a picture of a Romanian street sign or a traditional folk dance, they often fail. They might misread the text on the sign or not understand the cultural context of the dance.

The researchers found that simply translating the English questions into Romanian wasn't enough. The AI still struggled because it hadn't "seen" enough Romanian things or learned how to read Romanian text inside images (like on a bus ticket or a menu).

2. The Solution: A Complete Romanian Makeover

The team built a specialized version of these AI models, which they call RoVLM. To do this, they followed a strict "recipe":

  • Translating the Menu (Data): They took thousands of existing English image-and-text datasets and translated them into Romanian. But they didn't just translate the words; they used special tools to translate the text inside the images too. If a picture had a sign saying "Open," they edited the image to say "Deschis."
  • Adding Local Flavor (HoraVQA): They realized that translating English tests wasn't fair. So, they created a brand new test called HoraVQA. Think of this as a "Romanian Culture Quiz." Instead of asking about generic things, they asked questions about Romanian landmarks, traditional food, folk customs, and local history. They hired native speakers to write these questions and pick the photos, ensuring the test was truly grounded in Romanian life.
  • The "OCR" Challenge: A big part of the recipe was teaching the AI to read text inside pictures (OCR). They found that standard AI "eyes" (vision backbones) were too good at understanding the general vibe of a picture but bad at reading the tiny letters on a sign. They discovered that using a specific type of "eye" (called CLIP) was much better at reading Romanian text than the newer, more "multilingual" eyes.

3. The Results: Small Chef, Big Appetite

The most surprising part of their findings is that specialization beats size.

Usually, in AI, bigger is better. A giant model with billions of parameters is expected to beat a smaller one. However, the authors found that their smaller, Romanian-specialized models performed better than:

  • The original English-only versions of the same size.
  • Even some models that were twice as big but not specialized for Romanian.

It's like a small, local Romanian grandmother who knows exactly how to make mămăligă (polenta) perfectly, beating a giant, high-tech international food factory that tries to make it but gets the recipe wrong.

4. Key Lessons Learned

The paper highlights three main things that made their "Romanian Chef" successful:

  1. The Language Backbone: Changing the "brain" of the AI to a Romanian-specific one helped a little, but it wasn't the magic ingredient.
  2. The Vision Backbone: Choosing the right "eyes" (specifically CLIP) was crucial for reading text in images.
  3. The Data Mix: The most important factor was the data. They needed a mix of translated data, real Romanian documents (like PDFs), and synthetic charts. Without this specific "diet" of data, the model couldn't learn to read or understand Romanian culture.

Summary

In short, this paper proves that you can't just take a general AI and expect it to work perfectly in a low-resource language like Romanian. You have to feed it specific, culturally relevant food (data), teach it to read local signs (OCR), and test it on local culture (HoraVQA). When you do that, a smaller, specialized AI can outperform massive, generic giants.

The authors have released their "recipe" (code, data, and models) to the public, so anyone can try to cook up their own Romanian AI chef.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →