← Latest papers
💬 NLP

Multimodal Evaluation of Russian-language Architectures

This paper introduces MERA Multi, the first open multimodal evaluation framework for Russian-language architectures, featuring 18 culturally specific tasks across text, image, audio, and video modalities to assess model capabilities and establish a replicable methodology for diverse languages.

Original authors: Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov, Artem Safin, Maria Tikhonova, Alexander Kharitonov, Yulia Lyakh, Petr Surovtsev, Denis Shevelev, Vildan Saburov, Vasily Konovalov, Elisei Rykov, Ivan
Published 2026-01-27
📖 5 min read🧠 Deep dive

Original authors: Artem Chervyakov, Ulyana Isaeva, Anton Emelyanov, Artem Safin, Maria Tikhonova, Alexander Kharitonov, Yulia Lyakh, Petr Surovtsev, Denis Shevelev, Vildan Saburov, Vasily Konovalov, Elisei Rykov, Ivan Sviridov, Amina Miftakhova, Ilseyar Alimova, Alexander Panchenko, Alexander Kapitanov, Alena Fenogenova

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a group of very smart, new robots that can see pictures, hear sounds, watch videos, and read text all at the same time. These are called Multimodal Large Language Models (MLLMs). They are getting better every day, but until now, we only had a way to test them if they spoke English.

If you wanted to test a robot's ability to understand Russian culture, folklore, or specific Russian jokes, you had no ruler to measure it with. The existing tests were like trying to measure a Russian winter using a thermometer calibrated only for the Sahara desert.

This paper introduces MERA Multi, a brand-new, open "exam" designed specifically to test how well these AI robots understand the Russian language and culture across all senses: text, images, audio, and video.

Here is a breakdown of what they did, using some simple analogies:

1. The Problem: The "English-Only" Blind Spot

Think of the current AI world as a giant library. Most of the books (benchmarks/tests) are written in English. If you want to know if a student can read Russian, you can't just hand them an English book and say, "Read this." You need a test written in Russian that understands Russian context.

  • The Gap: Previous tests for Russian only looked at text. They didn't check if the AI could understand a Russian movie clip, a Russian song, or a picture of a Russian street scene.
  • The Solution: The authors built MERA Multi, the first "Russian-language driver's license test" for AI that checks all four senses.

2. The Exam: 18 Different Tasks

Instead of just one big test, they created 18 different mini-exams. Imagine a gym with 18 different stations:

  • The "Ear" Stations (Audio): Can the AI tell the difference between a Russian folk song and a pop song? Can it understand a spoken command like "Turn on the heater"?
  • The "Eye" Stations (Image): Can the AI look at a picture of a Russian math problem and solve it? Can it read a menu in a Russian restaurant from a photo? Can it spot if a picture is "weird" (like a penguin driving a car)?
  • The "Video" Stations: Can the AI watch a short clip of a Russian cooking show and explain the steps in order?
  • The "Brain" Stations (Reasoning): Can the AI look at a picture and understand the story behind it, not just the objects? (e.g., "Why is this person sad?" based on visual clues).

3. The "Cultural Translator"

The authors didn't just translate English tests into Russian. That would be like translating a joke about American baseball into Russian; it wouldn't make sense to a Russian speaker.

  • Custom Built: They built these tests from scratch, using Russian cultural references (like Soviet movies, Russian folklore, and local history).
  • The "Turing Test" Twist: Some tasks are designed to see if the AI can have a natural, human-like conversation in Russian, understanding sarcasm, idioms, and cultural nuances, just like a real person would.

4. The Grading System: The "Smart Teacher"

How do you grade an AI that gives a long, rambling answer?

  • The Human Baseline: First, real humans took the test to set a "gold standard" score.
  • The "Judge" AI: Since humans can't grade thousands of AI answers instantly, the authors trained a special "Judge AI." Think of this Judge as a strict but fair teacher. It doesn't just look for the exact right word; it checks if the meaning is correct.
  • Two Scores: They give the AI two scores:
    1. Exact Match: Did it say the exact right word?
    2. Semantic Score: Did it get the idea right, even if the words were slightly different?

5. Keeping the Test Fair (Anti-Cheating)

A big problem in AI testing is "data leakage." This is like a student memorizing the test questions before the exam because the test questions were accidentally used to teach the student in the first place.

  • Watermarks: The authors put invisible "watermarks" on their test data (like a hidden signature in a painting) to track if an AI model secretly learned from the test data during its training.
  • Leakage Detection: They built a "lie detector" tool that can tell if a model has seen the test questions before. If a model tries to cheat, the system can catch it.

6. The Results: Who Passed?

They tested over 50 different AI models (both free/open-source and paid/closed-source).

  • The Winners: The top performers were "Omni-models" (models that try to do everything) from the Qwen family. They scored the highest overall because they were good at all the different stations, not just one.
  • The Specialists: Some models were great at just one thing (like only understanding audio) but failed the others.
  • The Reality Check: Even the best models still struggle with complex Russian cultural nuances and difficult reasoning tasks. There is still a big gap between how well humans do and how well AI does.

Summary

MERA Multi is a comprehensive, culturally aware "report card" for Russian-speaking AI. It proves that to truly understand a language, an AI needs to understand the culture, the history, and the context, not just the dictionary definitions. It provides a fair, transparent way for researchers to see which AI models are actually smart and which ones are just guessing.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →