← Latest papers
💬 NLP

Measuring cross-language intelligibility between Romance languages with computational tools

This paper introduces a novel computational metric based on lexical, surface, and semantic similarity to measure and validate cross-language intelligibility among the five main Romance languages, finding that the results align with human experimental data and linguistic intuitions.

Original authors: Liviu P Dinu, Ana Sabina Uban, Bogdan Iordache, Anca Dinu, Simona Georgescu

Published 2026-02-10
📖 3 min read☕ Coffee break read

Original authors: Liviu P Dinu, Ana Sabina Uban, Bogdan Iordache, Anca Dinu, Simona Georgescu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The "Language Mirror" Problem: Can You Understand Your Neighbor?

Imagine you are at a massive international dinner party. You speak Spanish, and the person sitting next to you speaks Portuguese. You both know you’re "cousins" because your languages sound similar, but you aren't quite sure: if they start telling a story, will you catch the plot, or will you just be nodding politely while feeling totally lost?

This paper is essentially a high-tech attempt to measure exactly how much "understanding" exists between the major Romance languages (Spanish, French, Italian, Portuguese, and Romanian) using computers instead of just guessing.


The Core Idea: The Three Pillars of Understanding

The researchers realized that understanding a "cousin" language isn't just about words looking the same on paper. They created a new formula called the DLI (Lexical Intelligibility Index). Think of it like a three-legged stool; if one leg is missing, the whole thing falls over.

  1. The "Look" (Orthography): Does the word look like the one in your language? If you see "tempo" (Italian) and "tiempo" (Spanish), your brain goes, "Aha! I know that!"
  2. The "Sound" (Phonetics): This is the tricky part. A word might look familiar on paper, but if it sounds totally different when spoken, the "mirror" is broken. The paper notes that French is a great example: it might look like its cousins on paper, but once a French person starts speaking, the sounds change so much that the "cousins" often don't recognize them.
  3. The "Meaning" (Semantics): This is the "False Friend" trap. Imagine you and a friend use the word "cool" to mean "temperature," but your friend uses it to mean "fashionable." You’re using the same word, but you aren't actually communicating. The researchers used AI (word embeddings) to make sure the words actually mean the same thing, not just that they sound the same.

The Big Discovery: The "Asymmetry" Glitch

One of the most interesting things the researchers found is that intelligibility is asymmetrical. In science, "symmetrical" means if A understands B, then B understands A. But in language, it’s more like a one-way street.

The Romanian Paradox:
The study found that Romanian is a bit of a "rebel" in the Romance family.

  • The "One-Way Mirror": Romanian speakers are actually quite good at understanding other Romance languages. However, when other Romance speakers listen to Romanian, they often struggle to understand it.
  • Why? It’s partly because Romanian has picked up a lot of "Slavic" flavors (words from neighboring Eastern European languages) and has some unique historical twists. It’s like a cousin who moved to a different country and came back with a totally different accent and slang—you can understand them, but they sound like a stranger to you!

Why Does This Matter?

The researchers didn't just want to build a math formula; they wanted to see if a computer could "feel" language the way humans do. They compared their computer scores to real human tests (where people actually tried to fill in missing words in sentences).

The result? The computer was remarkably accurate. It matched the human experience with high correlation.

The Takeaway

By using AI to look at the look, sound, and soul (meaning) of words, we can finally map out the "family tree" of languages. This helps us understand how languages drift apart over centuries and how we might better bridge the gap between people who speak "cousin" languages but still struggle to find common ground.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →