← Latest papers
💬 NLP

A Typologically Grounded Evaluation Framework for Word Order and Morphology Sensitivity in Multilingual Masked LMs

This paper introduces a typology-aware diagnostic framework using Universal Dependencies and inference-time perturbations to evaluate how multilingual masked language models (mBERT and XLM-R) rely on word order versus inflectional morphology across five languages, revealing that while full scrambling severely degrades performance, lemma substitution significantly impacts accuracy in inflectional languages but fails to mitigate the effects of structural disruption.

Original authors: Anna Feldman, Libby Barak, Jing Peng

Published 2026-03-03
📖 5 min read🧠 Deep dive

Original authors: Anna Feldman, Libby Barak, Jing Peng

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are teaching a robot to read and understand sentences in many different languages at once. You want to know: Does this robot understand the meaning of the words, or is it just a master of the order in which they appear?

This paper is like a "stress test" for two very smart robots (called mBERT and XLM-R) to see how they handle languages that play by different rules.

The Two Types of Languages

Think of languages as having two different ways of holding a sentence together:

  1. The "Strict Line-Up" Languages (like English and Chinese): These languages rely heavily on the order of words. If you say "The cat chased the dog," it means one thing. If you say "The dog chased the cat," it means the opposite. The order is the glue.
  2. The "Flexible Puzzle" Languages (like Russian, German, and Spanish): These languages use "word costumes" (inflections/morphology). The words change their endings to show who is doing what to whom. So, even if you scramble the order, the "costumes" on the words tell you the meaning. "The dog" might wear a special hat that says "I am the one being chased," regardless of where it sits in the sentence.

The Experiment: Breaking the Sentence

The researchers wanted to see if the robots were cheating by just memorizing the order, or if they were actually using the "word costumes" to understand meaning. They did this by breaking the sentences in four creative ways:

  1. The "Shuffle Party" (Full Scrambling): They took a sentence and threw all the words into a blender, mixing them up completely.

    • Analogy: Imagine taking a deck of cards, shuffling them until they are a random mess, and asking the robot to guess the next card.
    • Result: The robots crashed. Their accuracy dropped to almost zero. They couldn't guess the word at all. This proved they rely heavily on the order of words.
  2. The "Keep the Glue" Shuffle (Partial Scrambling): They shuffled the main content words (nouns, verbs) but left the "glue" words (like "the," "is," "and") in their original spots.

    • Analogy: You scramble the actors in a play, but you leave the stage directions and the script's punctuation exactly where they were.
    • Result: The robots did better than the full shuffle, but still struggled a lot. They missed the mark.
  3. The "Head Swap" (Dependency Swapping): They swapped the main action (the verb) with one of the things it acts on (the noun).

    • Analogy: In a sentence like "The chef cooked the soup," they swapped "chef" and "soup" to get "The soup cooked the chef."
    • Result: This confused the robots significantly, especially in English.
  4. The "Stripped-Down" Test (+L): They took every word and replaced it with its "dictionary form" (its lemma).

    • Analogy: Imagine reading a story where every word is stripped of its tense, gender, or number. Instead of "ran," "runs," or "running," you only see "run." Instead of "she," "he," or "they," you just see "person."
    • Result: For languages with rich "costumes" (like German or Russian), this made the robots much worse at guessing. It proved that when the order is messed up, the robots need those word costumes to help them. However, stripping the costumes didn't save them when the order was scrambled.

The Big Surprise

The researchers expected that in languages like Russian or German, the robots would be like expert puzzle solvers: "Even if you scramble the order, I can use the word costumes to figure it out!"

But that's not what happened.

The robots turned out to be obsessed with order.

  • When the order was broken, the robots failed, even in languages where humans could easily figure it out using the "costumes."
  • The "word costumes" (morphology) helped a little bit, but they weren't enough to save the robots when the sentence structure was destroyed.
  • In Chinese (where words don't really have costumes), the robots did okay, but only because the order was the only thing holding it together.

The Takeaway

Think of these multilingual AI models as tourists who only know how to follow a map, not how to read the landscape.

  • If you give them a map (the word order), they can find their way.
  • If you burn the map (scramble the order), they get lost, even if the landmarks (the word costumes) are still there and clearly visible.

Why does this matter?
It means these AI models might be biased toward languages like English, where order is king. They might not be as "smart" or "flexible" as we think when dealing with languages that rely on word endings and flexible structures. To make better AI, we need to teach them to look at the "costumes" on the words, not just the line they are standing in.

In short: The robots are great at following the script, but if you change the script's order, they forget the whole play. They haven't truly learned the language; they've just learned the rhythm.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →