← Latest papers
💬 NLP

Multilingual Language Models Encode Script Over Linguistic Structure

This paper demonstrates that compact multilingual language models primarily organize their internal representations around surface-level orthographic cues rather than abstract linguistic structures, with true typological abstraction emerging gradually in deeper layers without forming a unified interlingua.

Original authors: Aastha A K Verma, Anwoy Chatterjee, Mehak Gupta, Tanmoy Chakraborty

Published 2026-04-08
📖 5 min read🧠 Deep dive

Original authors: Aastha A K Verma, Anwoy Chatterjee, Mehak Gupta, Tanmoy Chakraborty

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a multilingual AI model as a giant, bustling library where books from every country in the world are stored. The librarians (the AI's internal neurons) are trying to organize these books so they can answer questions in any language.

For a long time, researchers hoped these librarians had a secret, universal "Rosetta Stone" in their heads—a perfect, abstract way of understanding meaning that worked the same way for English, Hindi, Chinese, and Russian. They thought the librarians ignored the cover of the book (the script) and focused only on the story inside.

This paper, "Encode Script Over Linguistic Structure," flips that idea on its head. The authors (from IIT Delhi) went into the library to peek behind the scenes and found something surprising: The librarians care way more about the cover of the book than the story inside.

Here is the breakdown using simple analogies:

1. The "Cover" Matters More Than the "Story" (Orthography vs. Language)

Imagine you have a story about a cat.

  • Scenario A: The story is written in English (Latin alphabet).
  • Scenario B: The exact same story is written in Hindi (Devanagari script).
  • Scenario C: The story is written in Hindi, but someone transliterated it into Latin letters (like "Hindi" written as "Hindi").

The researchers found that the AI treats these three scenarios as three completely different libraries.

  • When the AI sees the Hindi script, it uses one set of "librarians."
  • When it sees the Romanized Hindi (Latin letters), it uses a totally different set of librarians.
  • Crucially, the "Romanized Hindi" librarians don't even talk to the "English" librarians, even though they both use the same alphabet. They are stuck in their own isolated rooms.

The Takeaway: The AI doesn't really understand "Hindi" as a concept. It understands "Hindi written in Devanagari" and "Hindi written in Latin letters" as two totally different things. It's obsessed with the visual shape of the letters (the script), not the language itself.

2. The "Shuffled Deck" Test (Word Order)

To see if the AI understands grammar (syntax), the researchers took sentences and shuffled the words (e.g., "The cat sat on the mat" became "mat on the sat cat the").

  • The Result: Surprisingly, the AI didn't panic much. The same "librarians" kept working.
  • The Metaphor: It's like if you gave a librarian a pile of books with the pages torn out and mixed up. They could still tell you, "Oh, this is a mystery novel," just by looking at the cover art and the font style, even if the story inside made no sense.

The Takeaway: The AI relies heavily on word statistics and visual cues (what words usually appear together) rather than deep grammatical rules. It's a master of patterns, not necessarily a master of logic.

3. The "Deep Dive" (Where is the Real Understanding?)

The researchers looked at the AI's "brain" layer by layer, from the bottom (where it sees the letters) to the top (where it generates answers).

  • Bottom Layers: These are obsessed with the script. They are like the front desk clerks who only care if the book is red, blue, or has a specific font.
  • Deeper Layers: As you go deeper, the AI does start to understand the relationships between languages (like how Hindi and Urdu are related, or how English and German share roots).
  • The Catch: Even though the deep layers can see these connections, the AI doesn't actually use them to generate text. It's like having a map of the world in the back room, but the delivery driver (the AI generating text) only looks at the street signs (the script) to know where to go.

4. The "Causal Test" (What actually makes the AI talk?)

Finally, the researchers tried to "break" the AI by turning off specific parts of its brain to see what happened.

  • Turning off the "Script-Sensitive" neurons: The AI started glitching. It would mix scripts inside a single word (e.g., writing "Hindi" with a mix of Hindi and English letters) or suddenly switch languages mid-sentence.
  • Turning off the "Grammar/Typology" neurons: The AI kept talking just fine.

The Big Revelation: The parts of the AI that are responsible for generating fluent text are the ones that are invariant to surface changes (neurons that stay calm whether the script changes or words are shuffled). The parts that hold the "deep linguistic knowledge" are just there for show; they aren't the ones doing the heavy lifting for generation.

Summary: The "Surface-Level" AI

Think of this multilingual AI not as a polyglot genius who speaks 100 languages fluently, but as a very talented costume designer.

  • If you give it a Latin script costume, it acts like a Latin speaker.
  • If you give it a Devanagari script costume, it acts like a Devanagari speaker.
  • If you give it a Romanized Hindi costume, it acts like a third, weird character that isn't quite English and isn't quite Hindi.

It doesn't have a single, unified "soul" that understands all languages equally. Instead, it has many different masks, and it switches between them based entirely on how the words look on the page.

Why does this matter?
If you want to build safer, more robust AI, you can't just assume it understands the meaning of a language. You have to realize that if you change the script (even slightly), the AI might completely forget what it's talking about or switch to a different "persona." It's a reminder that for these models, form often rules over function.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →