← Latest papers
💬 NLP

A Computational Approach to Language Contact -- A Case Study of Persian

This study investigates how historical language contact influences the intermediate representations of a monolingual Persian language model, revealing that while universal syntactic information remains largely unaffected, morphological features like Case and Gender are significantly shaped by contact-induced structural changes.

Original authors: Ali Basirat, Danial Namazifard, Navid Baradaran Hemmati

Published 2026-01-29
📖 5 min read🧠 Deep dive

Original authors: Ali Basirat, Danial Namazifard, Navid Baradaran Hemmati

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart robot that has only ever read books written in Persian. It has never been taught English, Arabic, Turkish, or Japanese. In fact, it doesn't even know those languages exist.

Now, imagine you hand this robot a sentence written in Turkish. You ask it: "Hey, what language is this? And what are the grammar rules here?"

You might expect the robot to say, "I don't know, I've only read Persian!" But this paper asks a deeper question: Does the robot's internal "brain" (its digital representations) secretly show signs that it understands the history of Persian, even if it was never taught those other languages directly?

Since Persian has been neighbors with, traded with, and influenced by many other languages for centuries, the authors wondered if the robot's "Persian brain" had absorbed some invisible fingerprints of those neighbors.

The Experiment: A Detective Story

The researchers used a tool called ParsBERT (a Persian-only AI model) and fed it sentences from 8 different languages, ranging from Arabic (a major historical neighbor) to Japanese (a stranger with no real contact).

They used two main "detective tools" to peek inside the robot's brain:

  1. The Information Meter: This measures how much the robot knows about a specific feature (like "Is this a noun?" or "Is this word masculine?"). Think of it like checking how loud a radio signal is.
  2. The Spotlight: This tries to find where in the robot's brain that knowledge is hiding. Is it stored in one specific neuron (a single lightbulb), or is it scattered across the whole room like a diffuse glow?

The Big Discovery: The "Universal" vs. The "Specific"

The results revealed a fascinating split in how the robot thinks, which the authors describe as a difference between universal rules and local habits.

1. The Universal Rules (The "Hard Drive")

When the researchers asked the robot about Universal Part-of-Speech (UPOS) tags—basically, "Is this a noun, a verb, or an adjective?"—the robot performed consistently well across all languages.

  • The Analogy: Imagine the robot has a universal dictionary of "types of words." Whether the word is in Persian, English, or Japanese, the robot knows that a "verb" is a verb.
  • The Result: The robot didn't care about language contact here. History didn't change how it saw the basic building blocks of sentences. These rules are too deep and universal to be shaken by centuries of borrowing words.

2. The Local Habits (The "Furniture")

When the researchers asked about Morphology (the specific shapes of words, like Case or Gender), the robot's behavior changed drastically based on the language's history with Persian.

  • The "Case" Test (Who did what to whom?):

    • The Problem: Persian doesn't use "cases" (changing word endings to show if someone is the subject or object). It uses word order instead.
    • The Result: When the robot looked at languages with complex case systems (like German or Russian, which have many word endings), it was completely lost. It couldn't find the information.
    • The Exception: When the robot looked at languages that also rely on word order or simple markers (like English or French), it did much better.
    • The Takeaway: The robot only "understands" the grammar of other languages if that grammar looks like Persian's. If the other language uses a system Persian never used (like complex case endings), the robot's brain doesn't have a map for it.
  • The "Gender" Test (Masculine vs. Feminine):

    • The Problem: Persian has no grammatical gender (words aren't "male" or "female").
    • The Result: The robot struggled with languages that have complex gender systems (like Russian with three genders). However, it did okay with languages that have simpler gender systems or rely on natural sex (like Arabic or French), likely because Persian has borrowed many words from Arabic, and those borrowed words sometimes kept their "flavor."
    • The Takeaway: The robot's understanding of gender is shaped by what it has actually seen in Persian. It can't invent a complex gender system it never learned.

The "Ghost" in the Machine

One of the most interesting findings was how this information was stored.

  • For languages with no contact (like Japanese), the robot had a few very specific "lightbulbs" that lit up only for Japanese. It was easy to spot.
  • For languages with heavy contact (like Arabic or Turkish), the information wasn't in a single lightbulb. Instead, it was a diffuse glow spread across the whole brain.
  • The Metaphor: Think of it like a house. If you invite a stranger (Japanese) in, they sit in a specific chair. If you invite a long-time family friend (Arabic) who has lived there for centuries, they don't sit in one chair; they have touched everything, moved the furniture, and left their scent all over the house. You can't point to one spot and say, "This is the Arabic part." The influence is everywhere, but it's mixed in with the Persian structure.

The Conclusion

The paper concludes that a language model trained on just one language (Persian) does show traces of its history, but only in a very selective way.

  • It keeps the universal rules (like what a verb is) safe and unchanged.
  • It adapts the specific rules (like gender or case) only if they match the "shape" of the Persian language.

If a language contact changed the deep structure of Persian (like adding complex case endings), the robot would have learned it. But since Persian mostly borrowed words and kept its own sentence structure, the robot's brain reflects that: it knows the borrowed words, but it still thinks in Persian grammar.

In short: The robot is a Persian speaker who knows a lot of foreign words, but it still thinks in Persian. It hasn't learned to speak the grammar of its neighbors, only to recognize the shapes of words it has borrowed from them.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →