← Latest papers
💬 NLP

\textit{Versteasch du mi?} Computational and Socio-Linguistic Perspectives on GenAI, LLMs, and Non-Standard Language

This interdisciplinary paper combines critical sociolinguistics and computational linguistics to examine how Large Language Models can be technically adapted to handle non-standard linguistic varieties like South Tyrolean dialects and Kurdish, aiming to address digital language divides and advance decolonial AI strategies.

Original authors: Verena Platzgummer, John McCrae, Sina Ahmadi

Published 2026-03-31
📖 5 min read🧠 Deep dive

Original authors: Verena Platzgummer, John McCrae, Sina Ahmadi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the internet and Artificial Intelligence (AI) as a massive, bustling Grand Library. For a long time, this library only stocked books written in a few specific, "perfect" languages (like standard English, German, or Mandarin). The librarians (the AI developers) decided that only these "perfect" books were worth reading, organizing, and understanding.

This paper, written by researchers Verena Platzgummer, John McCrae, and Sina Ahmadi, argues that this library is leaving millions of people out in the cold. It focuses on two specific groups of people who are being ignored: speakers of South Tyrolean dialects (a local, informal version of German spoken in Italy) and speakers of various Kurdish languages (spoken across the Middle East).

Here is the breakdown of the paper using simple analogies:

1. The Problem: The "Perfect" Library vs. Real Life

In the real world, people don't always speak in "perfect," textbook grammar. They speak in dialects, slang, and informal ways.

  • The Analogy: Imagine you walk into a fancy restaurant that only accepts reservations made in a specific, formal code. If you try to order in your local accent or slang, the waiter (the AI) doesn't just ignore you; they might not even know you exist.
  • The Reality: AI models (LLMs) are trained on "standard" language data found in textbooks and official documents. They treat informal dialects as "noise" or errors, rather than valid ways of speaking. This creates a digital language divide: if you don't speak the "standard" version, you can't fully participate in the digital world.

2. The Hidden Tax: The "Token" Toll Booth

One of the most technical parts of the paper explains why AI struggles with dialects, using a concept called Tokenization.

  • The Analogy: Think of AI as a delivery service that charges you based on how many "boxes" (tokens) it takes to pack your message.
    • If you speak Standard English, the AI packs your sentence into one or two neat, large boxes. It's cheap and fast.
    • If you speak a dialect (like South Tyrolean or a Kurdish variety), the AI doesn't have a "box" for those specific words. It has to smash your sentence into tiny, meaningless fragments (like breaking a word into individual letters) to fit it into its system.
  • The Result: You end up paying 3 to 15 times more (in computing power and money) to send the same message in your dialect compared to English. The paper calls this the "Tokenization Tax." It's like being charged extra just for speaking your mother tongue.

3. The Two Case Studies: Who is Being Left Out?

The authors look at two specific examples to show how deep the problem goes:

  • South Tyrolean Dialects (Italy):
    • People in this region speak a mix of German dialects in their daily lives (on WhatsApp, at home, in the market).
    • The Issue: Because this dialect doesn't have an official "ID code" (an ISO code) in the computer world, AI systems literally don't know it exists. It's like a city that doesn't appear on the GPS map. Even though people use it every day, the AI treats it as invisible.
  • Kurdish Varieties:
    • Kurdish isn't just one language; it's a family of many dialects spoken by 40 million people. Some are well-supported (like Northern and Central Kurdish), but others (like Southern Kurdish or Hawrami) are almost completely ignored.
    • The Issue: These languages have been suppressed by governments for decades, meaning there is very little written text available to train the AI. The AI sees them as "low quality" because it hasn't been fed enough data, creating a vicious cycle where the lack of data leads to bad AI, which leads to even less data.

4. The "Test" is Rigged

The paper also points out that the way we test AI is unfair.

  • The Analogy: Imagine a driving test where everyone is tested on driving a Ferrari on a highway. If you drive a tractor on a dirt road, you fail the test, even though you are a skilled driver in your own environment.
  • The Reality: AI benchmarks (tests) are mostly based on Western, English-speaking culture and standard grammar. They don't test if the AI understands local proverbs, cultural jokes, or informal speech. So, an AI might get a "perfect score" on a test but fail to understand a real human conversation in a dialect.

5. What Needs to Happen? (The Solution)

The authors argue that fixing this isn't just a technical problem; it's a political and social one. We can't just "tweak the code." We need a new approach:

  • Stop the "Extraction": Big Tech companies shouldn't just scrape data from the internet without permission. They need to work with communities.
  • "Nothing About Us Without Us": This is the golden rule. Communities speaking these dialects should own their own data and help build the AI tools. They should be the architects, not just the subjects.
  • New Rules for Companies: Big Tech should be required to report how well their AI works for different dialects (a "Dialect Gap Report") and pay a "tax" or fee to support the development of these languages.
  • Government Help: Governments need to fund the creation of data for these languages, treating them as a public good rather than a niche hobby.

The Bottom Line

The paper concludes that if we don't fix this, the digital world will become a place where only "standard" speakers have a voice, and everyone else is silenced. To have a truly democratic internet, AI needs to learn to understand the messy, beautiful, and varied ways humans actually speak, not just the clean, standardized versions found in textbooks.

In short: The AI is currently a snob that only talks to people who dress and speak like it. This paper asks the AI to take off its snobbery, learn the local slang, and treat every human voice with respect.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →