← Latest papers
💬 NLP

AI systems and the reproduction of (standard) language ideologies in World Englishes

This paper argues that AI systems reproduce standard language ideologies by privileging Inner Circle norms and marginalizing non-dominant Englishes across their design and usage, while simultaneously reigniting debates about linguistic legitimacy and ownership in World Englishes through a "standardisation paradox" that both homogenizes and pluralizes the language.

Original authors: Kingsley Ugwuanyi

Published 2026-07-31
📖 7 min read🧠 Deep dive

Original authors: Kingsley Ugwuanyi

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the English language as a massive, bustling city. For a long time, the city planners (teachers, editors, and dictionaries) decided that only one specific neighborhood—the "Standard" district—was the real, correct city. Everyone else's neighborhoods, with their unique slang, accents, and grammar, were treated like messy construction zones or "broken" versions of the real thing. This idea is called a language ideology: a set of beliefs about which way of speaking is "good" and which is "bad," often used to judge people's intelligence or worth.

Now, enter the new city planners: Artificial Intelligence (AI). These are giant computer brains called Large Language Models (LLMs) that have read almost everything on the internet to learn how to talk. You might think, "Great! If they read everything, they'll learn every neighborhood's style!" But here's the twist: these AI brains are built on old blueprints. They tend to think the "Standard" neighborhood is the only one that matters, and they often treat everyone else's way of speaking as a mistake. This paper asks a big question: Are these AI tools just copying the old, unfair rules of the city, or are they accidentally creating a space where all the different neighborhoods can finally be heard?


The AI City Planners and the "Delve" Detective Story

This paper, written by Kingsley Ugwuanyi, investigates how AI systems are acting like strict gatekeepers of the English language. The author argues that AI isn't a neutral tool; it's an ideological actor. Think of an AI like a very fast, very confident robot chef. If you give it a recipe book (training data) that mostly contains recipes from fancy, expensive restaurants in the UK and US, the robot will learn to believe that only those recipes are "real food." If you ask it to cook a dish from a street vendor in Nigeria or India, it might get confused, or worse, it might try to "fix" the dish by adding ingredients that don't belong, making it taste like the fancy restaurant version instead of the original.

The paper finds that these AI systems reproduce standard language ideologies at every step of their creation:

  1. The Training Data: The "recipe books" the AI reads are flooded with text from the "Inner Circle" (the US, UK, etc.). This makes the AI think that the way people in these countries speak is the default, and everything else is a glitch.
  2. The Testing: When scientists test if the AI is smart, they use exams written in standard American English. If the AI tries to answer a question written in African American Vernacular English (AAVE) or Nigerian English, it often fails or gives a rude, confused answer. It's like giving a driver's test in a language the driver doesn't speak and then saying they are a bad driver.
  3. The Human Touch: Even when humans help train the AI (by correcting its mistakes), they are often told to make the AI sound "correct." This usually means scrubbing away any unique flavor from non-standard Englishes, forcing the AI to sound like a generic, polished robot from the Global North.

The "Delve" Controversy: When a Word Becomes a Crime

To show how this plays out in the real world, the paper tells a fascinating story about a single word: "delve."

Recently, people started noticing that AI chatbots love to use the word "delve" (meaning to dig deep into a topic). Soon, a wave of complaints started online. People from the US and UK began saying, "If you use the word 'delve,' you must be a robot!" They treated the word like a secret code that gave away a fake writer.

The paper points out that this reaction is actually a trap.

  • The Accusation: Some commentators claimed that because AI uses "delve" so much, it must be copying African English speakers, who they claimed use the word too much. They called this an "indignity," suggesting it's bad for AI to sound like it's from Africa.
  • The Reality Check: The author checked the data and found that "delve" is actually used quite a bit in many parts of the world, including Asia and Africa, long before AI existed. The word wasn't "AI-sounding"; it was just a normal word that some people use.
  • The Backlash: When people from the Global South (like Nigeria and India) saw these complaints, they pushed back hard. They pointed out that they were taught to use words like "delve" in school. They argued that calling their English "fake" or "AI-generated" is just a new way of saying, "Your English isn't good enough for us." It's a modern version of colonial policing, where people from powerful countries decide what counts as "real" English and what counts as "suspicious."

The paper suggests that this isn't just about a word; it's about who gets to decide. When a wealthy American investor tweets that using "delve" proves you are cheating, he is using his power to police how others speak. The paper shows that AI has become a new battleground where these old arguments about who owns English are being fought again.

The "Standardization Paradox": A Glimmer of Hope?

Is it all doom and gloom? The paper suggests there is a strange twist called the "standardization paradox."

Imagine a giant mirror (the AI). On one side, the mirror is being polished to show only one perfect, standard face (the US/UK English). This makes English look more uniform and boring. But on the other side, because the mirror is so huge and reflects so many different sources of data, it also starts to show glimpses of all the other faces, too.

  • The Good News: Because AI is trained on massive amounts of data from everywhere, it is exposed to a wider variety of Englishes than ever before. Some researchers are trying to build "adapters" (like little plug-in modules) that can teach the AI to speak specific dialects without breaking the whole system.
  • The Bad News: The paper warns that this doesn't mean the AI is suddenly fair. The people training the AI are often paid very little and are forced to follow strict rules that erase their own language quirks. So, while the AI hears the diversity, it often speaks in a standardized, Western voice.

Why This Matters to You

The paper concludes that this isn't just a technical problem for computer scientists; it's a real-world problem for people. If an AI tutor doesn't understand a student's Nigerian English, that student might fail. If an AI hiring tool rejects a resume written in Indian English, that person loses a job. If a voice recorder can't understand an elder from an Indigenous community, their history might be lost.

The author argues that we need to stop treating AI as a neutral tool and start seeing it as a system that can either reinforce old prejudices or help us build a more inclusive world. The challenge is to design these systems so they recognize that English isn't one single language with one "correct" way to speak, but a family of many different, equally valid languages. Until we fix the "recipe book" and the "taste testers," the AI will keep serving up the same old, exclusive menu.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →