Standard Language Ideology in AI-Generated Language
This paper introduces the concept of "standard AI-generated language ideology" and a sociotechnical taxonomy to analyze how large language models reinforce standard language biases through mechanisms like legitimation and erasure, ultimately offering recommendations to protect linguistic diversity and ensure a more just AI future.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine that Large Language Models (LLMs) like ChatGPT are like a giant, super-smart librarian who has read almost everything on the internet. You ask this librarian a question, and they write a story or an answer for you.
This paper argues that this librarian isn't neutral. Instead, they have a very strong, invisible bias: they think that only one specific way of speaking is "correct," "smart," and "professional," while all other ways of speaking are "broken," "funny," or "wrong."
The authors call this bias Standard Language Ideology. In the real world, this usually means the librarian prefers "Standard American English" (the way rich, white, middle-class people in the US speak) and treats everything else as inferior.
Here is a breakdown of the paper's main ideas using simple analogies:
1. The "Gold Standard" Filter
Imagine the librarian has a filter on their desk. When you hand them a note written in a dialect like African American English (AAE) or a regional accent, the filter automatically rewrites it into "Standard American English" before they even read it.
- The Problem: The paper calls this Legitimation. The AI treats the "Standard" version as the only legitimate way to communicate.
- The Result: If you speak in your natural dialect, the AI might ignore you, misunderstand you, or give you a worse answer. It's like a restaurant that only serves food to people wearing a specific suit, and tells everyone else to change their clothes before they can order.
2. The "Bad Service" for Some Customers
Because the librarian has mostly read books written in "Standard" English, they are terrible at understanding other dialects.
- The Problem: This is Unequal Quality of Service. If you ask a question in a minoritized dialect, the AI might get confused, refuse to answer, or even flag your question as "hate speech" just because of the words you used.
- The Result: People who speak these dialects are forced to "code-switch" (pretend to speak like the "Standard" group) just to get a helpful answer. It's like being forced to wear a mask to enter a building just so the security guard will let you in.
3. The "Cartoon" Effect
Sometimes, the AI does try to speak your dialect, but it does it in a way that is mean or silly.
- The Problem: This is Stereotyping and Racialization. If you ask the AI to write in a specific dialect, it often produces a caricature—a cartoon version of that speech filled with stereotypes, insults, or "broken" grammar that no real person actually uses.
- The Result: It reinforces the idea that people who speak that way are uneducated or funny. It's like a comedian doing a bad impression of a group of people, but the AI presents it as a serious, factual representation of how they speak.
4. The "Thief" in the Library
The AI can learn to speak a minoritized dialect, but it does so without asking the people who actually speak it.
- The Problem: This is Appropriation. The AI takes the "cool" or "unique" parts of a culture's language, turns them into a product, and sells them to everyone else, while the original speakers get nothing.
- The Result: A white person can use the AI to sound like they belong to a Black community without having to deal with the racism that community faces. It's like a thief stealing a family's heirloom recipe, putting it on a box of cereal, and selling it for profit, while the family gets no credit or money.
5. The "Eraser"
Finally, the AI sometimes just deletes these dialects entirely.
- The Problem: This is Erasure. To avoid the risks of stereotyping or appropriation, developers might just program the AI to never speak in those dialects.
- The Result: If you want the AI to reply in your native dialect, it simply won't. It's like a library that decides to stop stocking books in a certain language because "it's too hard to manage," effectively making that language invisible in the digital world.
The Big Question: What Should the AI Do?
The paper asks a tough question: Should the AI just stick to "Standard" English to be safe, or should it try to speak every dialect?
- If it sticks to "Standard," it hurts people who don't speak that way.
- If it tries to speak every dialect, it risks making fun of them or stealing their culture.
The authors say there is no easy "technical fix" (like just adding more data) because this is a problem of power, not just code. The people who built the AI (mostly in big US tech companies) decided what "Standard" is, and they didn't ask the communities whose languages were ignored.
The Solution: Who Holds the Keys?
Instead of just trying to make the AI "less biased," the authors suggest we need to change who is in charge. They propose three main ideas:
- Measure the Harm: We need to test AI not just on how fast it is, but on how it treats different dialects. Is it being mean? Is it ignoring people?
- Let Communities Build Their Own: Instead of big companies taking data from communities, those communities should build their own AI tools. (The paper mentions groups in Africa and New Zealand who are already doing this).
- Share the Power: If big companies want to use a community's language, they need to sit down with that community, let them make the decisions, and share the profits. It's not about "consulting" the community; it's about letting them own the process.
In short: The paper argues that AI is currently reinforcing the idea that some ways of speaking are "better" than others. To fix this, we can't just tweak the software; we have to give the power back to the people whose languages are being used, ignored, or mocked.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.