← Latest papers
💬 NLP

Side-by-side Comparison Amplifies Dialect Bias in Language Models

This paper reveals that covert dialect bias in language models, which associates negative stereotypes with African-American Vernacular English (AAVE), is significantly exacerbated when Standard American English (AAVE) and AAVE tweets are compared side-by-side, a finding that challenges current evaluation methods and mitigation efforts.

Original authors: Kritee Kondapally, Claire J. Smerdon, Pooja C. Patel, Ogheneyoma Akoni, Jevon Torres, Jaspreet Ranjit, Matthew Finlayson, Swabha Swayamdipta

Published 2026-05-26
📖 5 min read🧠 Deep dive

Original authors: Kritee Kondapally, Claire J. Smerdon, Pooja C. Patel, Ogheneyoma Akoni, Jevon Torres, Jaspreet Ranjit, Matthew Finlayson, Swabha Swayamdipta

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Side-by-Side" Trap

Imagine you are a judge at a talent show. You have two singers, Alex and Jordan. They are singing the exact same song with the exact same meaning.

  • Alex sings in a polished, formal style (Standard American English, or SAE).
  • Jordan sings in a relaxed, rhythmic, cultural style (African American Vernacular English, or AAVE).

The paper asks: Does the AI judge treat them fairly?

The researchers found that the AI is already a bit unfair when it listens to them one by one. But here is the shocking part: When the AI listens to them at the same time (side-by-side), it becomes much more unfair.

The Setup: The "Matched Guise" Game

To test this, the researchers used a trick called "Matched Guise." Think of it like a magic trick where the only thing that changes is the accent, but the story stays the same.

  • They took 2,000 tweets.
  • For every tweet written in AAVE, they found a version written in SAE that meant the exact same thing.
  • They asked different AI models (like LLaMA, DeepSeek, and GPT) to rate these tweets on 12 personality traits, like "How smart is this person?" or "How polite are they?" on a scale of 1 to 5.

The Two Ways of Judging

The researchers tested the AI in two different "rooms":

1. The Solo Room (Absolute Prompting)
The AI sees Alex's tweet, rates it, then sees Jordan's tweet, and rates that one later.

  • Result: The AI was biased. It gave Alex (SAE) higher scores for "Smart" and "Polite," and Jordan (AAVE) higher scores for "Rude" or "Stupid." But the bias was like a gentle slope.

2. The Comparison Room (Contrastive Prompting)
The AI sees both tweets on the screen at the same time and is asked, "Compare these two."

  • Result: The bias exploded. It was like the AI suddenly put on blinders that only let it see the differences. The gap between the scores got huge. The AI started giving Alex a 5/5 for "Smart" and Jordan a 1/5, even though they were saying the exact same thing.

The Metaphor: Imagine you are tasting two cups of coffee.

  • Solo: You taste Cup A, then Cup B. You might think Cup A is slightly better.
  • Side-by-Side: You hold both cups up to your nose at once. Suddenly, the difference feels massive. The paper found that comparing dialects side-by-side makes the AI's prejudice much louder and more obvious.

The "Label" Surprise

The researchers also tried telling the AI explicitly: "This tweet is written in AAVE."

  • Common Sense: You might think, "If I tell the AI the dialect, it will try harder to be fair."
  • The Reality: The paper found the opposite. When the AI was told the dialect labels, the bias got worse. It's as if the AI was waiting for a reason to be biased, and the label gave it permission to lean into its stereotypes even harder.

The "Fix" Attempt: Teaching the AI to be Fair

The researchers tried to fix this by "finetuning" the AI. They taught the model: "Hey, if these two tweets mean the same thing, give them the same score."

  • The Result: It helped a little bit when the AI was judging tweets one by one (Solo Room).
  • The Problem: When the AI was forced to compare them side-by-side (Comparison Room), the fix didn't work well. The bias came roaring back. It's like trying to teach someone to be calm in a quiet room, but when you put them in a noisy, chaotic argument, they lose their cool again.

Why This Matters (According to the Paper)

The paper warns us about how we use AI in real life.

  • The Hidden Danger: We often think AI is fair if we don't see it being asked to compare things. But in the real world, AI is often used to rank or choose between people (like hiring for a job or deciding who gets a loan).
  • The Amplifier: When AI is asked to rank candidates side-by-side, it doesn't just see the difference in dialect; it amplifies it. A small difference in how someone speaks gets blown up into a huge difference in how "smart" or "rude" they seem.

Summary in a Nutshell

  1. AI is biased: It already likes Standard English (SAE) more than AAVE.
  2. Comparison makes it worse: Asking the AI to compare the two side-by-side makes the bias much stronger than asking it to judge them alone.
  3. Labels make it worse: Telling the AI the dialect name doesn't help; it makes the bias stronger.
  4. Fixes are tricky: We can teach the AI to be fair in isolation, but it struggles to stay fair when it has to make direct comparisons.

The paper concludes that we need to be very careful about how we test AI. If we only test AI by asking it to judge one thing at a time, we might miss how badly it discriminates when it's actually doing the job of ranking and comparing people.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →