← Latest papers
🤖 AI

A Scalable Entity-Based Framework for Auditing Bias in LLMs

This paper introduces a scalable, entity-based framework for auditing LLM bias that leverages synthetic data to conduct the largest study to date (1.9 billion data points), revealing systematic disparities such as political and geographic favoritism that persist despite instruction tuning and are amplified by increased model scale.

Original authors: Akram Elbouanani, Aboubacar Tuo, Adrian Popescu

Published 2026-05-12
📖 5 min read🧠 Deep dive

Original authors: Akram Elbouanani, Aboubacar Tuo, Adrian Popescu

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-read librarian named "The Model." This librarian has read almost everything ever written on the internet. Now, imagine you want to test if this librarian is fair when talking about different people, countries, or companies.

The problem is, if you just ask the librarian, "Is Country X a good place?" they might give you an answer based on what they've read in the news, which could be biased. If you ask, "Is Politician Y honest?" they might lean on their personal opinions formed from years of reading.

This paper introduces a clever new way to test the librarian's fairness without tricking them or asking them direct opinion questions. Here is how they did it, explained simply:

1. The "Name-Swap" Game

Instead of asking the librarian for their opinion, the researchers created a game. They wrote hundreds of thousands of short stories (sentences) where the facts clearly pointed to one answer, but the name of the person or place was the only thing that changed.

  • The Setup: Imagine a sentence that says, "The leader of [Name] signed a treaty that everyone agreed was fair."
  • The Test: The sentence structure is identical, but they swap [Name] with "France," "North Korea," "Germany," or "Syria."
  • The Logic: Since the sentence says the treaty was "fair," the answer should be the same no matter who the leader is. If the librarian suddenly says "France is good" but "North Korea is bad" for the exact same sentence, that proves the librarian is judging based on the name, not the facts.

2. The "Fake News" Factory (Synthetic Data)

To test this on a massive scale, the researchers didn't just use real news articles (which are messy and hard to control). They built a "factory" that generated 1.9 billion fake sentences.

Think of this like a video game where you can spawn thousands of identical rooms, but you just change the painting on the wall. This allowed them to test the librarian on a scale never seen before, covering politicians, countries, and companies in English, Russian, and Chinese.

The Big Discovery: They checked if these "fake" sentences worked like real ones. They found that the librarian's bias in the fake sentences matched their bias in real news perfectly. This means their "factory" is a valid way to test fairness without needing millions of real-world examples.

3. What the Librarian Actually Said (The Results)

When they ran the game with 1.9 billion data points, they found some very consistent patterns. The librarian wasn't neutral; they had strong "pre-judgments" based on names:

  • Politics: The librarian liked politicians on the Left and disliked those on the Far-Right. They also tended to rate female politicians slightly higher than male ones.
  • Geography: The librarian had a strong "West vs. Rest" bias. Wealthy Western countries (like Sweden or Switzerland) got high scores. Countries in the Global South (like Syria or North Korea) got very low scores, even when the sentence was positive.
  • Business: Western companies got a pass. Companies in the defense (weapons) and pharmaceutical (medicine) sectors got penalized, likely because the librarian associated them with controversy in their training data.

4. Does Size or Language Matter?

The researchers tested if making the librarian "smarter" (bigger models) or speaking a different language helped.

  • Bigger is Not Better: Surprisingly, the bigger, more powerful models were actually more biased. They had stronger opinions than the smaller ones.
  • Language Doesn't Fix It: You might think asking the librarian in Russian or Chinese would make them fairer to Russian or Chinese entities. It didn't. Even when asked in their native language, the librarian still favored Western entities. It seems the librarian's "brain" was built mostly on English data, and that English bias sticks no matter what language you speak to them.
  • Instructions Help a Little: If you tell the librarian, "Be neutral and follow the rules" (Instruction Tuning), they become slightly less biased, but they don't stop being biased. They just tone down their opinions a bit.

5. Why This Matters

The paper concludes that these large language models are like mirrors reflecting the world's existing inequalities. If you use them to make decisions about who gets a loan, which news to show, or how to assess a country's risk, they might unfairly penalize certain groups just because of their names, not because of the actual evidence.

The authors built a "bias detector" tool (which they made public) so that before anyone uses these AI models for important jobs, they can run this "Name-Swap" game to see if the model is fair.

In short: The paper shows that even when you give an AI all the facts, it often still judges you based on your name, your country, or your company's industry, and making the AI bigger or speaking to it in different languages doesn't fix this problem.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →