What Is The Political Content in LLMs' Pre- and Post-Training Data?
This paper investigates the political content in LLM training data and finds that pre-training corpora are systematically skewed toward left-leaning perspectives, with these biases persisting through post-training stages and strongly correlating with the resulting model's political stances, thereby highlighting the critical role of data composition in shaping model behavior.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are hiring a new assistant to help you write emails, answer questions, and give advice. You don't just hand them a blank notebook; you give them a massive library of books, articles, and blog posts to read before they start working. This is exactly how Large Language Models (LLMs) like the ones powering chatbots are built. They "read" terabytes of internet data to learn how to speak and think.
This paper asks a simple but crucial question: What kind of library did we give these AI assistants, and does that library make them biased?
Here is the breakdown of their findings, explained with some everyday analogies.
1. The "Left-Leaning" Library
The researchers looked at the "books" (data) used to train three different famous AI models (OLMo, Falcon, and Pythia). They wanted to know: Is the library balanced, or is it stacked?
- The Finding: The library is heavily stacked toward left-leaning viewpoints.
- The Analogy: Imagine you are trying to learn about politics by reading a stack of newspapers. If 80% of the papers are from The Guardian or HuffPost (left-leaning) and only 10% are from National Review or The Telegraph (right-leaning), you are going to learn a very specific version of the world.
- The Numbers: In the data they analyzed, there were 2.3 to 12 times more left-leaning documents than right-leaning ones. It's like a diet where you only eat one flavor of ice cream; eventually, you start thinking that's the only flavor that exists.
2. Where Do the Books Come From? (The Source)
The researchers also looked at where these documents came from.
- The Finding: Left-leaning texts mostly came from established news outlets (like The New York Times or The Washington Post). Right-leaning texts mostly came from blogs, personal websites, and fringe communities (like Townhall or FreeRepublic).
- The Analogy: Think of the left-leaning data as a formal dinner party hosted by professional journalists. The right-leaning data is more like a rowdy backyard BBQ with passionate individuals shouting their opinions. The AI "ate" both, but the formal dinner party provided much more food.
3. The "Echo Chamber" Effect
One of the big questions was: If we use different libraries for different AI models, will they have different biases?
- The Finding: Surprisingly, no. Even though the researchers used different data sources and different filtering methods for different models, the political bias was almost identical.
- The Analogy: Imagine three different chefs (different AI models) using three different recipes (different datasets) to make soup. You'd expect them to taste different. But instead, they all tasted exactly the same because they all used the same main ingredient: a massive, unfiltered dump of the internet that happens to be naturally left-leaning. The "filtering" didn't change the flavor much.
4. The "First Bite" vs. The "Dessert"
AI training happens in stages:
- Pre-training: The model reads the massive library (the main course).
- Post-training: Humans tweak the model to be more helpful and polite (the dessert).
- The Finding: The political bias is baked into the main course (pre-training). By the time the humans try to tweak the model with the "dessert" (post-training), the bias is already there. The dessert didn't change the flavor; it just made the AI speak more politely about its existing opinions.
- The Analogy: If you teach a child that "the sky is green" when they are 2 years old (pre-training), and then at age 10 you tell them to "be polite and helpful" (post-training), they will still think the sky is green. They will just be very polite about it. The bias was established early and stuck.
5. The Connection: Data = Behavior
Finally, the researchers checked if the AI's actual answers matched the books it read.
- The Finding: There was a very strong match (90% correlation). If the training data said "Climate Change is urgent," the AI said "Climate Change is urgent." If the data said "Immigration should be restricted," the AI was more likely to agree with restrictions (though the data was mostly against restrictions, so the AI was mostly against them too).
- The Analogy: The AI is a parrot. It doesn't have its own opinions; it just repeats the patterns it heard most often in the room. If the room was full of people shouting left-wing ideas, the parrot will sound left-wing.
The Big Takeaway
The paper concludes that we cannot fix AI bias just by tweaking the model later. The bias is in the ingredients.
If we want AI to be truly balanced and fair, we need to fix the library before the AI starts reading. We need to ensure the "books" we feed it represent a true mix of voices, not just the loudest or most common ones on the internet.
In short: The AI isn't "evil" or "biased" by choice. It's just a mirror reflecting the skewed library we gave it. If the library is unbalanced, the reflection will be too.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.