Fair Representation in Parliamentary Summaries: Measuring and Mitigating Inclusion Bias
This paper evaluates five large language models summarizing European Parliament debates, revealing systematic biases against speakers based on order, language, and political affiliation, and proposes a hierarchical summarization method to effectively mitigate these inclusion biases where prompting strategies fail.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the European Parliament as a massive, bustling town hall meeting where hundreds of people from different countries, speaking different languages, and holding different political views gather to debate important issues. The transcripts of these meetings are thousands of pages long—too long for any regular citizen to read.
Enter Large Language Models (LLMs), the "AI scribes." Their job is to read these massive transcripts and write a short, easy-to-read summary so the public can understand what happened. The goal is to make democracy more accessible.
However, this paper asks a critical question: Is the AI scribe being fair? Or is it accidentally (or intentionally) ignoring certain people, twisting their words, or only listening to the loud voices at the start and end of the meeting?
Here is a breakdown of the paper's findings using simple analogies.
1. The "Middle Curse": The Lost Voices
The Problem:
Imagine a long line of people waiting to speak at a town hall. The AI scribe tends to listen very carefully to the first few people and the last few people. But the people speaking in the middle? The AI often zones out on them. It's like a radio that has a strong signal at the start and end of a song, but the middle part is static.
In technical terms, this is called the "Lost-in-the-Middle" problem. The study found that speakers in the middle of a debate were systematically ignored or poorly represented in the summaries, regardless of how important their points were.
The Fix:
The researchers tried a new method called Hierarchical Summarization.
- Old Way (Flat): The AI tries to swallow the whole 50-page document at once and spit out a summary. It gets overwhelmed and forgets the middle.
- New Way (Hierarchical): The AI acts like a smart editor. First, it summarizes each individual speech into a tiny note. Then, it groups those notes by topic (e.g., "All points about taxes," "All points about education"). Finally, it combines those topic notes into one final summary.
- Result: This "step-by-step" approach fixed the problem. By breaking the task down, the AI stopped ignoring the middle speakers. It was like giving the scribe a checklist so they didn't miss anyone.
2. The "Language Barrier": The Translation Gap
The Problem:
The European Parliament has 24 official languages. The study found that if a politician spoke in a "low-resource" language (a language with less data available on the internet for the AI to learn from, like some smaller European languages), the AI did a worse job representing them.
Even worse, this bias happened even when the AI was summarizing official English translations of those speeches. It's as if the AI had a "bad ear" for certain accents or dialects, and that bias persisted even after the words were translated. The speakers using these languages were less likely to be included accurately in the final summary.
3. The "Political Filter": The Left-Wing Bias
The Problem:
The researchers also checked if the AI favored certain political parties. They found a subtle but clear bias: Left-of-center parties (generally more progressive or social-democratic groups) were represented more fully than right-of-center parties.
The AI didn't necessarily lie about what the right-wing parties said; it just omitted more of their points. It was like a camera that focused perfectly on the people on the left side of the stage but kept the people on the right side slightly out of frame.
- Note: This bias was hidden at first because the "Middle Curse" was so bad that it drowned out the political bias. Once they fixed the "Middle Curse" using the step-by-step method, the political bias became visible.
4. The "Magic Prompt" Didn't Work
The researchers tried a simple trick: they told the AI, "Hey, please make sure you give equal attention to everyone!"
Result: It didn't work. The AI ignored the instruction. This is like telling a distracted student, "Please pay attention to the whole book," without actually changing how they study. The bias is baked into how the model processes information, not just a lack of politeness.
The Big Takeaway
This paper is a warning and a guide for the future of AI in democracy.
- The Warning: If we just plug AI into government systems without checking, we might accidentally silence the middle voices, disadvantage speakers of smaller languages, and skew political representation.
- The Guide: We can't just rely on "magic prompts." We need to change how the AI works. By breaking big tasks into smaller, structured steps (like the hierarchical method), we can fix some of the biggest fairness issues.
In short: AI is a powerful tool for making democracy accessible, but like any tool, it needs to be calibrated carefully. If we don't, the "summary" of our democracy might only reflect the voices of the first, the last, the English-speaking, and the left-leaning, leaving everyone else in the dark.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.