← Latest papers
💬 NLP

Evaluating LLM-Driven Summarisation of Parliamentary Debates with Computational Argumentation

This paper proposes a novel computational argumentation framework to evaluate the faithfulness of LLM-generated summaries of parliamentary debates by focusing on the formal preservation of reasoning structures, addressing the limitations of existing metrics in capturing argumentative consistency.

Original authors: Eoghan Cunningham, Derek Greene, James Cross, Antonio Rago

Published 2026-04-22
📖 4 min read☕ Coffee break read

Original authors: Eoghan Cunningham, Derek Greene, James Cross, Antonio Rago

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the European Parliament as a massive, chaotic kitchen where hundreds of chefs (politicians) are arguing over a giant, complex recipe (a new law). They shout about adding salt, removing sugar, or changing the cooking time. The final recipe has dozens of specific instructions (provisions).

Now, imagine you are a regular person trying to understand what happened in that kitchen. You can't read the 500-page transcript of everyone shouting. So, you ask a super-smart robot (an AI or Large Language Model) to write a short summary for you.

The Problem:
The robot writes a nice, short summary. But is it telling the truth?

  • Did it say the chefs wanted to add salt when they actually wanted to remove it?
  • Did it hide the fact that half the kitchen was screaming "No!" about a specific ingredient?
  • Did it make a controversial ingredient look like everyone loved it?

Current ways of checking if a summary is good are like checking if the summary uses the same words as the original. If the robot uses different words but gets the meaning wrong, those old checks fail. They are like judging a movie summary by counting how many times the word "explosion" appears, rather than checking if the hero actually died.

The Solution: The "Argument Map"
This paper proposes a new way to check the robot's work. Instead of just looking at words, they turn the debate into a logical battle map (called a Quantitative Bipolar Argumentation Framework).

Think of it like a Tug-of-War game:

  1. The Rope: The specific parts of the law being debated (e.g., "Extend the health agency's power").
  2. The Teams:
    • Team Green (Support): Chefs pulling the rope one way, saying "This is good!"
    • Team Red (Attack): Chefs pulling the other way, saying "This is bad!"
  3. The Score: The final position of the rope depends on how many people are on each team and how hard they are pulling.

How the New System Works:
The researchers built a system that creates two maps:

  1. Map A: The map of the real debate (the full transcript).
  2. Map B: The map of the robot's summary.

Then, they compare the two maps using five simple rules (properties) to see if the summary is a "faithful" representation of the fight:

  • Rule 1: Did anyone get left out? (Relevance) Did the summary forget to mention a chef who was shouting loudly?
  • Rule 2: Is the score balanced? (Pro-Con Ratio) If 10 people said "Yes" and 2 said "No" in the real debate, does the summary show that same 10-to-2 split? Or did it accidentally show 10 "Yes" and 0 "No"?
  • Rule 3: Who won the tug-of-war? (Acceptability) In the real debate, did the "Yes" team win? Did the summary also show the "Yes" team winning?
  • Rule 4: Is the intensity right? (Balance) If the "Yes" team was winning by a little bit in reality, does the summary show them winning by a huge landslide? (This is a common AI mistake: making things look more certain than they are).
  • Rule 5: Is the exact score close? (Accuracy) If the real score was 55% "Yes," is the summary's score close to 55%, or is it wildly off?

What They Found:
They tested this on real European Parliament debates using different AI models (like Claude and Llama).

  • The Old Way (Word Counting): The AI summaries looked great! They used similar words and sounded smart.
  • The New Way (The Battle Map): The summaries failed the logic test.
    • The AI often ignored the "No" team. It would summarize a heated argument as if everyone agreed.
    • It turned controversial issues (where the rope was in the middle) into settled facts (where the rope was pulled all the way to one side).
    • Essentially, the AI was "smoothing over" the conflict, making the debate look much more peaceful and one-sided than it actually was.

The Takeaway:
If we want AI to help us understand democracy, we can't just ask it to "summarize the text." We have to ask it to "summarize the argument."

This paper gives us a ruler to measure if an AI is lying to us about who won the debate, who lost, and how hard they fought. It's like checking a referee's scorecard to make sure the game wasn't rigged, even if the referee's report looks well-written.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →