← Latest papers
💬 NLP

Information Satisfaction: A Reader-Centered Axis for Summarization Evaluation

This paper argues that current summarization evaluation metrics fail to capture individual user needs by ignoring reader personas, demonstrating through perturbation tests and expert human evaluation that both traditional and LLM-based metrics poorly correlate with human judgments of information satisfaction.

Original authors: Isabel Cachola, William Walden, Reno Kriz, Mark Dredze

Published 2026-08-17
📖 5 min read🧠 Deep dive

Original authors: Isabel Cachola, William Walden, Reno Kriz, Mark Dredze

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are in a library, but instead of books, the shelves are filled with millions of digital documents. Somewhere in this vast library, a robot is trying to write a short "summary sheet" for you. This field of science is called Natural Language Processing, and the specific task is Summarization: taking a long, complicated text and condensing it into a few sentences. For years, scientists have tried to teach computers how to write these summary sheets and then asked, "How good is it?" To answer this, they built metrics—mathematical rulers that measure things like how many words the robot used, how similar the words are to a "perfect" example, or how factually accurate the text sounds.

But here is the catch: a summary sheet isn't useful unless it fits the person holding it. If you are a medical student, you need a summary of a new vaccine that explains the chemistry and the test results. If you are a parent, you need a summary that explains if the vaccine is safe for your child and what the side effects are. The source document is the same, but your needs are totally different. For a long time, the "rulers" used to grade these summaries ignored the reader entirely. They just checked if the summary looked like a standard textbook answer, regardless of who was reading it. This paper asks a simple, vital question: Do our current rulers actually measure whether a summary satisfies your specific needs, or are they just measuring how well the robot mimics a generic style?

The authors of this paper decided to test the rulers by introducing a new concept called Information Satisfaction. Think of this as asking, "Did this summary actually give the reader the specific information they were looking for?" To do this, they didn't just look at the text; they looked at the Persona—the specific role and background of the reader, like "a biomedical researcher" versus "a family doctor."

First, the team put the popular "rulers" (the metrics) through a series of stress tests, like a video game level designed to break bad software. They took a perfect summary and started messing with it in specific ways. They added random, nonsense sentences to see if the ruler would notice the summary got worse. They chopped sentences out to see if the ruler noticed the summary got incomplete. They even rewrote the summary for a completely different audience (like changing a summary for a scientist into one for a journalist) to see if the ruler could tell the difference.

The results were a bit of a shock. Many of the most popular metrics, including the fancy new ones powered by massive AI models (called LLMs), failed the tests spectacularly. When the researchers added nonsense sentences, some of the AI judges actually gave the summary a higher score. When they shortened the summary, the scores went up and down in ways that made no sense. Most importantly, when they changed the summary to fit a different audience, the metrics often couldn't tell that the summary was now useless for the original reader. It was as if a teacher grading a math test couldn't tell the difference between a correct answer and a completely wrong one, or even between a test written for a 5th grader and one written for a PhD student.

Next, the researchers brought in real humans to see what they actually preferred. They recruited 20 experts from fields like medicine and computer science. Each expert was given a specific job (their persona) and a real question they had. The computer then generated four different summaries: some tailored to that expert's background, and some generic ones. The experts then played a "tournament," picking the best summary in head-to-head matchups.

Here is the twist: The experts consistently preferred the summaries that were tailored to their specific background. However, when the researchers compared the experts' choices to the scores given by the computer metrics, the computers were almost completely clueless. The metrics agreed with the human experts no better than if they had just been guessing. Even the new "Persona" metrics, which were specifically designed to check if the right information was included, failed to match human judgment. In fact, some of these new metrics agreed with humans less than random chance would suggest.

The paper concludes that we are currently using the wrong tools to measure the value of summaries. The old rulers and the new AI judges are great at checking if a summary sounds fluent or matches a reference text, but they are terrible at answering the most important question: "Does this help this specific person?" The authors suggest that until we build evaluation systems that truly understand who the reader is and what they need to learn, we won't be able to tell if a summary is actually good or just looks good on paper.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →