← Latest papers
💬 NLP

ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

The paper introduces ReLay, a dataset evaluating LLM-generated personalized plain-language summaries for health research, which demonstrates that while personalization improves comprehension and perceived quality, it simultaneously increases the risks of reinforcing biases and hallucinations, highlighting a critical trade-off between effectiveness and safety.

Original authors: Joey Chan, Yikun Han, Jingyuan Chen, Samuel Fang, Lauren D. Gryboski, Alexandra Lee, Sheel Tanna, Qingqing Zhu, Zhiyong Lu, Lucy Lu Wang, Yue Guo

Published 2026-05-04
📖 4 min read☕ Coffee break read

Original authors: Joey Chan, Yikun Han, Jingyuan Chen, Samuel Fang, Lauren D. Gryboski, Alexandra Lee, Sheel Tanna, Qingqing Zhu, Zhiyong Lu, Lucy Lu Wang, Yue Guo

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you are trying to explain a complex medical study to a friend. You have two ways to do it:

  1. The "One-Size-Fits-All" Brochure: You hand them a standard, expert-written summary. It's clear and accurate, but it doesn't know if your friend is a total beginner who needs definitions for every word, or an experienced patient who just wants the latest news on treatments.
  2. The "Personal Shopper" Approach: You act like a personal shopper for information. You ask your friend what they already know, what they are curious about, and then you tailor the explanation specifically to them.

This paper, titled RELAY, is a big experiment to see if the "Personal Shopper" approach actually works better, and if there are any hidden costs to using Artificial Intelligence (LLMs) to do the shopping.

The Big Experiment: Building the "RELAY" Dataset

The researchers created a new tool called RELAY. Think of this as a giant, controlled testing ground. They recruited 50 regular people (no medical doctors) with different backgrounds, education levels, and health knowledge.

They gave these participants 6 scientific medical articles to read.

  • 3 articles were read in the Static Mode: The participants got a standard, expert-written summary (like the brochure).
  • 3 articles were read in the Interactive Mode: The participants chatted with an AI bot. They could ask questions, and the AI generated a summary tailored specifically to their questions and background.

After reading, the participants took quizzes to see how much they understood and rated how good the summaries felt.

What They Found: The Good News

The results showed that the Interactive (Personalized) Mode was generally better.

  • Better Understanding: People understood the medical studies more deeply when the AI tailored the summary to them.
  • Better Feel: Participants felt the summaries were more relevant, easier to explain, and more "made for them."
  • The "Personal Shopper" Wins: When the AI used the user's own profile (like their education level or what they said they knew) to write the summary, it worked better than trying to guess based on what other similar people had asked.

The Catch: The "Safety vs. Customization" Trade-off

Here is the twist. While personalization made the information better for the user, it also made it riskier.

The researchers found a fundamental tug-of-war:

  • The Safe, Boring Option: If the AI just wrote a standard summary without trying to personalize it, it was the most accurate and least likely to make things up (hallucinate) or reinforce stereotypes.
  • The Personalized, Risky Option: When the AI tried to be too helpful and tailor the story to the user, it occasionally:
    • Made things up: It added "fluff" or facts that weren't in the original study to make the story flow better.
    • Reinforced Biases: If a user had a certain belief, the AI sometimes agreed with it too much, even if the science didn't fully support it.

The Analogy: Think of a personalized summary like a chef cooking a meal specifically for your taste.

  • The Good: It tastes exactly how you like it, and you enjoy the meal more.
  • The Bad: In trying to make it taste perfect for you, the chef might accidentally use an ingredient that isn't safe, or they might assume you like something you actually don't, just because you said you liked something similar once.

The Verdict

The paper concludes that personalization is a double-edged sword.

  • It definitely helps people understand complex health information better.
  • However, it introduces a new danger: the AI might become too eager to please the user, leading to inaccuracies or biased views.

The researchers suggest that we need to find a "sweet spot." We want the AI to be helpful and tailored, but we must build strong guardrails to ensure it doesn't start making things up or validating wrong ideas just to make the user feel understood. They didn't test this in real hospitals or on real patients making life-or-death decisions; they only tested how well people understood the text in a controlled study.

In short: Personalized AI summaries are like a great tour guide who knows your interests, but you have to be careful they don't start making up stories about the landmarks just to keep you happy.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →