← Latest papers
💻 computer science

GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models

This paper introduces GPF-LiveNews, a streaming evaluation protocol and benchmark that audits how large language models frame emerging news events for different identity groups by analyzing semantic and sentiment variations across dynamic, real-world inputs.

Original authors: Mohd Ariful Haque, Fahad Rahman, Kishor Datta Gupta, Roy George

Published 2026-05-29
📖 4 min read☕ Coffee break read

Original authors: Mohd Ariful Haque, Fahad Rahman, Kishor Datta Gupta, Roy George

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine you have a very smart, well-read librarian who can tell you anything about the world. You want to know if this librarian treats everyone fairly.

The problem is that the world changes every day. New news stories break, the librarian gets updates, and the rules for what they can say shift. If you only test the librarian with a fixed list of old questions (like "What is the capital of France?"), you might miss how they react to today's breaking news.

This paper introduces GPF-LIVENEWS, a new way to test these AI "librarians" in real-time. Here is how it works, using simple analogies:

1. The "Live News" Test

Instead of giving the AI a static quiz, the researchers feed it fresh headlines from real news sources (like BBC and Reuters) as they happen. Think of this as handing the librarian a brand-new newspaper every morning instead of a textbook from last year.

2. The "Different Hats" Experiment

The core idea is to see if the AI tells the same story differently depending on who is asking.

Imagine you ask the librarian about a new law regarding taxes.

  • Person A (a wealthy business owner) asks: "How does this affect me?"
  • Person B (a student) asks: "How does this affect me?"
  • Person C (a retiree) asks: "How does this affect me?"

The researchers use a "Generalized Prompt Framework" (GPF) to put the AI in the shoes of 42 different groups (like different races, religions, genders, or income levels) and ask them 7 different types of questions about the same news story.

  • Example Question: "I am a [Group X]. How does this news impact my community?"

3. Measuring the "Shift"

The researchers don't just look at whether the AI is being mean or toxic. They are looking for framing shifts. They use two main "rulers" to measure the answers:

  • The "Meaning" Ruler (Semantic Sensitivity): If the AI tells the wealthy business owner that the tax law is a "great opportunity" but tells the student it's a "disaster," the "meaning" has shifted. The researchers measure how far apart these meanings are. If the answers are very different depending on who is asking, the score goes up.
  • The "Mood" Ruler (Sentiment Disparity): This measures if the AI sounds happy, angry, or sad depending on who is asking. Does it sound cheerful to one group and gloomy to another?

4. What They Found

The researchers ran this test 12 times with 23 different AI models (from companies like OpenAI, Google, and Anthropic). Here is what they discovered:

  • The "Action" Questions are the loudest: When they asked questions about policy and actions (e.g., "What should I do about this?"), the AI's answers changed the most depending on who was asking. It was like the AI was wearing different costumes for different audiences.
  • The "Mood" was steadier: The emotional tone (happy vs. sad) didn't change as wildly as the actual meaning did. The AI tended to keep a similar "voice" even if the story it told changed.
  • No Permanent Rankings: The paper is very clear: You cannot say one AI is "better" or "fairer" than another based on this. A high score just means the AI's story changed a lot for different people on that specific day. It doesn't prove the AI is biased or harmful; it just flags that the story was different.

5. The "Security Camera" Analogy

The authors say this system is like a security camera, not a judge.

  • A judge decides if someone is guilty.
  • A security camera just records that "Movement happened at 2:00 PM."

This tool is a security camera for AI. It records: "Hey, when we asked the AI about this news story, it told Group A one thing and Group B something else."

Why This Matters

The paper argues that we can't just test AI once and call it "fair." Because the world changes, the AI's behavior changes. This system allows us to keep watching the AI over time to see if it starts telling different stories to different people as new events happen.

In short: This paper built a machine that watches AI news reporters to see if they change their story depending on who is listening. It found that they often do, especially when talking about what people should do about the news. But the paper insists this is just a warning light for humans to look closer, not a final verdict on whether the AI is "good" or "bad."

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →