Are LLMs More Skeptical of Entertainment News?
This paper reveals that certain large language models exhibit a genre-specific skepticism toward legitimate entertainment news, misclassifying it as fake more frequently than hard news due to biases against the genre's epistemic legitimacy rather than stylistic features, thereby highlighting the need for genre-stratified evaluation in automated credibility assessment.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you have a team of four very smart, high-tech librarians (the AI models). Their job is to stand at the door of a massive library and decide which books are "Real News" and which are "Fake News."
The researchers wanted to see if these librarians treat all "Real News" books equally. Specifically, they wanted to know: Do these AI librarians unfairly flag "Entertainment News" (like celebrity gossip) as fake, even when it's true, just because it looks different from "Hard News" (like politics or crime)?
Here is what they found, broken down simply:
1. The "Fashion Police" Effect
The researchers gave the librarians a mix of real stories: some about serious politics (Hard News) and some about celebrity drama (Entertainment News).
- The Result: Two of the librarians (DeepSeek-V3.2 and GPT-5.2) acted like strict fashion police. They looked at the celebrity stories and thought, "This looks too dramatic, too emotional, or too focused on private lives. It must be fake!" They flagged about 10% more real celebrity stories as fake compared to real political stories.
- The Other Librarians: The other two librarians (Claude Opus 4.6 and Gemini 3 Flash) didn't care about the style. They treated the celebrity stories and political stories fairly, flagging them at the same rate.
The Takeaway: Being "smart" doesn't mean being fair. Some AIs have a built-in bias against the style of entertainment news, assuming that if a story feels like gossip, it must be a lie.
2. The "Makeover" Experiment (Style vs. Substance)
The researchers wondered: Is the AI just judging the "clothes" the story is wearing? Entertainment news often uses colorful, emotional language, while hard news is dry and factual.
To test this, they took 50 real celebrity stories and used an AI to rewrite them into "boring, serious political news" style. They kept the facts exactly the same but changed the tone.
- The Result: The makeover didn't work well.
- For one AI, the fake rate stayed exactly the same.
- For another AI, the makeover actually made things worse—it flagged more of the rewritten stories as fake!
- The Metaphor: It's like trying to trick a bouncer by putting a tuxedo on a rock star. The bouncer (the AI) still recognized the rock star's energy and said, "You don't belong here," even though the outfit was perfect. The bias wasn't just about the writing style; it was deeper.
3. The "Specialist" Prompt
The researchers tried giving the librarians a new job description. Instead of just "News Checker," they told one AI: "You are now a specialized expert in entertainment news. Your job is to fact-check celebrity stories."
- The Result: This worked like magic for one librarian (DeepSeek-V3.2). It stopped flagging real celebrity stories as fake, and it didn't start missing actual fake news.
- The Catch: It didn't work for the other librarian (GPT-5.2). No matter what instructions they gave, that AI kept being suspicious of celebrity stories.
The Takeaway: You can sometimes fix the bias with the right instructions, but it depends on which AI you are using. There is no "one-size-fits-all" fix.
4. Why Does This Happen?
The researchers looked at the notes the AIs wrote when they made mistakes. They found two main reasons why the AIs were skeptical:
- "I can't prove it": The AIs thought, "This story is about a celebrity's private life. I can't go check their diary, so it must be fake."
- "This genre is weak": The AIs seemed to have a hidden rule that entertainment news is just a "lesser" type of journalism, so they trust it less than political news.
The Big Picture
The main lesson from this paper is that average scores can be misleading.
If you look at a test score that says "This AI is 95% accurate," you might think it's perfect. But this paper shows that the AI might be 99% accurate on political news and only 85% accurate on entertainment news. It's "cheating" by being too harsh on one specific type of story.
In short: These AI systems aren't just checking if facts are true; they are also judging which types of journalism are allowed to be true. Sometimes, they unfairly decide that celebrity news doesn't deserve the same trust as political news, even when both are telling the truth.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.