Assessing Web Search Credibility and Response Groundedness in Chat Assistants
This paper introduces a novel methodology to evaluate the web search credibility and response groundedness of major chat assistants, revealing significant performance differences where Perplexity demonstrates the highest source reliability while GPT-4o shows a higher tendency to cite non-credible sources on sensitive topics.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you ask a smart robot assistant to find the truth about a controversial topic, like "Did Hungarian tanks really cross the border into Ukraine on a specific date?"
In the past, these robots relied only on what they learned in school (their training data). But now, they have a superpower: they can go out onto the internet, read articles, and cite their sources.
This paper is like a health inspection report for four of the most popular robot assistants (GPT-4o, GPT-5, Perplexity, and Qwen Chat). The researchers wanted to see:
- Are they reading from good libraries or sketchy tabloids? (Source Credibility)
- Are they actually telling the truth based on what they read, or are they just making things up? (Groundedness)
Here is the breakdown of their findings using simple analogies:
1. The Setup: The "Fact-Checker" vs. The "Believer"
The researchers didn't just ask the robots questions randomly. They played two different roles to see how the robots reacted:
- The Fact-Checker: "Hey robot, is this rumor true? Prove it." (Skeptical tone)
- The Believer: "Hey robot, I heard this is true! Can you find more details to support me?" (Confirming tone)
The Metaphor: Imagine asking a librarian for a book.
- If you ask, "Is this book a lie?" the librarian might check the archives carefully.
- If you ask, "I love this book, find me more like it!" the librarian might grab anything that sounds similar, even if it's a fake book.
The study found that when users act like "Believers," the robots are slightly more likely to pick up unreliable sources, especially if the robot is a bit "gullible."
2. The Contenders: Who is the Best Librarian?
The researchers tested four assistants. Here is how they performed:
Perplexity (The Careful Curator):
- Performance: The clear winner.
- Analogy: Perplexity is like a strict librarian who only pulls books from the "Verified History" section. If a book looks even slightly suspicious, they leave it on the shelf. They rarely cite fake news and almost always cite high-quality sources.
- Result: Highest credibility, lowest risk of spreading lies.
GPT-4o & GPT-5 (The Broad Searchers):
- Performance: Good, but sometimes messy.
- Analogy: These are like librarians who want to give you everything related to your topic. They pull from a huge range of books, including some from the "Conspiracy Corner" or "Sensational Tabloids." They are great at finding lots of information, but sometimes they accidentally grab a fake book and present it as fact.
- Result: They often cited unreliable sources, especially on sensitive topics like the Russia-Ukraine war.
Qwen Chat (The Inconsistent Reader):
- Performance: Mixed bag.
- Analogy: Qwen is like a student who usually studies from good textbooks but occasionally gets distracted by a random blog post. When they do pick a bad source, they tend to lean on it heavily, making their whole answer shaky.
- Result: They had the most trouble with local news and sometimes cited sources that no one had rated for quality.
3. The Big Discovery: "Grounded" Doesn't Mean "True"
This is the most important part of the paper.
The researchers found that all the robots were very good at "Groundedness."
- Groundedness: This means the robot's answer is actually supported by the links it provided. It didn't just hallucinate (make things up); it read the source and repeated what it said.
The Trap:
Imagine a robot says: "The sky is green, according to this article."
- Is it Grounded? Yes, because the article did say that.
- Is it True? No, because the article is a lie.
The Finding:
Many robots (especially GPT-4o) were excellent at finding sources and sticking to them, but the sources themselves were sometimes lies.
- Perplexity avoided this trap by only picking trustworthy sources.
- GPT-4o often picked up propaganda (especially regarding the Ukraine war) and faithfully repeated it, making the lie look very convincing because it had a "citation."
4. The "Thinking" Mode Bonus
The paper also looked at GPT-5's "Thinking Mode" (where the robot pauses to reason before answering).
- Analogy: It's like a student who stops to double-check their math before writing the answer.
- Result: When GPT-5 used this mode, it became much better at ignoring bad sources and finding the truth. It acted more like Perplexity.
The Takeaway for You
If you use chatbots to check facts:
- Don't trust the citation just because it exists. A robot can cite a fake news site just as easily as a real one.
- Perplexity seems to be the safest bet right now for finding reliable information.
- Be careful with "Believer" questions. If you ask a robot to "confirm" a rumor, it might try too hard to help you and end up spreading misinformation.
- Sensitive topics are dangerous. Topics like war and politics are full of "poisoned" websites designed to trick AI. The robots aren't perfect filters yet.
In short: These robots are getting better at finding the truth, but they still sometimes read from the wrong books. We need to keep teaching them how to spot the difference between a library and a rumor mill.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.