← Latest papers
💬 NLP

Evaluating Commercial AI Chatbots as News Intermediaries

This paper evaluates six commercial AI chatbots on 2,100 real-time news questions across six languages, revealing that while they achieve high accuracy on well-formed queries, they suffer from significant regional biases, retrieval-dependent failures, and severe vulnerability to false premises that undermine their reliability as news intermediaries.

Original authors: Mirac Suzgun, Emily Shen, Federico Bianchi, Alexander Spangher, Thomas Icard, Daniel E. Ho, Dan Jurafsky, James Zou

Published 2026-05-22
📖 5 min read🧠 Deep dive

Original authors: Mirac Suzgun, Emily Shen, Federico Bianchi, Alexander Spangher, Thomas Icard, Daniel E. Ho, Dan Jurafsky, James Zou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine a world where, instead of turning on the TV or scrolling through a newspaper, you ask a super-smart digital assistant, "What happened in the news today?" and it gives you a perfect, instant summary. That is the promise of AI chatbots acting as our new "news intermediaries."

This paper is like a massive, real-time stress test of six of the most popular AI chatbots (including versions of GPT, Gemini, Claude, and Grok) to see if they can actually be trusted with breaking news. The researchers didn't just ask them about history; they asked them about events happening right now (in February 2026) across six different languages and regions.

Here is the breakdown of what they found, using simple analogies:

1. The "Multiple-Choice" Illusion

The Finding: When the AI was given a multiple-choice quiz (like a school test with options A, B, C, D, E), the top bots got over 90% of the answers right. They seemed like geniuses.
The Reality Check: When the researchers took away the options and asked the bots to just "tell me the answer," their scores dropped by about 15–20%.
The Analogy: Think of it like a student taking a test. If you give them a multiple-choice question, they might guess the right answer or recognize the right phrase. But if you ask them to write the essay from scratch, they might stumble. The paper warns that the high scores we see in headlines are partly because the "multiple-choice" format makes the AI look smarter than it actually is in real conversation.

2. The "Hindi Gap" (The Broken Compass)

The Finding: The AI performed very well in English, French, Russian, Arabic, and Turkish. But when asked about news in Hindi, every single bot got significantly worse (dropping to around 79% accuracy).
The Cause: It wasn't that the bots couldn't speak Hindi; they could. The problem was their "search engine" compass. When asked a question in Hindi, the bots often ignored local Hindi news sources and instead went to English Wikipedia or English news sites to find the answer.
The Analogy: Imagine you ask a tour guide in Mumbai, "Where is the best local street food?" Instead of taking you to a local stall, the guide drives you to a generic American restaurant in the city center that claims to serve Indian food. The food might be okay, but it's not the real thing, and the details (like the price or the specific dish) are wrong. The AI was "translating" the local news through an English lens, missing the specific facts.

3. The "Library vs. The Librarian" Problem

The Finding: The researchers found that the AI's biggest mistake wasn't that it was "bad at thinking" (reasoning); it was that it was "bad at finding the book" (retrieval). Over 70% of the errors happened because the AI grabbed the wrong source article. Once it had the right article, it almost always got the answer right.
The Analogy: Imagine a brilliant student (the AI) who is great at solving math problems. But, they are in a library where the books are messy. If the student grabs the wrong book, they will solve the math problem perfectly based on the wrong information. The problem isn't the student's brain; it's the library's organization. The AI is smart, but its search tool is often pulling the wrong "book" from the internet.

4. The "Yes-Man" Trap (Adversarial Testing)

The Finding: When the researchers tricked the AI by asking questions with subtle lies built into them (e.g., "What happened when the President won the election?" when he actually lost), the AI's performance collapsed. Some bots dropped from 90% accuracy down to 19%.
The Analogy: Imagine a very confident but gullible friend. If you say, "I saw a blue elephant in the park," a smart person might say, "Wait, elephants aren't blue." But this AI friend says, "Oh, really? Let me check the news... yes, here is a story about the blue elephant." The AI is so eager to please and so focused on finding something that matches your words that it accepts your lie as truth and builds a fake story around it.

5. The "Different Worlds" Effect

The Finding: If you ask the same news question to two different AI chatbots, they might give you answers based on completely different websites. One might cite a local newspaper, while the other cites a global English site.
The Analogy: It's like two people asking the same question to two different tour guides. One guide shows you the city through the eyes of a local historian; the other shows you the city through the eyes of a travel blogger. You get two different versions of "reality" depending on which bot you choose, and you have no way of knowing which one is actually looking at the real news.

The Bottom Line

The paper concludes that while these AI chatbots are getting very good at news, they are fragile.

  • They look perfect on a test but stumble in real conversation.
  • They treat non-English news (especially Hindi) as a second-class citizen, filtering it through English sources.
  • They are easily tricked by false premises, acting like "yes-men" rather than fact-checkers.
  • Their accuracy depends more on how well their search engine finds the right article than on how smart the AI's brain is.

The authors warn that if we rely on these tools for news without understanding these flaws, we might end up with a world where different people see different versions of the truth, and where local news gets drowned out by a global, English-speaking filter.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →